New Universal-3.6 Pro Realtime is now available Learn more
Insights & Use Cases

LiveKit voice agent with AssemblyAI Universal-3 Pro Realtime

Wire Universal-3.6 Pro Realtime into a LiveKit agent as a native streaming STT plugin — setup, agent_context, accuracy benchmarks, and the turn-detection setting that trips people up.

Abstract green cylinder illustration

Written by

Kelsey Foster

Published on

30 September 2026

LiveKit is how much of the voice-agent market ships—and Universal-3.6 Pro Realtime is a native plugin. Here's how to wire them together and avoid the one gotcha.

The model and endpoint

  • Model: Universal-3.6 Pro Realtime, id universal-3-6-pro (the streaming parameter is speech_model, singular).
  • LiveKit plugin: requires livekit-agents 1.8.0+ (1.8.3+ via LiveKit Inference). The plugin still defaults to universal-3-5-pro, so pass model="universal-3-6-pro" explicitly.
  • Endpoint: wss://streaming.assemblyai.com/v3/ws. The old /v2/realtime/ws returns HTTP 410 — don't use it.
  • Audio: PCM16 mono 16kHz; auth is your raw API key (no Bearer prefix on streaming STT).
  • Pricing: $0.45/hr base for streaming, or the managed Voice Agent API at a flat $4.50/hr if you'd rather not assemble STT + LLM + TTS yourself.

Streaming setup in Python

The core streaming loop looks like this — it's the same speech layer the LiveKit plugin runs under the hood:

# pip install "assemblyai>=1.0.0"
import os
from assemblyai.streaming.v3 import (
    RealTimeEvents, RealTimeParameters, RealTimeSessionParameters,
    RealTimeTranscriber, RealTimeTranscriberOptions, TurnEvent,
)

def on_turn(_, event: TurnEvent):
    tag = "FINAL" if event.end_of_turn else "partial"
    print(f"{tag}: {event.transcript}")

client = RealTimeTranscriber(
    RealTimeTranscriberOptions(),
    api_key=os.environ["ASSEMBLYAI_API_KEY"],
)
client.on(RealTimeEvents.Turn, on_turn)
client.connect(RealTimeParameters(
    sample_rate=16000,
    speech_model="universal-3-6-pro",
    agent_context="What's your email address?",
))
# feed 16kHz mono PCM16 chunks (50-1000ms) via client.stream(chunk)
# after each agent reply, update the context:
client.set_params(RealTimeSessionParameters(agent_context="Sure, what date would you like to book?"))
client.disconnect(terminate=True)  # ALWAYS terminate
Drop It Into Your LiveKit Agent

Universal-3.6 Pro Realtime is a native streaming STT plugin in LiveKit Agents. Get a free API key and wire it into your agent in minutes.

Sign up free

agent_context / Context Carryover

Universal-3.6 Pro Realtime takes an agent_context parameter — you feed it what your agent just said and it biases recognition toward the expected reply. Across 20,000 voice-agent files, this cut WER by 10.2%. Rolling context is on by default (prior finalized user turns carry over automatically), and LiveKit supports it out of the box, so you get it without predefining key terms.

The turn-detection gotcha (read this)

Do not use turn_detection="stt" with the Pro model. Use LiveKit's English turn-detection model instead. This is the single most common way a LiveKit + AssemblyAI setup goes wrong, so wire it up correctly from the start.

How accurate is it?

On AssemblyAI's English voice-agent benchmark, Universal-3.6 Pro Realtime posts 5.19% WER versus Deepgram Flux EN at 13.50%. Independently, it posts a 0.96% pooled semantic WER on Pipecat's open STT benchmark, with the final transcript landing a median 307 ms after the speaker stops, and the lowest WER on Coval's STT benchmark (2.2%, 7-day window). Full numbers are in our Universal-3.6 Pro Realtime research post.

Diarization and deployment

Universal-3.6 Pro Realtime does streaming diarization with revision for up to 10 speakers — useful for multi-party rooms. For deployment, the plugin runs anywhere your LiveKit agent runs; migrate any host-specific deploy steps (e.g., a Fly.io walkthrough) from your existing tutorial content here and update model ids to universal-3-6-pro.

Build it

Talk to a live agent to hear it in action, or get your free API key and drop Universal-3.6 Pro Realtime into your LiveKit agent.

Ship Your Voice Agent on LiveKit

Streaming at $0.45/hr base, or the managed Voice Agent API at a flat $4.50/hr. Grab a free API key and get Context Carryover out of the box.

Sign up free

Frequently asked questions

Does AssemblyAI work with LiveKit out of the box?

Yes. AssemblyAI is a native streaming STT plugin in LiveKit Agents, running Universal-3.6 Pro Realtime when you pass model="universal-3-6-pro", and it supports Context Carryover out of the box — no custom integration required.

Which turn-detection setting should I use?

Use LiveKit's English turn-detection model. Do not set turn_detection="stt" with the Pro model — that combination is not the recommended path.

Which model and endpoint does the plugin use?

The LiveKit plugin still defaults to universal-3-5-pro; pass model="universal-3-6-pro" (livekit-agents 1.8.0+) to run Universal-3.6 Pro Realtime (speech_model=universal-3-6-pro) over wss://streaming.assemblyai.com/v3/ws. Streaming is $0.45/hr base; the managed Voice Agent API is a flat $4.50/hr.

How do I pass conversation context?

Use agent_context: seed it at connection time and update it after each agent reply. It cut WER by 10.2% across 20,000 voice-agent files, and prior finalized user turns carry over automatically.

How accurate is it for voice agents?

5.19% WER on AssemblyAI's English voice-agent benchmark (versus Deepgram Flux EN at 13.50%), a 0.96% pooled semantic WER on Pipecat's open STT benchmark, and the lowest WER on Coval's independent STT benchmark at 2.2%.