New Universal-3.6 Pro Realtime is now available Learn more
Insights & Use Cases

Migrating from OpenAI Realtime API to AssemblyAI Voice Agent API

OpenAI Realtime API migration guide: learn how to move to AssemblyAI Voice Agent API with less session code, simpler audio streaming, and tool support.

Abstract green cone illustration

Written by

Kelsey Foster

Published on

22 September 2026

If you are running a voice agent on the OpenAI Realtime API and you are considering a move, here is the concrete result. Most of your agent’s configuration stops living in your client code and moves server-side into a stored agent you bind by ID. The local state machine you maintain to mirror the service’s session state goes away. Your billing goes from token accounting across a variable-length conversation to one flat $4.50/hr, billed per second. And you keep a WebSocket-based architecture — wss://agents.assemblyai.com/v1/ws — so the shape of your application does not change.

This guide walks the migration in the order you will actually do it: what maps to what, the real event flow, the code before and after, what you delete, how the pricing math changes, and where the honest limits are.

What actually maps to what

Both APIs solve the same problem: a bidirectional WebSocket carrying audio in and audio out, with a language model in the middle and tool calls available mid-conversation. The differences are in where configuration lives and how much orchestration is yours.

Concern OpenAI Realtime API AssemblyAI Voice Agent API
Connection One WebSocket to OpenAI’s realtime endpoint One WebSocket: wss://agents.assemblyai.com/v1/ws
Auth Bearer token Authorization: Bearer <API_KEY>; short-lived token in the browser
Where config lives In your client code, sent per session Server-side stored agent bound by agent_id, or inline if you prefer
Speech recognition Bundled, not separately tunable Universal-3.5 Pro Realtime, with prompt, keyterms and Voice Focus exposed
Language model OpenAI model, bundled Voice Agent LLM bundled, or bring your own over an OpenAI-compatible endpoint
Speech synthesis Bundled Voice Agent TTS, bundled
Audio format Base64 PCM Base64 PCM16 mono 24 kHz, or 8 kHz G.711 μ-law/A-law for telephony
Turn detection You configure thresholds Adaptive by default, and entity-aware — nothing to tune
Billing model Token-based, varies with conversation length and modality Flat $4.50/hr ($0.075/min), billed per second
Reconnection Your responsibility session.resume within a 30-second window, context preserved
End-to-end latency Not published as a single figure ~1 second, caller stops speaking to agent starts speaking

The row that determines how much code you delete is the configuration row. Everything else is a rename.

The biggest structural change: configuration moves server-side

On a realtime API where configuration travels with every session, your client holds the system prompt, the tool schemas, the voice selection and the turn-detection thresholds. That means your client also holds whatever secrets those tools need, and changing your agent’s behavior means shipping a release.

AssemblyAI’s documented recommendation is the opposite. From the session-configuration reference, verbatim: “Most agents should create a stored agent instead.” You POST the agent once to https://agents.assemblyai.com/v1/agents, and your client binds it by ID:

{
  "type": "session.update",
  "session": { "agent_id": "7ad24396-b822-4dca-871a-be9cc4781cf9" }
}

agent_id is valid only on the first session.update and is mutually exclusive with the inline fields — sending both returns agent_id_not_first.

Three things follow, and they are the practical case for the migration:

  • Secrets leave the client. Tools that hit an HTTP API can run server-side, so AssemblyAI makes the request and your client never handles the round trip or the credential.
  • One agent, many transports. The same agent_id serves your API integration, a browser client and a Twilio phone number. Twilio connects over SIP, so there is no media server or webhook to run.
  • Behavior changes without a deploy. Update the stored agent and the next call picks it up.

Inline configuration over session.update remains available and is the right call for genuinely dynamic or one-off agents. The field set is identical either way.

The real event flow

Before the code, the sequence you are porting to:

connect
  → session.update                 (config, or just agent_id)
  ← session.ready                  (save session_id; only now send audio)
  → input.audio  (stream)          (base64 PCM16 mono 24 kHz)

  ← input.speech.started
  ← transcript.user.delta          (full text so far - replace, do not concatenate)
  ← input.speech.stopped
  ← transcript.user                (final)

  ← reply.started
  ← reply.audio                    (base64 chunk in `data`)
  ← transcript.agent.delta
  ← transcript.agent
  ← reply.done                     (status: completed | interrupted)

tool flow:
  ← tool.call                      (arguments already a dict)
  ← reply.done                     ← send tool.result here
  → tool.result                    (result must be a JSON *string*)

teardown:
  → session.end
  ← session.ended

The full surface is larger than this path — seven client-to-server events and fourteen server-to-client — but the events above are what a working agent needs. Barge-in handling, session resume, mid-call reconfiguration and injected conversation context are additions on top of a loop that already runs.

Get An API Key Before You Port

The free tier is 185 hours of pre-recorded transcription plus 333 hours of streaming — enough to build the minimum loop and shadow a slice of real calls before you commit to a cutover.

Sign up free

The code: before and after

Before — the shape of a Realtime API integration

A typical OpenAI Realtime integration in Python looks structurally like this. The specific event names are in OpenAI’s documentation; what matters for the comparison is the shape.

import asyncio
import json
import os

import websockets

OPENAI_KEY = os.environ["OPENAI_API_KEY"]

# Local mirror of server-side session state. This exists only because
# the client has to reconstruct what the service already knows.
state = {
    "response_active": False,
    "audio_committed": False,
    "pending_tool_calls": {},
    "current_item_id": None,
    "interrupted": False,
}


async def handle_events(ws):
    async for message in ws:
        event = json.loads(message)
        event_type = event["type"]

        # A production dispatch table for this API is long. Most
        # branches log and return; a handful mutate `state` above;
        # the ones you skipped fire on a bad connection.
        if event_type in LIFECYCLE_EVENTS:
            update_session_state(state, event)
        elif event_type in TRANSCRIPTION_EVENTS:
            await on_transcript(event)
        elif event_type in RESPONSE_EVENTS:
            await on_model_output(state, event)
        elif event_type in TOOL_EVENTS:
            await dispatch_tool(state, event, ws)
        elif event_type in ERROR_EVENTS:
            await on_error(event)
        else:
            pass        # the branch that eventually becomes an incident


async def main():
    headers = {"Authorization": f"Bearer {OPENAI_KEY}"}
    async with websockets.connect(REALTIME_URL, additional_headers=headers) as ws:
        # Prompt, tools, voice, turn detection - all shipped from the client,
        # on every session.
        await ws.send(json.dumps(SESSION_UPDATE_PAYLOAD))
        await asyncio.gather(send_audio(ws), handle_events(ws))

Note what the state dictionary is for. It is not application state — it is a shadow copy of the service’s state, maintained by you and wrong whenever an event you did not handle changes something on the server. Most hard bugs in a realtime voice agent are a divergence between that mirror and reality.

After — the same integration on the Voice Agent API

import asyncio
import base64
import json
import os

import websockets

API_KEY = os.environ["ASSEMBLYAI_API_KEY"]
AGENTS_WS_URL = "wss://agents.assemblyai.com/v1/ws"

# Everything about the agent lives server-side. This is the whole config.
SESSION_UPDATE = {
    "type": "session.update",
    "session": {"agent_id": os.environ["AGENT_ID"]},
}


async def send_audio(ws, mic_queue):
    """Base64 PCM16 mono 24 kHz. Real time, not faster - the server
    drops frames beyond ~1s of audio per second of wall clock."""
    while True:
        chunk = await mic_queue.get()                 # ~50 ms of int16 PCM
        await ws.send(json.dumps({
            "type": "input.audio",
            "audio": base64.b64encode(chunk).decode(),
        }))


async def run(ws, speaker, mic_queue, tools):
    session_id = None
    last_event = None
    pending_tools = []

    async def flush_if_idle():
        # tool.result goes out only when reply.done is the latest event.
        if last_event != "reply.done" or not pending_tools:
            return
        for tool in pending_tools:
            await ws.send(json.dumps({
                "type": "tool.result",
                "call_id": tool["call_id"],
                "result": json.dumps(tool["result"]),   # a JSON *string*
            }))
        pending_tools.clear()

    async for raw in ws:
        event = json.loads(raw)
        etype = event.get("type")

        if etype == "session.ready":
            session_id = event["session_id"]           # keep for session.resume
            asyncio.create_task(send_audio(ws, mic_queue))

        elif etype == "transcript.user":
            log_user(event["text"])

        elif etype == "reply.audio":
            speaker.write(base64.b64decode(event["data"]))

        elif etype == "transcript.agent":
            log_agent(event["text"])

        elif etype == "tool.call":
            result = await tools[event["name"]](**event["arguments"])
            pending_tools.append({"call_id": event["call_id"], "result": result})
            await flush_if_idle()

        elif etype in ("reply.started", "input.speech.started"):
            last_event = etype
            if etype == "input.speech.started":
                speaker.abort(); speaker.start()       # flush on barge-in

        elif etype == "reply.done":
            last_event = etype
            if event.get("status") == "interrupted":
                pending_tools.clear()                  # agent moved on
                speaker.abort(); speaker.start()
            else:
                await flush_if_idle()

        elif etype == "session.error":
            on_error(event["code"], event["message"])

    return session_id


async def main():
    headers = {"Authorization": f"Bearer {API_KEY}"}
    async with websockets.connect(AGENTS_WS_URL, additional_headers=headers) as ws:
        await ws.send(json.dumps(SESSION_UPDATE))
        try:
            await run(ws, speaker, mic_queue, TOOLS)
        finally:
            # Closing without this keeps the session - and billing -
            # alive for a 30 second resume window.
            await ws.send(json.dumps({"type": "session.end"}))


asyncio.run(main())

What is missing is the point. There is no state dictionary, because there is no session state you need to mirror. And there is no configuration payload, because the configuration is a stored agent.

If you would rather configure inline

The same fields, sent over the socket instead of stored:

{
  "type": "session.update",
  "session": {
    "system_prompt": "You are a support agent for an online retailer. Keep every reply to one or two short sentences. Lead with the answer.",
    "greeting": "Hey, what can I do for you?",
    "tools": [
      {
        "type": "function",
        "name": "lookup_order",
        "description": "Get status and ship date for an order. Use whenever the caller gives an order number. Returns status (string) and ships_on (ISO-8601 date).",
        "parameters": {
          "type": "object",
          "properties": {
            "order_number": { "type": "string", "description": "Digits only, e.g. 41820933." }
          },
          "required": ["order_number"]
        },
        "execution_mode": "interactive",
        "timeout_seconds": 120
      }
    ],
    "input": {
      "format": { "encoding": "audio/pcm" },
      "transcription_mode": "balanced",
      "transcription_prompt": "A customer service call about order status. Expect order numbers read as digit strings.",
      "keyterms": ["expedited", "backorder"],
      "voice_focus": "near-field",
      "voice_focus_threshold": 0.85,
      "turn_detection": {
        "interrupt_response": true
      }
    },
    "output": { "voice": "alba", "format": { "encoding": "audio/pcm" }, "volume": 100 }
  }
}

Five details that generated code gets wrong. Tool definitions carry "type": "function". voice_focus takes near-field or far-field with a hyphen, not an underscore, and its threshold defaults to 0.85 on this API (0.7 on streaming speech-to-text — different surfaces). greeting, output.voice and output.format become immutable once session.ready arrives, while system_prompt, tools, keyterms, turn_detection, transcription_mode, transcription_prompt and output.volume stay mutable mid-call. session.tools updates replace the array rather than merging into it.

And the fifth, which is the one migrators reach for first: do not port your silence thresholds across. min_silence and max_silence are deliberately absent from the example above. Left unset they are adaptive, and the agent slows down on its own when it is collecting a value your tools need — a phone number, an email address, an order number. The docs are explicit that setting either field turns that adaptive, entity-aware behaviour off for the rest of the session, and that the way to improve turn-taking is better tool descriptions rather than VAD knobs. If you arrive from a stack where you had to hand-tune those windows, this is one more thing you get to delete rather than translate.

Tool calls

Your tool implementations port unchanged. Only the envelope and the timing differ, and the timing is the part worth reading twice.

async def lookup_order(order_number: str) -> dict:
    """Your existing tool. Unchanged by the migration."""
    row = await db.fetch_order(order_number)
    if row is None:
        return {
            "error": (
                f"No order found for {order_number}. The customer lookup "
                "succeeded. Ask the caller to re-read just the order number."
            )
        }
    return {"status": row.status, "ships_on": row.ships_on.isoformat()}

Three protocol facts. tool.call delivers arguments already parsed as a dict — no json.loads. tool.result takes the original call_id and a result that must be a JSON string, not an object. And you send it when reply.done is the latest event you have received, which is why the example above accumulates into pending_tools and drains in the reply.done handler — your tool may well return mid-turn.

Worth designing deliberately: the error text. The model reads it verbatim. “Lookup failed” makes the agent re-ask for everything. Naming the failing field, saying what did succeed, and telling the agent what to ask for next gets a clean recovery — which is why the stub above returns a sentence rather than a code. Good tool descriptions do double duty here: they are also what the turn detector uses to decide how long to wait while a caller reads out a value.

Reconnection, which you previously owned

Sessions survive a dropped connection for 30 seconds. Save session_id from session.ready, and on reconnect send session.resume as your first message instead of session.update:

if session_id:
    await ws.send(json.dumps({"type": "session.resume", "session_id": session_id}))
else:
    await ws.send(json.dumps(SESSION_UPDATE))

If the window has expired you get a session.error with session_not_found or session_forbidden; clear the saved ID and start fresh. Conversation context is preserved across the resume, which is the part you would otherwise have rebuilt yourself.

Run Your Own Audio Through It First

Before you port the loop, hear how Universal-3.5 Pro Realtime handles the order numbers, names and addresses your agent actually has to get right.

Try playground

What the flat $4.50/hr actually removes

Token-based billing on a voice agent is harder to model than token-based billing on anything else, because a conversation’s token count is a function of how long the caller talks, how much history you carry, how chatty the model is, and how many tool round trips a turn takes — none of which you control precisely. Teams end up building cost dashboards to answer a question that should be arithmetic.

The Voice Agent API is $4.50/hr, or $0.075/min, billed per second. Verbatim from our pricing page: “Every feature is included in the $4.50/hr rate. There are no per-layer add-ons, concurrency fees, or per-agent subscriptions.”

Read that as three specific removals:

  • No per-layer add-ons. Speech-to-text, the LLM and text-to-speech are in the number. You are not summing three vendors’ rates, and you are not discovering at month end that one layer scaled differently from the others.
  • No concurrency fees. Concurrent sessions do not carry a surcharge. A Monday-morning spike costs the same per minute as a Sunday-night trickle.
  • No per-agent subscriptions. Ten agents and one agent cost the same per minute of conversation. You can deploy a narrow agent for a single workflow without a seat-cost argument.

There are no minimums and no commitment. What a call costs is minutes multiplied by $0.075. Full detail is on the pricing page.

Two things to put in your migration checklist. First, because billing is per second rather than per token, carrying full conversation history on every turn — usually the right call for quality, and expensive under token billing — costs nothing extra here. Second, the opposite trap: if you close the WebSocket without sending session.end, the server holds the session for a billable 30-second resume window. Send session.end on every intentional hangup.

What you gain on accuracy

The speech layer is Universal-3.5 Pro Realtime, and unlike a fully bundled realtime API, its inputs are exposed to you. session.input.transcription_prompt takes up to 1750 characters describing the situation. session.input.keyterms takes up to 100 exact strings to bias toward — pair it with any lookup tool you ship, so the transcript does not mangle a company or product name before your tool ever sees it. Both are mutable mid-call, so you can narrow them as the conversation narrows.

On the Pipecat open STT benchmark, run against real agent conversations rather than read audiobooks, lower is better:

Metric Universal-3.5 Pro Realtime Deepgram Flux ElevenLabs Scribe v2 Google Chirp3
Word error rate 6.99% 15.58% 9.76% 9.04%
Entity error rate 15.31% 50.50% 39.70% 21.51%
Names 16.92% 39.21% 38.03% 22.10%
Places 6.28% 14.86% 34.06% 10.04%
Phone numbers 3.55% 10.41% 4.78% 4.95%

If your agent takes an order number, an address or a callback number, the entity rows are the ones to read. Aggregate word error rate tells you how the transcript reads; entity error rate tells you whether the transaction completed.

On short-form English audio, Universal-3.5 Pro posts a 3.87% mean normalized word error rate, first against eight competitors — relevant here because voice agent turns are, almost by definition, short-form audio. Methodology for both is at assemblyai.com/benchmarks.

What you give up, and when not to migrate

Being straight about this is more useful than a feature table.

Streaming language coverage caps at 18 languages. English, Spanish, French, German, Italian, Portuguese, Arabic, Danish, Dutch, Finnish, Hebrew, Hindi, Japanese, Mandarin, Norwegian, Swedish, Turkish and Vietnamese, with native code-switching between them. If your agent needs a language outside that set in real time, this is a blocker and no configuration works around it.

You are choosing a bundle by default. Three models behind one connection is the reason the price is one number. If your agent depends on a particular model’s behavior, you can point the agent at an OpenAI-compatible endpoint of your own — but evaluate that path before you commit rather than after.

If you only need transcription, this is the wrong product. The Voice Agent API bundles reasoning and synthesis. If you already have those, use Universal-3.5 Pro Realtime directly over wss://streaming.assemblyai.com/v3/ws at $0.45/hr base. Same model, a fraction of the price, and you keep your orchestration.

If your audio is short and bounded, you may not need streaming at all. Push-to-talk, voice commands, IVR responses and voicemail are better served by the Sync API, which returns a finished transcript in about 134 ms at p50 from a single HTTP request.

A migration plan that does not require a big-bang cutover

  1. Inventory your event handlers. Mark which branches contain real logic and which only maintain the state mirror. The second group is what you delete, and in most integrations it is the majority.
  2. Create a stored agent first. Move the system prompt, voice, greeting and tool schemas out of your client and into a POST /v1/agents body before you touch the socket code. This is the change that shrinks everything downstream.
  3. Port the tool implementations. They are unchanged. Only the envelope and the send timing differ. Spend the time you save on the tool descriptions — they drive turn-taking here.
  4. Build the minimum loop against the docs. session.update → session.ready → input.audio → transcript.user / reply.audio / reply.done. Get one full turn working before adding barge-in, resume or mid-call reconfiguration.
  5. Do not port your silence thresholds. Leave min_silence and max_silence unset and let the adaptive default handle entity capture. Only reach for them if you have measured a specific problem the default causes.
  6. Re-tune the system prompt for voice. Prompts rarely transfer verbatim between models. The most common adjustment is response length — a model that writes two paragraphs is unusable out loud. Drop markdown entirely; TTS reads asterisks literally.
  7. Set transcription_prompt and keyterms before you benchmark. Comparing a tuned incumbent against an untuned replacement gives you the wrong answer.
  8. Shadow a slice of traffic. Run both agents against the same calls, compare transcripts on the entities that matter to your business, and measure end-to-end turn latency the same way on both.
  9. Cut over by segment. One workflow, one region, or one hour of the day. There is no commitment, so a partial deployment costs nothing to keep running.

The free tier — 185 hours of pre-recorded transcription plus 333 hours of streaming — is enough to shadow a meaningful slice before deciding anything.

The build guide walks through a complete agent, and the launch post covers the design decisions. The two pages to keep open during the port are the events reference and session configuration.

Migrating Under Constraints

Self-hosted deployment, EU data residency, or a Business Associate Addendum for PHI — talk it through with someone who has run these cutovers before you plan yours.

Talk to AI expert

Frequently asked questions

How hard is it to migrate from the OpenAI Realtime API to the AssemblyAI Voice Agent API?

Most of the work is deletion. Both APIs are bidirectional WebSockets carrying audio in and audio out, so your architecture does not change — you swap the endpoint to wss://agents.assemblyai.com/v1/ws, move your prompt, tools and voice into a server-side stored agent, and drop the local state mirror your event handlers existed to maintain. Tool implementations port unchanged, and your turn-detection thresholds do not need porting at all.

What does the AssemblyAI Voice Agent API cost compared to the OpenAI Realtime API?

A flat $4.50/hr, or $0.075/min, billed per second, versus token-based billing that varies with conversation length, history size and modality. From our pricing page: “Every feature is included in the $4.50/hr rate. There are no per-layer add-ons, concurrency fees, or per-agent subscriptions.” One thing to watch during migration: closing the socket without sending session.end leaves a billable 30-second resume window on every call.

Which speech recognition model does the Voice Agent API use?

Universal-3.5 Pro Realtime, bundled with the Voice Agent LLM and Voice Agent TTS behind the single agents WebSocket. It posts a 6.99% word error rate and a 15.31% entity error rate on the Pipecat open STT benchmark against real agent conversations, and Universal-3.5 Pro posts a 3.87% mean normalized WER on short-form English audio, first against eight competitors.

Do I have to rewrite my tool calls?

No — the tool functions themselves are unchanged, but the envelope and the timing differ. tool.call hands you arguments already parsed as a dict, and your tool.result must echo the original call_id with result as a JSON string, sent when reply.done is the latest event you have received. Return specific error text rather than a code, because the model reads it verbatim when deciding how to recover.

Should I port my turn-detection settings across?

No. min_silence and max_silence are adaptive when left unset, and the agent automatically slows down while a caller is reading out a phone number, an email address or an order number. Setting either field turns that behaviour off for the rest of the session, so porting thresholds from a stack that required them makes your agent worse rather than better. The documented way to improve turn-taking here is to write better tool descriptions.

What if I only need transcription, not a full agent?

Use Universal-3.5 Pro Realtime directly over wss://streaming.assemblyai.com/v3/ws at $0.45/hr base — the same model, without the bundled LLM and TTS you would not be using. If your audio is short and bounded rather than open-ended, the Sync API returns a finished transcript in about 134 ms at p50 from a single HTTP request, versus 5 to 6 seconds through async submit-and-poll.