There's a particular kind of engineering meeting that happens right after a team decides to put an LLM on top of their audio. Someone opens a doc. Someone else starts listing providers. And within twenty minutes the conversation has stopped being about the product and started being about API keys, rate limits, four separate billing relationships, and which provider's outage page you're supposed to check at 2am.
That meeting is the actual cost of building voice AI apps. Not the transcription. Not the prompt. The plumbing.
This post is about removing that plumbing — and about what you build once it's gone. We'll cover the categories of voice AI applications that are generating real returns, the two architectures you can choose between, and then the part the earlier version of this post was missing entirely: working code, named parameters, and what it costs per million tokens.
Voice AI applications that are already working
Voice AI isn't speculative anymore. AssemblyAI processes more than 600 million inference calls a month across 1,500 corporate customers, and we crossed one million developers on August 21, 2026. That's not a projection about a market in 2030. That's throughput.
What are all those calls doing? Broadly, five things.
Meeting intelligence
Notetakers and meeting platforms take a recording, produce a transcript with speaker labels, and then turn that transcript into something a human would have written: a summary, a decision log, a list of action items with owners. The hard part was never the summary. It's getting the transcript accurate enough that the summary is worth reading — particularly on names, company names, and the acronyms every team invents for itself.
Granola builds in this category. As Jonathan Kim, Software Engineer at Granola, put it:
"The speed difference is immediately noticeable — our users see their conversations transcribed almost instantaneously. It feels so much more responsive than what we were using before."
Sales intelligence
Sales platforms analyze calls for objections, competitor mentions, talk-time ratios, and next steps, then push structured output into a CRM. The LLM layer here is doing extraction, not summarization — pulling specific fields out of an unstructured conversation and returning them in a shape a database can accept.
Siro does exactly this for field sales teams, and their co-founder and CEO Jake Cronin describes the reaction:
"On 10 out of 10 onboarding calls, our customers are at some point telling us 'wow that insight was crisp' — and that's because of the accuracy we're getting from AssemblyAI."
Siro reports a 36% improvement in close rate and a 90% reduction in customer complaints and support tickets.
Customer service and contact centers
Contact centers run the highest volume of any category on this list, and they run it against the tightest compliance requirements. The workload is QA scoring, sentiment tracking, compliance checking, and increasingly real-time agent assist — surfacing an answer to the agent while the customer is still talking. Calabrio operates here at enterprise scale.
Healthcare documentation
Ambient scribes listen to a clinical encounter and produce a structured note. This category has the least tolerance for entity errors of anything on the list, because a wrong medication name or a wrong dosage is not a transcription defect — it's a patient safety event. Commure runs an ambient product on AssemblyAI, and their tech lead Gautam Pradeep describes the trajectory:
"We've integrated the newest models from AssemblyAI for pre-recorded audio ASR in our ambient product, and it's been excellent. We're now exploring Universal-3.5 Pro for async and realtime speech-to-text capabilities for new use cases. What's been just as important is the reliability of the platform itself—both technically and in terms of partnership."
If you're building in this category: AssemblyAI enables covered entities and their business associates subject to HIPAA to use the AssemblyAI services to process protected health information (PHI). AssemblyAI is considered a business associate under HIPAA, and we offer a standard Business Associate Addendum (BAA) that is required under HIPAA to ensure that AssemblyAI appropriately safeguards PHI. The BAA can be signed in minutes without a sales call. Medical Mode is a separate capability — activated with domain: "medical-v1" — and it reaches a 3.2% missed entity rate on medical entities.
Real-time voice agents
The newest category and the fastest-growing one. A caller speaks, a model transcribes, an LLM reasons, a voice responds, and the whole loop has to close in about a second or the conversation feels broken. This is a fundamentally different engineering problem from the four above, and it's worth being precise about why. Everything else on this list is a pipeline you can retry. A voice agent is a live conversation you can't.
Evaluate real-time speech-to-text with low latency and strong accuracy. Launch pilots quickly with clear docs and developer-friendly APIs.
Two architectures, and how to pick
Every voice AI app resolves to one of two shapes.
Cascade. Speech-to-text, then an LLM, then optionally text-to-speech. Three discrete stages, each independently swappable, each independently observable. You can log the transcript. You can inspect the prompt. You can see exactly which stage produced a bad answer.
Speech-to-speech. A single multimodal model takes audio in and emits audio out with no text intermediate. It's elegant, it preserves prosody, and it gives you almost nothing to debug. When it gets a phone number wrong, there is no transcript to look at.
Here's my position, and it's not a neutral one: for anything with a compliance requirement, an audit trail, or a support team that has to explain what happened on a call, cascade wins and it isn't close. You need the transcript. Not as an artifact — as infrastructure. It's what you search, what you store, what you show a regulator, and what you feed back into evaluation when something goes wrong. Speech-to-speech is genuinely better for open-ended, low-stakes, emotionally expressive interaction. That's a real category. It's just a smaller one than the discourse suggests.
We go deeper on the tradeoff in our breakdown of voice agent architectures.
The three surfaces you'll actually call
AssemblyAI exposes three products, and the useful way to think about them is by what owns the clock.
Speech-to-Text API — you own the clock
Submit a file, get a transcript back. Universal-3.5 Pro is the async flagship at $0.21/hr, and it became the default async model on September 2, 2026. It runs 18 languages with native code-switching — English, Spanish, French, German, Italian, Portuguese, Arabic, Danish, Dutch, Finnish, Hebrew, Hindi, Japanese, Chinese, Norwegian, Swedish, Turkish, and Vietnamese — and anything outside those 18 falls back automatically to Universal-2 for 99 languages total.
It uses an LLM-based decoder, which is why it holds up on the things that break conventional acoustic models: proper nouns, account numbers, drug names, product SKUs. Its hallucination rate runs about 30% lower than Whisper's. The full methodology is on our benchmarks page.
LLM Gateway — one key, many models
LLM Gateway is a single API across Anthropic, OpenAI, Google, and Qwen. One key, one bill, one integration, with automatic cross-provider fallback, streaming with tool calling, structured JSON output, and prompt caching.
The live pricing page lists 33 models: 9 from Anthropic, 12 from OpenAI, 9 from Google, and 3 from Qwen. That number matters more than it looks. It means you can start a project on Gemini 2.5 Flash Lite because it's cheap, discover your extraction task needs stronger reasoning, and move to Claude Opus by changing a string — without a new contract, a new key, a new SDK, or a new line item on the invoice.
Voice Agent API — the clock owns you
One WebSocket at wss://agents.assemblyai.com/v1/ws that bundles Universal-3.6 Pro Realtime for transcription, a Voice Agent LLM tuned for spoken conversation rather than text chat, and Voice Agent TTS. Flat $4.50/hr, billed per second on connected conversation time, roughly one second end to end.
Every feature is included in the $4.50/hr rate. There are no per-layer add-ons, concurrency fees, or per-agent subscriptions.
Code: transcribe, then analyze
Enough architecture. Here's the actual pattern, in about thirty lines.
Step one is the transcript. Set the model explicitly — don't rely on the account default, because defaults change and your reproducibility shouldn't depend on ours.
import assemblyai as aai
aai.settings.api_key = "YOUR_API_KEY"
config = aai.TranscriptionConfig(
speech_models=["universal-3-5-pro"],
speaker_labels=True,
)
transcript = aai.Transcriber(config=config).transcribe(
"https://storage.googleapis.com/aai-web-samples/meeting.mp4"
)
if transcript.status == "error":
raise RuntimeError(transcript.error)
print(f"Model used: {transcript.speech_model_used}")
# Build a speaker-attributed transcript for the LLM
dialogue = "\n".join(
f"Speaker {u.speaker}: {u.text}" for u in transcript.utterances
)Note speech_models — plural, an array. That's the async parameter. Streaming uses speech_model, singular. Getting this backwards is the most common integration mistake we see, and the response field speech_model_used tells you which model actually ran, which is worth logging in production.
Step two is the analysis. LLM Gateway is OpenAI-compatible, so if you already have the OpenAI client installed you're one base_url away from working.
from openai import OpenAI
client = OpenAI(
base_url="https://llm-gateway.assemblyai.com/v1",
api_key="YOUR_API_KEY", # the same AssemblyAI key
)
response = client.chat.completions.create(
model="qwen3.5-4b-32k-fast", # swap for any of the 33 models on /pricing
messages=[
{
"role": "system",
"content": (
"You analyze meeting transcripts. Be specific. "
"Never invent a name or a number that isn't in the transcript."
),
},
{
"role": "user",
"content": (
"List every action item in this meeting with its owner "
"and any stated deadline.\n\n" + dialogue
),
},
],
max_tokens=1000,
)
print(response.choices[0].message.content)That's the whole pattern. Audio in, structured insight out, two API calls, one key.
The same call in cURL, if you'd rather see the wire format:
curl -X POST https://llm-gateway.assemblyai.com/v1/chat/completions \
-H "authorization: YOUR_API_KEY" \
-H "content-type: application/json" \
-d '{
"model": "qwen3.5-4b-32k-fast",
"messages": [
{"role": "user", "content": "Summarize the decisions in this call: ..."}
],
"max_tokens": 1000
}'EU customers swap the host for https://llm-gateway.eu.assemblyai.com/v1. Same request shape, data stays in region. The EU endpoint serves Anthropic Claude and Google Gemini; OpenAI models are US-only, so check coverage before you switch hosts. Full setup is in the LLM Gateway quickstart.
Structured extraction, not prose
Summaries are the demo. Structured fields are the product. If the output has to land in a database, ask for JSON and enforce the shape in the prompt:
schema_prompt = """Return ONLY valid JSON matching this shape:
{
"competitors_mentioned": [string],
"objections": [{"objection": string, "speaker": string}],
"next_step": string | null,
"committed_date": string | null
}
Use null when the transcript does not state a value. Do not guess."""
response = client.chat.completions.create(
model="qwen3.5-4b-32k-fast",
messages=[
{"role": "system", "content": schema_prompt},
{"role": "user", "content": dialogue},
],
max_tokens=800,
)One thing to know before you ship this: qwen3.5-4b-32k-fast does not support response_format, so the JSON shape here is enforced by the prompt alone, not by the API. Parse defensively and handle the case where the model returns prose. If you need schema-guaranteed output, check /pricing for a model that supports structured outputs and swap the model string.
The "use null, do not guess" instruction earns its place. Extraction models are agreeable by default, and an agreeable model will invent a deadline rather than admit the call didn't have one.
What it costs
LLM Gateway is billed per million input and output tokens, at the provider's rate, on your AssemblyAI invoice. There is no gateway markup line item and no separate provider relationship to manage.
One rule you need before you read the table: listed prices are for global routing. In-region US or EU routing costs 10% more, because provider costs are higher in-region. If you want the listed rates, pass model_region as global in the request body. Global routing is currently live for Anthropic Claude models only, with Google Gemini 3 coming, so everything else bills in-region. If you have a data residency requirement that rules out global routing, budget the extra 10%.
A representative slice of the 33 models, as of September 3, 2026:
| Model | Provider | Input / 1M tokens | Output / 1M tokens |
|---|---|---|---|
| GPT-5 Nano | OpenAI | $0.05 | $0.40 |
| Gemini 2.5 Flash Lite | $0.10 | $0.40 | |
| Qwen3.5 4B Fast | Qwen | $0.10 | $0.50 |
| Qwen3 32B | Qwen | $0.15 | $0.60 |
| GPT-5 mini | OpenAI | $0.25 | $2.00 |
| Gemini 2.5 Flash | $0.30 | $2.50 | |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 |
| GPT-5 | OpenAI | $1.25 | $10.00 |
| Gemini 2.5 Pro | $1.25 | $10.00 | |
| Claude Sonnet 4.5 | Anthropic | $3.00 | $15.00 |
| Claude Opus 4.5 | Anthropic | $5.00 | $25.00 |
Transcription is billed separately and by the hour: $0.21/hr for Universal-3.5 Pro async, $0.15/hr for Universal-2, $0.45/hr for Universal-3.6 Pro Realtime, $4.50/hr for the bundled Voice Agent API. Everything is pay-as-you-go, billed per second, with no minimums, no upfront commits, and unlimited concurrency. The free tier gives you 185 hours of pre-recorded transcription and 333 hours of streaming. Full detail lives on the pricing page, and we break down how to model total cost in our guide to speech-to-text API pricing.
A worked example, because "per million tokens" is hard to feel. A 60-minute sales call runs roughly 9,000 words, call it 12,000 tokens. Transcribing it on Universal-3.5 Pro costs $0.21. Running an extraction prompt over it with GPT-5 mini costs about $0.003 in, plus output. The transcription is the expensive half by two orders of magnitude — which is a useful thing to know before you spend a sprint optimizing prompt length.
Experience natural, real-time conversations that go far beyond IVR menus. Test streaming transcription speed and accuracy on your own audio.
Summarization: the surface changed
This section needs a specific correction, because the old way of doing it stopped working.
Legacy auto chapters and legacy summarization were deprecated. If you have code calling auto_chapters or the old boolean summarization parameter, it needs to move. The replacement is Speech Understanding Summarization, and the migration guides are here: summarization migration and auto chapters migration.
Here's the parameter, named plainly, because the previous version of this post said summarization was "enabled with a single parameter" and then never said which one. It's a nested object on the transcription request:
{
"audio_url": "https://storage.googleapis.com/aai-web-samples/meeting.mp4",
"speech_models": ["universal-3-5-pro"],
"speech_understanding": {
"request": {
"summarization": {
"summary_type": "paragraph"
}
}
}
}
summary_type takes "paragraph" or "bullets". There's an optional effort field that takes "low" (the default) or "medium" — medium thinks harder and returns a denser summary. It does not change the price. Summarization is priced à la carte at $0.03 per hour of audio regardless of effort. Full reference in the summarization docs.
So when do you use Speech Understanding Summarization and when do you use LLM Gateway? Simple rule: Speech Understanding when you want a good summary with no prompt engineering, LLM Gateway when the output shape is yours. A support platform that needs every ticket summarized the same way should use Speech Understanding. A platform that needs "summarize this, then classify it against our 14 internal issue categories, then return JSON" needs the gateway.
Question answering over a transcript
The Q&A pattern is the same two calls as above with a different prompt. What makes it work well isn't the model — it's giving the model something to refuse with. Ask an LLM "what did the customer say about pricing?" against a call where pricing never came up, and without an explicit escape hatch it will manufacture something plausible. Add "if the transcript does not address this, say so" and the failure mode disappears.
For real-time equivalents, the mechanics are different again — we cover feeding live conversation state into a model in our post on conversation context for voice agents.
Meeting intelligence: what to build first
If you're deciding where to start, this is roughly how the meeting-intelligence applications sort by effort and payoff.
| Application type | Primary API features | Business impact | Implementation time |
|---|---|---|---|
| Executive briefings | Speech Understanding (Summarization), LLM Gateway | Significant reduction in prep time | 1–2 weeks |
| Project coordination | LLM Gateway (structured extraction) | Faster action item tracking | 2–3 weeks |
| Sales reviews | Speech Understanding, LLM Gateway, speaker diarization | Improved pipeline accuracy | 3–4 weeks |
Who's building this way
A note on how to read this section, because the previous version of this post got it wrong: everyone named below is an AssemblyAI customer. The earlier draft mixed customers with market examples and competitors in the same lists, which made the whole set impossible to trust.
Fireflies runs AssemblyAI in its voice agent pipeline. Foysal Osmany, Software Engineer at Fireflies:
"We were searching for the best realtime ASR model for our voice agent pipeline in Fireflies. The new Universal 3.5 Pro speech model from Assembly is best so far in terms of accuracy, latency and language switching."
Metaview builds recruiting intelligence, where context is the whole game. Shahriar Tajbakhsh, Co-founder and CTO:
"Since moving to AssemblyAI, we've seen a meaningful improvement in the confidence tail of our production transcripts....What stands out is not just the model quality, but the way [they] let us bring real meeting context into transcription, from calendar titles to organizations, domains, and participant names, so recruiting conversations come through with the nuance our customers depend on."
LiveKit makes Universal-3.5 Pro available through LiveKit Inference. David Zhao, Co-founder at LiveKit:
"We're excited to make AssemblyAI's Universal-3.5 Pro available on LiveKit Inference. What really stands out is their pace of innovation with Context Carryover — it intelligently applies conversation context to improve transcription accuracy in a way most speech models don't, removing the need for users to predefine key terms."
Also building on this stack: Uber, Veed, CallRail, Calabrio, HeyGen, ClickUp, Notta, Gorgias, HoneyBook, Recall.ai, Apollo.io, and AlphaSense. Retell integrates AssemblyAI as its high-accuracy transcription option. Orchestration frameworks including LiveKit, Pipecat, and Vapi support AssemblyAI as a drop-in STT layer, which matters if you already have a voice stack and don't want to rebuild it.
Four things that separate the teams who ship
Pick a use case where the failure mode is visible. "Summarize meetings" is unfalsifiable — a mediocre summary and a great one look similar in a demo. "Extract the committed follow-up date" is testable. You'll know within a week whether it works.
Fix transcription before you tune prompts. This is the most common wasted sprint in voice AI. Teams spend three weeks on prompt engineering to fix an output problem that was a transcription problem the whole time. If the model heard "Metoprolol" as "meta prolyl," no prompt saves you. Check the transcript first. Every time.
Route by task, not by brand loyalty. You have 33 models behind one key. Classification and routing tasks run fine on GPT-5 Nano at $0.05 per million input tokens. Nuanced extraction over a messy 90-minute call is worth Claude Sonnet. The gateway exists so this is a one-line decision rather than a procurement cycle.
Instrument the boundary. Log the transcript, the prompt, the model that actually ran, and the raw response. When something goes wrong in production — and it will — the difference between a ten-minute fix and a two-day investigation is whether you can see which of the two stages broke.
The real advantage isn't the model
Here's the thing most teams figure out about eighteen months in, and it's not what they expected going in.
The models will keep changing. Universal-3.5 Pro replaced Universal-3 Pro as the async default this month. The best extraction model on the pricing table today will not be the best one next quarter. That churn is permanent, and it's good — it's the market working.
What determines whether you can take advantage of it is how many things have to change when the answer changes. If moving from Gemini Flash to Claude Sonnet means a new vendor contract, a new key in your secrets manager, a new SDK, a new billing relationship, and a security review, you won't do it. You'll stay on the model you started with, and eighteen months later you'll be running a stack that was optimal in 2026.
One API and one bill isn't a convenience feature. It's what keeps the option to change your mind cheap enough to keep exercising.
Evaluate real-time speech-to-text with low latency and strong accuracy. Launch pilots quickly with clear docs and developer-friendly APIs.
Frequently asked questions
What is AssemblyAI's LLM Gateway and how does it work?
LLM Gateway is a single API that gives you access to 33 large language models from four providers using one AssemblyAI key and one bill. It's OpenAI-compatible: you point a standard OpenAI client at https://llm-gateway.assemblyai.com/v1, pass your AssemblyAI API key, and change the model string to switch providers. It adds automatic cross-provider fallback, streaming with tool calling, structured JSON output, and prompt caching on top of the underlying models. EU customers use https://llm-gateway.eu.assemblyai.com/v1 for in-region processing.
Which LLM providers does AssemblyAI's LLM Gateway support?
Four: Anthropic, OpenAI, Google, and Qwen, totaling 33 models on the live pricing page. The breakdown is 9 Anthropic models (Claude Haiku 4.5 through the Opus line), 12 OpenAI models (GPT-5 Nano through GPT-5.5), 9 Google models (the Gemini Flash Lite, Flash, and Pro families), and 3 Qwen models. Switching between them is a change to the model field — no new key, no new contract, no new SDK.
How does AssemblyAI's LLM Gateway consolidated billing work across providers?
Every model runs on your single AssemblyAI account and appears on one invoice, billed per million input and output tokens at the provider's rate. The one rule to know: listed prices are for global routing, and in-region US or EU routing costs 10% more. Pass model_region as global in the request body to get the listed rates, which is currently supported on Anthropic Claude models. Transcription is billed separately by the hour — $0.21/hr for Universal-3.5 Pro async — and everything is pay-as-you-go with no minimums or upfront commits.
What are the best speech-to-text APIs with built-in LLM cleanup for dictation apps?
For dictation you want a low-latency transcription surface plus an LLM pass in the same account, which is exactly the Sync API and LLM Gateway pairing. The Sync Speech-to-Text API returns a finished transcript in a single HTTP POST at about 134 ms p50 on a two-second clip, for clips from 80 ms to 2 minutes at $0.45/hr. You then run the raw text through LLM Gateway for punctuation cleanup, formatting, or tone adjustment — one key, one bill, two calls. Note that the Sync API does not support PII redaction, speaker diarization, Speech Understanding, or Medical Mode.
What is the best one-API solution for voice agents?
AssemblyAI's Voice Agent API is a single WebSocket at wss://agents.assemblyai.com/v1/ws that handles transcription, LLM reasoning, and speech synthesis for a flat $4.50/hr billed per second, with roughly one second of end-to-end latency. Every feature is included in that rate — prompting, Voice Focus, advanced turn detection, interruption detection, recordings and transcripts, and BYO-Twilio SIP trunking with no per-minute markup. There are no per-layer add-ons, concurrency fees, or per-agent subscriptions. If you already run LiveKit, Pipecat, or Vapi, the alternative is dropping in Universal-3.6 Pro Realtime as your STT layer at $0.45/hr and keeping the rest of your stack — see our comparison of the best voice agent APIs.
What is the best API for conversation intelligence and call analytics?
The combination you want is accurate transcription with speaker labels, then an LLM layer for extraction — which is Universal-3.5 Pro at $0.21/hr plus LLM Gateway on one key. Universal-3.5 Pro uses an LLM-based decoder built for exactly the entities call analytics depends on: account numbers, product names, and customer surnames. Diarization is our most accurate yet, and Speech Understanding adds sentiment at $0.02/hr, entity detection at $0.08/hr, and topic detection at $0.15/hr à la carte. See conversation intelligence solutions for the full architecture.