New Universal-3.6 Pro Realtime is now available Learn more

What is real-time agent assist? How AI transforms live customer support

Real-Time Agent Assist (RTAA) uses AI to guide customer service agents during live calls with instant knowledge, compliance alerts, and next-best actions. Here's how it works and why it matters.

What is real-time agent assist? How AI transforms live customer support

Written by

Kelsey Foster

Published on

22 September 2026

Real-time agent assist is software that listens to a live call, transcribes it as it happens, and pushes information to the human agent while the conversation is still going — the right knowledge base article, the customer’s account state, a compliance prompt, a next-best-action, a summary of what has been said so far. It is not post-call analytics moved earlier. The defining constraint is that the output has to land while the agent can still act on it, which in practice means the transcript has to be complete and correct within a few hundred milliseconds of the customer finishing a sentence.

That constraint is what makes agent assist an infrastructure problem rather than a UI problem. Most teams that try to build it discover the same three things in the same order: the transcript arrives too late to be useful, the turn boundaries are wrong so the assist fires mid-sentence, and the account numbers and names — the exact tokens the assist logic needs to look anything up — are the tokens the model gets wrong. This guide walks through each of those, with the numbers and the parameters.

What real-time agent assist actually does

Strip away the category marketing and agent assist does four jobs, all of them downstream of a live transcript:

  • Retrieval. The customer describes a problem; the system matches it against documentation, policy, or past tickets and surfaces the answer on the agent’s screen before the agent has to search for it.
  • State lookup. The customer reads out an account number, an order ID, a member ID or a policy number. The system captures it and pulls the record.
  • Guidance and compliance. The system checks whether required disclosures were read, whether a prohibited phrase was used, or whether the agent has missed a step in a script, and prompts accordingly.
  • Summarization and after-call work. The system keeps a running summary so the wrap-up note is largely written by the time the call ends.

Three of those four break immediately if the transcript is late, and all four break if the transcript is wrong about the specific entities in it. That ordering matters for how you spend your engineering budget.

The three layers, and which latency number belongs to which

The single most common source of confusion in this category is that “latency” refers to at least two different measurements, and they get quoted against each other as if they were the same thing. They are not. Be explicit about which layer you are talking about.

Layer 1 — streaming speech-to-text

This is the time from the customer speaking to your application holding a finished, formatted transcript of what they said. On Universal-3.5 Pro Realtime, transcripts are emitted continuously as audio arrives and finalized at the end of each turn. This is the number that governs whether a retrieval or a lookup can fire while the agent is still in the same beat of the conversation — and it is the one you should measure on your own audio rather than take from any vendor’s headline, because the span each vendor reports is rarely the same span.

Layer 2 — your assist logic

Retrieval, a database call, a rules engine, or a model call that turns the transcript into something worth showing. This is your code and your infrastructure, and it is the layer you control most directly. In an agent assist product, this layer does not have to respond to the customer — a human is doing that — so it has more headroom than the equivalent layer in a voice bot.

Layer 3 — a full conversational turn, if a machine is speaking back

If you are building something that replies rather than something that assists a human, you are measuring a different thing: speech in to speech out. The Voice Agent API runs approximately 1 second end-to-end for a complete turn, because that budget covers transcription, the language model, and text-to-speech, all three. It bundles Universal-3.5 Pro Realtime, a Voice Agent LLM and a Voice Agent TTS behind one WebSocket at wss://agents.assemblyai.com/v1/ws.

So: a transcription figure describes one layer. Approximately 1 second is a full-turn figure. An agent assist product that has a human doing the talking only needs layer 1 and layer 2. Quoting the full-turn number as if it described the transcription layer, or the transcription number as if it described a whole conversational loop, is how architecture decisions go wrong.

Turn detection is the part most teams underestimate

An assist that fires halfway through a customer’s sentence is worse than no assist at all, because the agent learns to ignore it. Getting the timing right is a turn-detection problem, and turn detection built on silence thresholds alone will always be either too eager or too slow — a customer pausing to read a number off a card looks identical, in silence terms, to a customer who has finished speaking.

Universal-3.5 Pro Streaming does not endpoint on silence alone. When a speaker pauses, the model evaluates what has been said so far to judge whether the turn reads as complete or as a mid-thought pause; if it reads as complete the turn ends, otherwise a partial is emitted and the turn continues. Formatting is always on, so punctuation and casing are applied to streaming output by default — which is what makes that check possible in the first place.

The parameters you will be tuning:

Parameter What it controls Default
mode Latency and accuracy preset: min_latency, balanced, or max_accuracy. Universal-3.5 Pro Streaming only. It sets the defaults for the two silence fields, so change this before you change anything else. balanced
min_turn_silence Silence in milliseconds before an end-of-turn check runs. Set by mode: 128 (min_latency), 128 (balanced), 512 (max_accuracy)
max_turn_silence Maximum silence before the turn is forced to end regardless of content. Set by mode: 640 / 1280 / 2560
vad_threshold Confidence threshold for classifying a frame as speech. Raise it in noisy environments to reduce false speech detection. 0.2

Enabling speaker_labels swaps in diarization-tuned values — min_turn_silence 640 and max_turn_silence 768, with continuous partials disabled — so re-tune after you turn it on rather than before. Full detail is in the turn detection documentation. For agent assist, balanced is usually the right preset: you are not racing a text-to-speech engine, and the small extra time buys transcript quality that your retrieval layer depends on.

One naming trap worth flagging, because it costs people an afternoon. On the streaming speech-to-text API the preset is mode. On the Voice Agent API the equivalent field is session.input.transcription_mode, with the same three values. Same concept, different key, different product.

A minimal streaming configuration for an agent assist deployment looks like this. Note that streaming takes speech_model in the singular; the pre-recorded API takes speech_models as a list, and mixing them up is a common first-day error.

CONNECTION_PARAMS = {
    "sample_rate": 16000,
    "speech_model": "universal-3-5-pro",
    "mode": "balanced",
    "language_codes": ["en"],
    "voice_focus": "near-field",
    # "voice_focus_threshold": 0.7,  # Optional, defaults to 0.7. Higher is more aggressive.
    "agent_context": "Asking the customer to confirm the account number on their statement.",
}

Two of those fields deserve their own explanation.

voice_focus takes near-field for headsets, handsets and other close-talking microphones, or far-field for conference rooms, drive-thru speakers, laptop mics and other distant capture. Contact centre audio is almost always near-field; in-store, branch or kiosk assist is not. voice_focus_threshold controls how aggressively background audio is suppressed, from 0.0 to 1.0, and defaults to 0.7 for both variants on this surface. It requires voice_focus to be set or the connection returns a validation error. Voice Focus is an add-on at $0.10/hr.

agent_context is the parameter that does the most work in this use case, and it is covered in the section below. Context Carryover — rolling memory across the conversation — is on by default, so the model does not lose the thread of a long call.

If you need to attribute lines to speakers as the call runs, streaming speaker diarization supports up to 10 speakers with revision, correcting attributions within roughly half a second of the stream ending. It is a $0.12/hr add-on. Enabling it changes several turn-detection defaults, so re-tune after you turn it on rather than before.

Hear Where The Turns Land

Talk at the playground and watch turn boundaries, formatting and finalization happen live, before you write a line of WebSocket code.

Try playground

Entity accuracy is the metric that decides whether agent assist works

Here is the thing that word error rate hides. Agent assist is, mechanically, a lookup system. Somebody reads an account number aloud. Somebody spells a surname. Somebody gives an address, an order ID, a member number, a phone number. Every one of those is a string that has to be exactly right or the lookup returns nothing — and a lookup that returns nothing is worse than an assist that stayed quiet, because the agent now has to recover from a wrong answer in front of the customer.

A model can post a respectable overall word error rate and still be unusable here, because entities are a small fraction of total words and a large fraction of the words that matter. What you want to compare is entity error rate. These figures are from the Pipecat open speech-to-text benchmark, run on real agent conversations. Lower is better.

Metric Universal-3.5 Pro Realtime Deepgram Flux ElevenLabs Scribe v2 Google Chirp3
Word error rate 6.99% 15.58% 9.76% 9.04%
Entity error rate 15.31% 50.50% 39.70% 21.51%
Names 16.92% 39.21% 38.03% 22.10%
Places 6.28% 14.86% 34.06% 10.04%
Phone numbers 3.55% 10.41% 4.78% 4.95%

Read the names row rather than the word error rate row. Names are the hardest category in every system, and they are the category agent assist depends on most, because the customer’s name and the agent’s read-back of it are the spine of an identity-verification flow. Full methodology and the rest of the comparisons are on the benchmarks page.

What agent_context changes

Entity accuracy is not fixed at the model level. You can move it by telling the model what is being asked. If the system knows the agent just said “can you confirm the last four digits of your card”, a mumbled “five two eight one” resolves very differently than it would in a vacuum.

Measured across a benchmark of 10,000+ voice agent audio files, passing the agent’s own spoken reply as agent_context cut word error rate by 8.9%, rising to 16.4% when a context prompt is added on top. The published per-class breakdown is measured on that combined configuration:

  • Place-name entities −30.7%
  • Fabrications −27.0%
  • Medical entities −26.2%
  • Name entities −21.8%
  • Short-utterance errors −20.5%
  • Entity errors overall −11.1%

The short-utterance line is the one to notice for agent assist specifically. “Yes.” “Twelve.” “B as in bravo.” Those are the replies that carry the most operational weight and the least acoustic information, and they are where context does the most lifting. The prompting and keyterms documentation covers how to feed it, and keyterms prompting is included at no extra cost on Universal-3.5 Pro Realtime — so product names, plan names and internal jargon can be biased for free.

Redacting PII in the stream

Agent assist is deployed in exactly the two industries where the entities it is best at capturing are the entities you are least allowed to store: financial services and healthcare. A system that transcribes a card number perfectly and then writes it to a log has created a problem, not solved one.

Streaming PII Text Redaction handles this at the transcription layer rather than leaving it to your application code. It is an add-on at +$0.12/hr, priced on the pricing page alongside the rest of the streaming options. Redacting in the stream means the sensitive string never reaches your retrieval index, your analytics warehouse, or your quality-management recordings in the first place — which is a materially easier posture to defend than redacting on the way out.

The practical pattern is to run redaction on the transcript that gets persisted and displayed, while your lookup logic consumes the value it needs in-memory and discards it. That keeps the assist functional without widening the blast radius.

See The Entity Accuracy On Your Own Calls

The free tier includes 185 hours of pre-recorded transcription and 333 hours of streaming, which is enough to run a real pilot rather than a demo.

Sign up free

What this looks like when it is working

Rather than quote industry-wide averages, here is what customers building on this stack have reported publicly:

  • Siro — 90% reduction in customer complaints and support tickets.
  • Siro — 36% improvement in close rate.
  • Calabrio — 80% increase in customer satisfaction.

Two of those are sales-conversation outcomes and one is a contact-centre outcome, which is a fair reflection of where real-time assist lands commercially: the same architecture serves coaching, quality management and live support, and teams generally start with one and expand into the others once the transcription layer is in place.

Agent assist in healthcare

Healthcare contact centres and clinical support desks are one of the strongest use cases for real-time assist, and also the one with the most constraints. Three things to get right.

Medical vocabulary. Drug names, procedure names and dosages are entities, and they fail in the same way account numbers fail. Medical Mode is a single parameter — domain: “medical-v1” — paired with speech_model: “universal-3-5-pro” on streaming. It posts a 3.2% Missed Entity Rate and 87% fewer entity errors than the base model. It covers English, Spanish, German and French, on both pre-recorded and streaming. Pricing is a $0.15/hr add-on, which is $0.60/hr combined with streaming. See the streaming Medical Mode documentation for the request shape.

Context on top of the domain model. These are complementary levers, not alternatives. Medical Mode tunes the model for clinical vocabulary generally; contextual prompting tells it about this specific encounter. On a public benchmark of 20,000 real voice-agent calls, detailed context cut medical-term entity errors by 43%, and scenario-level context by 24%. Use both.

PHI handling. AssemblyAI is a business associate under HIPAA, not a covered entity. AssemblyAI enables covered entities and their business associates subject to HIPAA to use the AssemblyAI services to process protected health information (PHI). AssemblyAI is considered a business associate under HIPAA, and we offer a standard Business Associate Addendum (BAA) that is required under HIPAA to ensure that AssemblyAI appropriately safeguards PHI. The BAA can be reviewed and signed self-serve without a sales call — details are in the BAA FAQ and the Business Associate Addendum itself.

Supporting the same deployments: PHI redaction across audio and transcripts, SOC 2 Type 2, ISO 27001:2022, and PCI DSS v4.0. For teams with data residency requirements, EU endpoints are available at streaming.eu.assemblyai.com at the same price, with data staying in the EU, and self-hosted deployment into a customer VPC is supported. More on the security page.

What it costs to run

Agent assist economics are per-hour-of-audio, and the useful exercise is to price a concentrated configuration rather than the base rate alone.

Component Rate Notes
Universal-3.5 Pro Realtime $0.45/hr Base streaming, 18 languages, keyterms prompting included
Streaming diarization +$0.12/hr Up to 10 speakers, with revision
Streaming PII Text Redaction +$0.12/hr Redact in the stream, not after
Voice Focus +$0.10/hr near-field or far-field
Medical Mode +$0.15/hr $0.60/hr combined with streaming
Voice Agent API (full turn) $4.50/hr flat Only if a machine is speaking back

On the Voice Agent API, from the pricing page: “Every feature is included in the $4.50/hr rate. There are no per-layer add-ons, concurrency fees, or per-agent subscriptions.” Billing is per second, with no minimums — which matters for contact centres, where volume is spiky by definition and per-seat licensing models tend to price the peak rather than the average.

How to build it

  1. Open a streaming connection to wss://streaming.assemblyai.com/v3/ws with speech_model: “universal-3-5-pro”. Start on mode: “balanced”.
  2. Feed agent_context on every turn. If your agent desktop knows what prompt the agent is on, that string is the highest-leverage input you have.
  3. Load your keyterms. Product names, plan names, competitor names, internal codes. Included on Universal-3.5 Pro Realtime.
  4. Tune turn detection against recorded calls before touching live traffic. Change the mode preset first; only then reach for the individual thresholds.
  5. Add redaction before you add persistence. It is much easier than retrofitting it into a warehouse that already has a year of transcripts in it.
  6. Measure entity accuracy, not word accuracy, against a held-out set of your own calls, weighted toward the strings your lookups actually use.

If you want to feel the turn-detection behaviour before writing any code, the playground will let you talk at it and watch the boundaries land. If you are scoping a contact-centre rollout with procurement and security in the room, talk to an AI expert.

Start building

The transcription layer is the part you should validate first, because everything above it inherits its latency and its entity accuracy. Run your own calls through it, measure the entity error rate on the strings your product actually looks up, and tune turn detection against recordings before you go near live traffic.

Further reading: AI voice agents · voice agent solutions · Voice Agent API documentation

Scoping A Contact-Centre Rollout

Talk through turn-detection tuning, streaming PII redaction, EU data residency and volume pricing with someone who has taken agent assist to production.

Talk to AI expert

Frequently asked questions

What is real-time agent assist?

Real-time agent assist is software that transcribes a live call as it happens and surfaces information to the human agent during the conversation — knowledge base answers, customer account state, compliance prompts, or a running summary. It differs from post-call analytics in that its output has to arrive fast enough for the agent to act on it before the moment passes, which typically means a complete transcript within a few hundred milliseconds of the customer finishing a sentence.

How fast does transcription need to be for agent assist to work?

Fast enough that the transcript is finished and formatted before the agent has moved on — in practice, within the natural conversational pause, a few hundred milliseconds. Universal-3.5 Pro Realtime emits transcripts continuously as audio arrives and finalizes at end of turn, with end-of-turn decided from what has been said rather than from a silence timer. Measure the span on your own audio: vendors report different spans under the same word, so a published headline is a starting point rather than an answer.

What is the difference between real-time agent assist and a voice agent?

Agent assist supports a human who is doing the talking; a voice agent does the talking itself. That changes the latency budget: agent assist needs only the speech-to-text layer to be fast, while a voice agent needs a complete speech-in-to-speech-out turn, which the Voice Agent API runs at approximately 1 second end to end because it also covers the language model and text-to-speech. Many teams deploy both, using the same streaming transcription layer underneath.

Why do entity errors matter more than word error rate in agent assist?

Because agent assist is a lookup system, and a lookup fails completely on one wrong character in an account number, order ID or surname. Entities are a small share of total words, so a model can post a good overall word error rate while being unreliable on exactly the strings the product depends on. On the Pipecat open benchmark, Universal-3.5 Pro Realtime posts a 15.31% entity error rate against 50.50% for Deepgram Flux, 39.70% for ElevenLabs Scribe v2 and 21.51% for Google Chirp3.

Can real-time agent assist be used in healthcare with protected health information?

Yes, with a Business Associate Addendum in place. AssemblyAI is a business associate under HIPAA and offers a standard BAA, which can be reviewed and signed self-serve without a sales call. For clinical vocabulary, Medical Mode adds domain: “medical-v1” to a Universal-3.5 Pro Realtime stream at $0.60/hr combined, posting a 3.2% Missed Entity Rate across English, Spanish, German and French. PHI redaction across audio and transcripts, SOC 2 Type 2, ISO 27001:2022 and PCI DSS v4.0 support these deployments.

How much does real-time agent assist cost to run?

The transcription layer starts at $0.45/hr on Universal-3.5 Pro Realtime, with keyterms prompting included. Common additions are streaming diarization at $0.12/hr, streaming PII Text Redaction at $0.12/hr, Voice Focus at $0.10/hr, and Medical Mode at $0.15/hr. If you also need a machine to speak back, the Voice Agent API is a flat $4.50/hr with every feature included and no concurrency fees. Current rates are on the pricing page.