New Universal-3.6 Pro Realtime is now available Learn more
Insights & Use Cases

Voice agents in noisy environments (Drive-Thrus, Contact Centers, Field)

Voice agents in noisy environments: learn how they work, where they fit, and how to build them for drive-thrus, contact centers, and field teams at scale.

Abstract green half-sphere illustration

Written by

Kelsey Foster

Published on

22 September 2026

If you are deploying a voice agent into a drive-thru, a contact center floor, a service van, or a self-service kiosk, three settings decide whether it works. voice_focus picks the noise model — near-field for headsets and phones, far-field for rooms and open spaces. voice_focus_threshold sets how hard that model works, from 0.0 to 1.0, defaulting to 0.85 on the Voice Agent API. And transcription_mode decides how much response time you are willing to spend on transcript quality: min_latency, balanced, or max_accuracy.

Those are three separate dials, not one. The first two are fixed when you open the connection; the third you can change mid-call. The rest of this post is about which combination fits which deployment, why that connect-time constraint matters more than it sounds, and what to do about the part that actually breaks in noise: entity accuracy.

Noise does not break words, it breaks entities

The intuitive model of noise is that it degrades everything a little. That is not what the data shows, and it is why teams misdiagnose the problem.

Common words survive background noise reasonably well, because the model has enormous linguistic context to fall back on. If half of “can I get a large” is masked by an engine, the sentence still resolves. Entities have no such safety net. A name, a street address, an order number, a phone number, a confirmation code — these are, by construction, sequences the model cannot predict from context. When the acoustics get hard, that is where the errors concentrate.

The Pipecat open STT benchmark, which uses real agent conversations rather than read speech, shows the gap clearly. Universal-3.5 Pro Realtime posts a 6.99% word error rate and a 15.31% entity error rate — entity errors run more than double the word error rate on the same audio. Break that down further and the pattern holds: 16.92% on names, 6.28% on places, 3.55% on phone numbers. For comparison on the same benchmark, Deepgram Flux posts 15.58% WER and 50.50% entity error, ElevenLabs Scribe v2 posts 9.76% and 39.70%, and Google Chirp3 posts 9.04% and 21.51%. The full table is on the benchmarks page.

The operational point: if you are monitoring your deployment on word error rate alone, a noisy site can look acceptable while it is failing on every order number. Instrument entity accuracy separately, on the entity types your product actually depends on.

The three dials, and which one you can still change mid-call

Before the scenario table, the constraint that shapes how you use it.

session.input.voice_focus selects the noise model. near-field is the default and suits audio captured close to the speaker — headsets, handsets, a phone held or clipped near the mouth — where the dominant problem is other people talking nearby. far-field suits audio captured at a distance, where reverberation and distance attenuation dominate: rooms, kiosks, drive-thru lanes, conference tables.

session.input.voice_focus_threshold is the strength dial, from 0.0 to 1.0, defaulting to 0.85 on the Voice Agent API. Higher values suppress more aggressively. It requires voice_focus to be set. This is the parameter that resolves the tradeoff every noisy deployment runs into — suppress too little and background speech leaks into the transcript, suppress too much and you clip the quiet or distant speaker you actually wanted. Most deployments should start at the default and only move once they have real audio to judge against.

One caveat if you also build directly on streaming speech-to-text: that API’s voice_focus_threshold defaults to 0.7, not 0.85, and its latency/accuracy preset is keyed as mode rather than transcription_mode. The values are the same on both surfaces; the defaults and the key names are not. Every number in this post is the Voice Agent API’s.

Both are set at connect. A mid-session change is accepted, but it only takes effect on the next speech-to-text reconnect — so in practice you are choosing the noise model when you open the connection, not adjusting it turn by turn. That is the design constraint worth planning around. If your kiosk sits in a lobby that is quiet at 7am and loud at noon, you either pick a setting that holds across the day or cycle the session when conditions change. For a drive-thru, where conditions vary by vehicle rather than by hour, it means picking a setting that holds for the worst realistic case rather than the average one.

session.input.transcription_mode is the latency/accuracy control — min_latency, balanced (the default), or max_accuracy — and unlike the two above, it is freely mutable mid-session. That asymmetry is genuinely useful. You can run a call on balanced for the conversational parts and switch to max_accuracy for the stretch where the customer reads out a confirmation number or an address, then drop back. You cannot do the equivalent with the noise model. Adapt accuracy on the fly; pick the noise model once, correctly.

near-field vs far-field vs max_accuracy: which setting for which environment

The mapping below is our recommendation rather than documented canon — the parameters and their ranges are documented, but which value suits which room is engineering judgment. Treat the thresholds as starting points to tune against your own recordings.

Environment voice_focus voice_focus_threshold transcription_mode Why
Drive-thru far-field 0.85–0.95 max_accuracy Microphone is a fixed distance from a car window, competing with engine noise, HVAC, road noise, and passengers. The order is entity-dense — sizes, modifiers, counts — and a wrong order costs more than a half-second pause. Set toward the high end because the noise floor is relentless, and spend the latency.
Contact center floor near-field 0.85 (default) balanced Agent or customer is on a headset or handset, so the speech is close-mic’d. The dominant problem is cross-talk from neighbouring desks, which near-field is built to reject. Fast turn-taking, so balanced keeps the rhythm natural — switch to max_accuracy for the address-capture stretch, then switch back.
Field and mobile near-field 0.85–0.9 balanced, or max_accuracy on capture-only flows A phone held or clipped near the speaker is near-field audio even outdoors. Wind, traffic, and machinery are loud but the source is close. If the audio is being captured for later review rather than answered live, there is no latency budget to protect — run max_accuracy throughout.
Self-service kiosk far-field 0.85–0.95 max_accuracy Microphone is in a housing, the speaker stands back from it, and the room reverberates. Bystander speech is the hardest part, which argues for the high end. Kiosk interactions are short and transactional, so a slightly slower, more accurate response reads as normal.
Conference room / multi-speaker far-field 0.7–0.85 balanced Distance mic in a reverberant space, but every participant is a wanted speaker — this is the one case where aggressive separation works against you, so stay at or below the default. Pair with streaming diarization if you need to attribute turns.
Quiet desk / web widget near-field 0.85 (default) balanced or min_latency The baseline case. min_latency is worth trying where the interaction is simple command-and-response and the acoustic conditions are genuinely clean.

Two notes. First, balanced is the default transcription_mode and near-field is the default voice_focus, so an unconfigured session is already running a near-field model in balanced mode — which is exactly wrong for a drive-thru. Second, Voice Focus is a $0.10/hr add-on on the streaming path and is included in the flat rate on the Voice Agent API, where every feature is included in the $4.50/hr with no per-layer add-ons, concurrency fees, or per-agent subscriptions.

When to move off 0.85

Do not tune the threshold from intuition. Record twenty real interactions at the site, transcribe them at the default, and read the failures. If background speech is appearing in transcripts as though the customer said it, raise the threshold — higher values suppress more aggressively. If the customer’s own words are dropping out — especially quiet speakers, children, or anyone standing further back than the design assumed — lower it. Those two failure modes look completely different in a transcript, which is what makes this dial tunable at all rather than a matter of taste.

The latency budget, by layer

“How fast is it” is two questions, and blending them into one range is how teams end up with a budget that does not add up.

Streaming speech-to-text returns transcripts continuously as the audio arrives, finalizing at the end of each turn. That is the transcription layer alone: audio in, finished text out. End-of-turn detection does not run on a silence timer — when a speaker pauses, the model evaluates what has been said so far to judge whether the turn is complete. That matters more in noise than it sounds, because a silence-only detector in a loud room either fires early on a background gap or never fires at all. Turn detection exposes a minimum and maximum turn silence and a voice-activity threshold, all tunable — check the streaming turn-detection documentation for the current defaults rather than hard-coding values from a blog post, ours included, because they are set by the mode preset and they change again when speaker labels are on.

TheVoice Agent API delivers roughly one second end-to-end. That is the whole loop — speech-to-text, the language model, and text-to-speech — over a single WebSocket.

Use the right number for the right conversation. If you are building the stack yourself and comparing transcription vendors, measure the transcription layer on your own audio and compare like with like. If you are budgeting the user-perceived pause before the agent starts speaking, roughly one second end to end is the figure that matters. Moving to max_accuracy spends against whichever budget you are tracking, which is why the table above ties the mode to the interaction type — and why the ability to change it mid-session is worth using rather than picking one value for the whole call.

Measure It On Your Own Audio

333 hours of streaming are included on the free tier, which is more than enough to profile a noisy site end to end before you commit to a configuration.

Sign up free

Barge-in is semantic, not volume-triggered

This is the behaviour that separates a voice agent that survives a noisy room from one that does not, and it is easy to miss because it is invisible when it works.

A naive interruption model listens for energy on the input and stops the agent talking whenever it hears some. In a drive-thru or on a contact center floor, that model fails constantly: a passenger laughs, a neighbouring agent finishes a sentence, a door closes, and your agent stops mid-word for no reason the customer can see.

Interruption on the Voice Agent API is semantic. It is driven by what was said, not by how loud it was. Back-channels — “uh-huh”, “mm-hm”, “makes sense” — do not interrupt the agent, because acknowledging that you are listening is not the same as trying to take the floor; “wait, stop” does. That distinction is doing a lot of work in exactly the environments this post is about, where the microphone is picking up far more sound than the one person you care about.

If brief back-channels are still cutting the agent off, the documented lever is interruption_delay — how long after speech begins a barge-in can interrupt, from 0 to 1000 ms. It follows the transcription mode by default (0 for min_latency, 500 for balanced and max_accuracy). Raise it rather than reaching for the VAD threshold.

When a genuine interruption does land, the server tells you clearly: reply.done arrives with status: “interrupted”, and transcript.agent carries interrupted: true with its text trimmed to what the user actually heard. Handle both if you are logging what the agent actually delivered versus what it intended to say — in a noisy deployment those diverge more often, and the difference is worth having in your analytics rather than discovering it in a complaint.

Fixing the entity problem: context and keyterms

Once the noise model and the mode are set correctly, the remaining headroom in a noisy deployment is almost entirely in entity accuracy. Two levers move it, and they do different jobs.

Transcription context passes the agent’s own spoken reply into the transcription request. When the agent has just asked “what’s the confirmation number on your order?”, the model is no longer transcribing a mumbled string in isolation; it knows what shape of answer is coming. This matters most in exactly the conditions this post is about: short replies, bad acoustics, no surrounding sentence to disambiguate against.

The field depends on the surface, and this is worth getting right because the two APIs do not share it. On the Voice Agent API it is session.input.transcription_prompt, capped at 1750 characters. On the streaming speech-to-text API it is agent_context, which is a streaming field and is not available on the Voice Agent API. Do not copy agent_context into a Voice Agent session config — it does not belong there.

The published measurement comes from the streaming API, across a benchmark of 10,000+ voice agent audio files: passing the agent’s reply as context cut word error rate by 8.9%, rising to 16.4% when a context prompt is added on top. The per-class breakdown is measured on that combined configuration: place-name entities down 30.7%, fabrications down 27.0%, medical entities down 26.2%, name entities down 21.8%, short-utterance errors down 20.5%, and entity errors overall down 11.1%. The short-utterance and place-name numbers are the ones to watch for a noisy deployment — those are the failure modes a drive-thru or a kiosk produces all day, and they are the two categories context helps most.

Context Carryover, a rolling memory across turns, is on by default, so the model retains what has already been established in the conversation rather than treating each turn as cold.

Keyterms — session.input.keyterms, up to 100 strings — is the other lever, and the one people under-use. It biases the model toward an explicit list: your menu items, your product SKUs, your store names, the street names in your service area, the medication names in your formulary. It is included at no extra cost on Universal-3.5 Pro Realtime. If your drive-thru sells something with a proper noun in the name, it belongs in the keyterms list before you do anything else.

The two compose. Context tells the model what this turn is about; keyterms tell it what vocabulary exists in your world. Full parameter documentation is in prompting and keyterms.

Language coverage in noisy environments

Noise and code-switching compound, and multilingual deployments in loud settings are where stacks fail hardest. A bilingual customer at a drive-thru window switching languages mid-order is a genuinely difficult input, and a pipeline that detects language per session cannot handle it.

Universal-3.5 Pro Realtime supports 18 languages — English, Spanish, French, German, Italian, Portuguese, Arabic, Danish, Dutch, Finnish, Hebrew, Hindi, Japanese, Mandarin, Norwegian, Swedish, Turkish, and Vietnamese — with native code-switching inside a single utterance rather than routing between per-language models. On the async code-switching benchmark across five language pairs, Universal-3.5 Pro posts a 7.69 average normalized WER against 8.77 for ElevenLabs Scribe v2 and 12.22 for Deepgram Nova-3 Multilingual.

Streaming caps at those 18 languages. If you have seen a 99-language figure associated with AssemblyAI, that belongs to universal-2, a previous-generation async model at $0.15/hr kept for legacy pre-recorded integrations. It is not a voice agent model and there is no 99-language streaming path.

Configuration is a single decision: pass language_codes as a list to bias toward a known set, or omit it for auto-detection and free code-switching. On mixed-language traffic, omitting it is usually the better choice.

Connecting to the Voice Agent API

One WebSocket, one endpoint, one flat rate. The handshake has an order to it: connect, send session.update immediately, wait for session.ready, and only then start streaming input.audio. Sending audio before session.ready is an error, and it is the most common way a first integration fails.

import asyncio
import base64
import json
import os

import websockets

# One WebSocket for the whole agent: STT, LLM, and TTS.
# Flat $4.50/hr, every feature included.
ENDPOINT = "wss://agents.assemblyai.com/v1/ws"

# Tuned for a noisy, fixed-location deployment (drive-thru, kiosk).
# voice_focus and voice_focus_threshold are set at connect; a mid-session
# change only takes effect on the next speech-to-text reconnect.
# transcription_mode IS freely mutable — send another session.update
# later in the call to change it.
SESSION_UPDATE = {
    "type": "session.update",
    "session": {
        "input": {
            "voice_focus": "far-field",        # near-field (default) | far-field
            "voice_focus_threshold": 0.9,      # 0.0-1.0, default 0.85 here
            "transcription_mode": "max_accuracy",  # min_latency | balanced | max_accuracy
            "keyterms": [                      # up to 100 strings
                "Baja Crunchwrap",
                "Willowbrook Plaza",
                "curbside",
            ],
            # Max 1750 chars. Tell the model what answer to expect.
            # NOTE: this is the Voice Agent field. On streaming STT the
            # equivalent is agent_context, which does NOT work here.
            "transcription_prompt": (
                "The agent is taking a drive-thru order and will ask for a "
                "confirmation number, a name, and pickup preferences."
            ),
            # min_silence / max_silence are deliberately NOT set. Left unset
            # they are adaptive and the agent slows down on its own while a
            # customer reads out an order number. Setting them turns that off.
            "turn_detection": {
                "interrupt_response": True,
                "interruption_delay": 500,     # ms, 0-1000; raise if back-channels cut in
            },
        }
    },
}


async def run_agent(audio_source):
    headers = {"Authorization": f"Bearer {os.environ['ASSEMBLYAI_API_KEY']}"}

    async with websockets.connect(ENDPOINT, additional_headers=headers) as ws:
        # 1. Configure the session immediately on connect.
        await ws.send(json.dumps(SESSION_UPDATE))

        # 2. Wait for session.ready before sending a single audio frame.
        async for message in ws:
            if json.loads(message)["type"] == "session.ready":
                break

        async def send_audio():
            # PCM16, mono, 24 kHz, base64-encoded in each input.audio event.
            async for chunk in audio_source:
                await ws.send(json.dumps({
                    "type": "input.audio",
                    "audio": base64.b64encode(chunk).decode("ascii"),
                }))

        async def receive_events():
            async for message in ws:
                event = json.loads(message)

                if event["type"] == "transcript.user":
                    print("user:", event)

                elif event["type"] == "reply.done":
                    # Semantic barge-in: back-channels do not interrupt.
                    if event.get("status") == "interrupted":
                        print("customer took the floor mid-reply")

                elif event["type"] == "session.error":
                    print("error:", event)

        await asyncio.gather(send_audio(), receive_events())


asyncio.run(run_agent(audio_source))

The event surface

There are seven client→server events — input.audio, session.update, session.resume, session.end, tool.result, reply.create, and conversation.message — and fourteen server→client events: session.ready, session.updated, session.ended, input.speech.started, input.speech.stopped, transcript.user.delta, transcript.user, reply.started, reply.audio, transcript.agent.delta, transcript.agent, reply.done, tool.call, and session.error.

For a noisy deployment, four of those earn special attention. input.speech.started and input.speech.stopped tell you what the endpointer thinks is happening, which is the fastest way to diagnose a threshold that is set wrong — if speech events fire on an empty lane, your suppression is too weak. reply.done with status: “interrupted” and transcript.agent with interrupted: true tell you when the agent got cut off, which is the signal you want trending on a dashboard per site.

One migration note: the speech model parameter is singular on the streaming path (speech_model) while the async API takes the plural speech_models as a list, and the legacy realtime endpoint is retired and should not appear in your code. Full details in the Voice Agent API documentation.

A fuller walkthrough of the build is in how to build with the Voice Agent API, and the launch post covers the architecture rationale. If you need per-speaker attribution in a room, streaming speaker diarization handles up to 10 speakers with revision, correcting within about half a second of the stream ending.

Try It With Your Noisiest Recording

Run a clip from your actual site through the playground before you tune anything. That single test will tell you more than any benchmark table.

Try playground

What good looks like in production

Siro, one of the voice AI teams building on AssemblyAI, runs speech capture in one of the least forgiving acoustic environments there is — real field sales conversations, recorded on a phone in someone’s pocket, outdoors, in cars, in customers’ homes. Their published results: a 90% reduction in customer complaints and support tickets and a 36% improvement in close rate. That is what near-field capture plus disciplined entity handling looks like once it is tuned, and it is a useful counterweight to the assumption that difficult audio is simply a cost of doing business in the field.

The pattern across deployments that work is unremarkable. Pick voice_focus to match the physical microphone placement and get it right at connect, because changing it later only takes effect on a reconnect. Start the threshold at whatever the default is on your surface and tune it against recordings from the actual site. Use transcription_mode dynamically — balanced for conversation, max_accuracy for the entity-dense stretches. Leave the silence windows unset so the adaptive endpointer can do its job. Pass a transcription prompt on every session, keep keyterms current with your real vocabulary, and measure entity accuracy rather than word accuracy. None of that is exotic. It is mostly a matter of not leaving the defaults in place because the demo sounded fine — and knowing which defaults are worth leaving alone.

Sizing A Multi-Site Rollout

Talk through noise model selection, threshold tuning per site, and entity accuracy targets with someone who has configured these deployments before.

Talk to AI expert

Frequently asked questions

What is the best setting for a voice agent in a noisy environment?

Set voice_focus to far-field with voice_focus_threshold toward the high end of its range, and transcription_mode to max_accuracy, for fixed-location noise like drive-thrus and kiosks where the microphone sits at a distance from the speaker. Use near-field at the default threshold with balanced for headsets, handsets, and phones, which covers most contact center and field traffic. These are three independent dials and all three are worth setting deliberately.

What is the difference between near-field and far-field?

near-field, the default, is for audio captured close to the speaker — headsets, handsets, phones held or clipped near the mouth — where the main problem is background speech from nearby people. far-field is for audio captured at a distance — rooms, kiosks, drive-thru lanes, conference tables — where reverberation and distance attenuation dominate. Both are values of session.input.voice_focus, which is a $0.10/hr add-on on streaming and included in the Voice Agent API’s flat rate.

What does voice_focus_threshold do?

It sets how aggressively the selected noise model separates the target speaker from everything else, on a scale of 0.0 to 1.0 — higher values are more aggressive. The default depends on the surface: 0.85 on the Voice Agent API and 0.7 on the streaming speech-to-text API. It requires voice_focus to be set. Raise it when background speech is leaking into transcripts as though the customer said it; lower it when the customer’s own words are being clipped, which typically shows up first with quiet speakers or anyone standing further back than the deployment assumed.

Can I change noise settings mid-call?

Not effectively. voice_focus and voice_focus_threshold are set at connect; a mid-session change is accepted but only takes effect on the next speech-to-text reconnect, so you cannot adjust them turn by turn. transcription_mode is different: it is freely mutable mid-session, so you can run a call on balanced and switch to max_accuracy for the stretch where a customer reads out an address or a confirmation number, then switch back.

Does AssemblyAI have a noise suppression endpoint?

No. There is no separate noise suppression endpoint and no standard/aggressive intensity setting. Noise handling is session.input.voice_focus on the existing streaming and Voice Agent connections, set to near-field or far-field, with strength controlled by session.input.voice_focus_threshold. The min_latency / balanced / max_accuracy values belong to session.input.transcription_mode on the Voice Agent API — or to mode on streaming speech-to-text — a separate control governing the latency-accuracy tradeoff rather than noise.

Will background noise make the agent interrupt itself?

No, because barge-in is semantic rather than volume-triggered. Interruption is driven by what was said, not by how loud the input was, so back-channels like “uh-huh” and “mm-hm” do not stop the agent — and neither does ambient noise that carries no intent to take the floor. If brief back-channels are still cutting in, raise interruption_delay (0–1000 ms) rather than the VAD threshold. When a real interruption happens, the server emits reply.done with status: “interrupted” and transcript.agent with interrupted: true.

Why do voice agents get names and order numbers wrong in noisy rooms?

Because entities carry no linguistic context to recover from. Common words survive background noise because the surrounding sentence disambiguates them; a confirmation code or a surname does not. On the Pipecat benchmark, Universal-3.5 Pro Realtime’s entity error rate of 15.31% runs more than double its 6.99% word error rate on the same audio. The levers are transcription context — session.input.transcription_prompt on the Voice Agent API, agent_context on streaming, where it cut word error rate by 8.9% across 10,000+ files and 16.4% with a context prompt on top, with short-utterance errors down 20.5% and place-name errors down 30.7% on that combined configuration — and keyterms, which bias the model toward your own vocabulary at no extra cost.