Skip to main content
Real calls are noisy: keyboard clacks, café chatter, a car engine, room echo. That noise makes the agent mishear words and, worse, can trip false turn-taking and interruptions. input.voice_focus isolates the primary speaker and suppresses background audio before it reaches transcription, so the agent responds to the caller and not the room. Set it when you create or update the agent. Match the model to how the caller is captured: a browser agent on a headset wants near-field; a phone or in-room agent wants far-field.
Voice Focus is applied when the speech-to-text connection opens, so set it at agent-create or connect time. For how the models and threshold behave under the hood, see Voice Focus in the Streaming docs.