Overview
This guide covers using AssemblyAI’s Universal-3.6 Pro speech-to-text model as the transcriber in a Bolna voice agent. Bolna is a voice AI platform for building and running phone and web voice agents. Each agent is a pipeline of a transcriber (speech-to-text), an LLM, and a synthesizer (text-to-speech). Bolna handles the orchestration, including telephony, turn-taking, and interruptions. You configure agents in the Bolna dashboard or through the Bolna API.Universal-3.6 Pro is our flagship next-generation streaming model for voice agents — multilingual and promptable.Available on Bolna: set the transcriber
provider to "assembly" and model to "universal-3-6-pro".Turn detection
How Bolna uses AssemblyAI’s end-of-turn signal to decide when the caller is done.
Accuracy
Prompting and key terms for names, brands, and jargon.
Languages
How Bolna steers Universal-3.6 Pro to your agent’s language.
Running your agent
Web calls, phone calls, and the audio format Bolna uses for each.
Bolna AssemblyAI transcriber docs
View Bolna’s AssemblyAI transcriber reference.
For a standalone voice agent without Bolna, see the AssemblyAI Voice Agent API, which handles STT, LLM routing, and TTS in a single WebSocket connection.
Quickstart
Get a working, talking agent in a few minutes, then optimize from there.1
Create a Bolna account
Sign up or sign in at platform.bolna.ai.
2
Set AssemblyAI as the transcriber
- Dashboard
- API
- Open Build → Agent Studio, then select an agent or click New Agent.
- Open the Languages tab and go to Transcription.
- Set Provider to AssemblyAI and Model to
universal-3-6-pro. - Optionally, fill in Keywords and Context. See Accuracy.
- Click Save agent.
3
Run and test
- Dashboard
- API
Open the Test options dropdown next to Get a call from agent. Choose Web Call (Beta) to talk to the agent in your browser, or Get a call from agent to receive a phone call.
Parameters reference
Set these fields ontools_config.transcriber in the Bolna agent config. In the dashboard, they are in the Transcription section of the Languages tab.
string
required
Must be
"assembly".string
required
The streaming model, passed to AssemblyAI as
speech_model.
"universal-3-6-pro" is the recommended flagship model. "universal-3-5-pro"
is also supported. If you leave it out, Bolna falls back to a Deepgram model
and the request fails language validation, so always set it. See Speech
model comparison.string
default:"en"
The caller’s language as an ISO 639-1 code, such as
en, es, or hi. See
Languages.string
Comma-separated terms to boost. Sent to AssemblyAI as
keyterms_prompt. Up to 100 terms; Bolna
rejects an agent with more.string
A short, plain description of the call. Sent to AssemblyAI as
prompt. Up to 1,750 characters; Bolna
rejects an agent with a longer context.boolean
default:"true"
Keep
true for real-time transcription.sampling_rate or encoding. Bolna sets them from the call channel. See Running your agent.
Turn detection
Bolna uses AssemblyAI’s built-in turn detection:- While the caller speaks, AssemblyAI streams partial transcripts. Bolna treats them as interim.
- When AssemblyAI sends a
Turnmessage withend_of_turn: true, Bolna treats that transcript as final and sends it to the LLM.
min_turn_silence or max_turn_silence, so AssemblyAI’s default timing applies. In the dashboard, the Endpointing slider under Engine → Response Latency is hidden for AssemblyAI for this reason.
Linear Delay (task_config.incremental_delay, 900 ms on new agents) still applies. After the first exchanges, Bolna holds the agent’s reply until that long has passed since the caller stopped speaking, and cancels the reply if the caller keeps talking within that window. Lower it for snappier replies; raise it if callers pause mid-thought and the agent talks over them.
For how AssemblyAI decides that a turn has ended, see Turn detection.
Accuracy
Prompting
Describe the call incontext. Bolna sends it to AssemblyAI as prompt:
Key terms
List proper nouns, product names, and SKUs inkeywords. Bolna sends them to AssemblyAI as keyterms_prompt:
Languages
Bolna steers Universal-3.6 Pro to the agent’s language. It takes the base code oflanguage (for example, en from en-IN) and sends it as a single-element language_codes list, which makes each session monolingual: the model transcribes in the agent’s language instead of code-switching. If the code is not one AssemblyAI supports, Bolna leaves out language_codes and the model detects the language itself.
Codes Bolna sends as language_codes:
When you configure an agent, the Bolna dashboard shows the languages available for each model. For multilingual agents, set the transcriber separately for each language on the Languages tab. See Bolna’s Multilingual config reference for the API equivalent.
Running your agent
Web calls
Test in the browser with Web Call (Beta) in the dashboard. Bolna streams 16 kHzlinear16 audio to AssemblyAI.
Phone calls
Phone audio is 8 kHz. Its encoding depends on the telephony provider:
Bolna converts
mulaw audio to 16-bit PCM before sending it, so AssemblyAI always receives 16-bit PCM at the call’s sample rate.
Troubleshooting
Migrating from another STT provider
On Bolna, switching to AssemblyAI only changes the transcriber. The LLM, voice, prompts, and telephony stay the same.
Migrating a production deployment? Talk to our team.
Speech model comparison
The U3 Pro family is recommended for all new Bolna agents. The legacy
universal model is kept for existing agents.