Skip to main content

Overview

This guide covers using AssemblyAI’s Universal-3.6 Pro speech-to-text model as the transcriber in a Bolna voice agent. Bolna is a voice AI platform for building and running phone and web voice agents. Each agent is a pipeline of a transcriber (speech-to-text), an LLM, and a synthesizer (text-to-speech). Bolna handles the orchestration, including telephony, turn-taking, and interruptions. You configure agents in the Bolna dashboard or through the Bolna API.
Universal-3.6 Pro is our flagship next-generation streaming model for voice agents — multilingual and promptable.Available on Bolna: set the transcriber provider to "assembly" and model to "universal-3-6-pro".
AssemblyAI provides the speech-to-text and the turn detection in your Bolna agent: Once you have an agent running, tune it for what matters most to your use case:

Turn detection

How Bolna uses AssemblyAI’s end-of-turn signal to decide when the caller is done.

Accuracy

Prompting and key terms for names, brands, and jargon.

Languages

How Bolna steers Universal-3.6 Pro to your agent’s language.

Running your agent

Web calls, phone calls, and the audio format Bolna uses for each.

Bolna AssemblyAI transcriber docs

View Bolna’s AssemblyAI transcriber reference.
For a standalone voice agent without Bolna, see the AssemblyAI Voice Agent API, which handles STT, LLM routing, and TTS in a single WebSocket connection.

Quickstart

Get a working, talking agent in a few minutes, then optimize from there.
1

Create a Bolna account

Sign up or sign in at platform.bolna.ai.
To use the API, generate a Bolna API key. In the left sidebar, expand Developers and select API Keys. Every API request goes to https://api.bolna.ai with an Authorization: Bearer <key> header.
2

Set AssemblyAI as the transcriber

  1. Open Build → Agent Studio, then select an agent or click New Agent.
  2. Open the Languages tab and go to Transcription.
  3. Set Provider to AssemblyAI and Model to universal-3-6-pro.
  4. Optionally, fill in Keywords and Context. See Accuracy.
  5. Click Save agent.
The provider value is "assembly", not "assemblyai". The Bolna API rejects any other value.
3

Run and test

Open the Test options dropdown next to Get a call from agent. Choose Web Call (Beta) to talk to the agent in your browser, or Get a call from agent to receive a phone call.
Speak after you hear the greeting. Phone calls use Bolna wallet credits.

Parameters reference

Set these fields on tools_config.transcriber in the Bolna agent config. In the dashboard, they are in the Transcription section of the Languages tab.
string
required
Must be "assembly".
string
required
The streaming model, passed to AssemblyAI as speech_model. "universal-3-6-pro" is the recommended flagship model. "universal-3-5-pro" is also supported. If you leave it out, Bolna falls back to a Deepgram model and the request fails language validation, so always set it. See Speech model comparison.
string
default:"en"
The caller’s language as an ISO 639-1 code, such as en, es, or hi. See Languages.
string
Comma-separated terms to boost. Sent to AssemblyAI as keyterms_prompt. Up to 100 terms; Bolna rejects an agent with more.
string
A short, plain description of the call. Sent to AssemblyAI as prompt. Up to 1,750 characters; Bolna rejects an agent with a longer context.
boolean
default:"true"
Keep true for real-time transcription.
You don’t set sampling_rate or encoding. Bolna sets them from the call channel. See Running your agent.

Turn detection

Bolna uses AssemblyAI’s built-in turn detection:
  • While the caller speaks, AssemblyAI streams partial transcripts. Bolna treats them as interim.
  • When AssemblyAI sends a Turn message with end_of_turn: true, Bolna treats that transcript as final and sends it to the LLM.
There are no AssemblyAI turn-silence settings to tune on Bolna: Bolna doesn’t send min_turn_silence or max_turn_silence, so AssemblyAI’s default timing applies. In the dashboard, the Endpointing slider under Engine → Response Latency is hidden for AssemblyAI for this reason. Linear Delay (task_config.incremental_delay, 900 ms on new agents) still applies. After the first exchanges, Bolna holds the agent’s reply until that long has passed since the caller stopped speaking, and cancels the reply if the caller keeps talking within that window. Lower it for snappier replies; raise it if callers pause mid-thought and the agent talks over them. For how AssemblyAI decides that a turn has ended, see Turn detection.

Accuracy

Prompting

Describe the call in context. Bolna sends it to AssemblyAI as prompt:
Keep it to a sentence or two about the domain and what callers usually ask. It is not a prompt for the LLM. See Prompting for how the model uses context.

Key terms

List proper nouns, product names, and SKUs in keywords. Bolna sends them to AssemblyAI as keyterms_prompt:
In the dashboard, use the Context and Keywords fields in the Transcription section. For writing tips, see Bolna’s Keywords and context guide.

Languages

Bolna steers Universal-3.6 Pro to the agent’s language. It takes the base code of language (for example, en from en-IN) and sends it as a single-element language_codes list, which makes each session monolingual: the model transcribes in the agent’s language instead of code-switching. If the code is not one AssemblyAI supports, Bolna leaves out language_codes and the model detects the language itself. Codes Bolna sends as language_codes: When you configure an agent, the Bolna dashboard shows the languages available for each model. For multilingual agents, set the transcriber separately for each language on the Languages tab. See Bolna’s Multilingual config reference for the API equivalent.

Running your agent

Web calls

Test in the browser with Web Call (Beta) in the dashboard. Bolna streams 16 kHz linear16 audio to AssemblyAI.

Phone calls

Phone audio is 8 kHz. Its encoding depends on the telephony provider: Bolna converts mulaw audio to 16-bit PCM before sending it, so AssemblyAI always receives 16-bit PCM at the call’s sample rate.

Troubleshooting

Migrating from another STT provider

On Bolna, switching to AssemblyAI only changes the transcriber. The LLM, voice, prompts, and telephony stay the same. Migrating a production deployment? Talk to our team.

Speech model comparison

The U3 Pro family is recommended for all new Bolna agents. The legacy universal model is kept for existing agents.

Resources