To summarize a meeting with an LLM, you transcribe the recording into speaker-labeled text, then send that transcript to a language model with a prompt that specifies exactly what to extract. In Python, that is two API calls: one to AssemblyAI's speech-to-text API for the transcript, one to the LLM Gateway for the summary.
The part that decides whether this works is the first call, not the second. A summarization prompt cannot recover a name the transcriber got wrong, and it cannot attribute a decision to the right person if the speaker boundaries are in the wrong place. Every hallucinated action item in a meeting summary traces back to a word the model never had. This guide builds the pipeline end to end and shows where transcript quality actually determines output quality.
What you need to get started
You need an AssemblyAI API key, Python 3.9 or later, and two packages:
pip install -U assemblyai requests
Set your key once at the top of your script:
import assemblyai as aai
aai.settings.api_key = "YOUR_API_KEY"
Step 1: Transcribe the meeting audio
Meeting audio is the hard case for speech-to-text: overlapping speakers, cross-talk, varying microphone quality, people dialing in from cars. Pass a TranscriptionConfig and name the model explicitly rather than relying on a default — you want to know exactly what produced your transcript, and you want that pinned when you compare output across runs.
import assemblyai as aai
aai.settings.api_key = "YOUR_API_KEY"
MEETING_URL = "https://storage.googleapis.com/aai-web-samples/meeting.mp3"
config = aai.TranscriptionConfig(
speech_models=["universal-3-5-pro"],
)
transcript = aai.Transcriber(config=config).transcribe(MEETING_URL)
if transcript.status == aai.TranscriptStatus.error:
raise RuntimeError(f"Transcription failed: {transcript.error}")
print(transcript.text[:500])
universal-3-5-pro is the flagship async model at $0.21/hr. It handles 18 languages with native code-switching out of the box, which matters for meetings where a participant switches languages mid-sentence — a common failure point for models that require you to declare one language per file. If you need coverage beyond those 18, universal-2 supports 99 languages total at $0.15/hr with lower tier accuracy. See the Universal-3.5 Pro async release notes for the full benchmark set.
transcript.text is a single flat string. That is enough for a basic summary and not enough for a good one — we will add speaker labels shortly.
Step 2: Generate the summary with the LLM Gateway
The LLM Gateway gives you one endpoint and one API key for models from OpenAI, Anthropic, Google and others, plus models AssemblyAI hosts directly. No separate provider accounts, SDKs, or bills.
Write the summarization prompt
Treat the prompt as an output specification, not a request. Vague prompts produce vague summaries, and the failure mode is confident invention — a model asked to "summarize the meeting" will fill gaps rather than report them.
Four things belong in every meeting prompt: a role and scope, an explicit output structure with named sections in order, an instruction to write "not discussed" rather than guess at missing information, and a rule to preserve product names and acronyms as spoken instead of normalizing them.
PROMPT = """You are summarizing a recorded team meeting for someone who did not attend.
Produce exactly these sections, in this order:
## Decisions
Decisions that were actually made. Not proposals, not open questions.
## Action items
One bullet per item: owner, task, due date, current status.
If an owner or date was never stated, write "unassigned" or "no date".
## Open questions
Questions raised and left unresolved.
## Risks and blockers
Anything described as blocking, at risk, or dependent on something outside the team.
Rules:
- Use only what is in the transcript. If a section has no content, write "Not discussed."
- Preserve product names, acronyms and technical terms exactly as spoken.
- Do not infer intent that was not stated.
Transcript:
"""
Send the transcript to the LLM Gateway
The Gateway speaks the standard chat-completions shape, so this is an ordinary POST:
import requests
GATEWAY_URL = "https://llm-gateway.assemblyai.com/v1/chat/completions"
response = requests.post(
GATEWAY_URL,
headers={
"authorization": aai.settings.api_key,
"content-type": "application/json",
},
json={
"model": "claude-sonnet-4-6",
"messages": [
{"role": "user", "content": PROMPT + transcript.text}
],
"max_tokens": 2048,
"temperature": 0.0,
},
timeout=120,
)
response.raise_for_status()
summary = response.json()["choices"][0]["message"]["content"]
print(summary)temperature: 0.0 is deliberate. Summarization is an extraction task: the same transcript should produce the same summary every time, and you do not want creative latitude anywhere near an action-item list.
To get structured data instead of Markdown — writing action items straight into a task tracker, say — the Gateway supports structured outputs via response_format with a json_schema, and fallback chains so a provider outage downgrades to a second model instead of dropping the job.
Transcribe with Universal-3.5 Pro and summarize through the LLM Gateway using a single API key. Free credits to start, no card required.
Choose the right LLM model
Model choice for meeting summarization is a tradeoff between reasoning depth, latency and cost. A nightly batch job over yesterday's calls has different constraints than a summary that has to appear before the participants close their laptops.
| Model | Best for | Tradeoff |
|---|---|---|
qwen3.5-4b-32k-fast |
Turn summarization, live formatting, transcript rewrite. Self-hosted on AssemblyAI GPUs, latency-optimized, 32,768-token context. Averages 612ms. $0.10 per 1M input / $0.50 per 1M output. | Strongest on shorter audio. For a full-length meeting, chunk the transcript or pair it with a larger model for the final pass. |
claude-sonnet-4-6 |
The default for full-meeting summaries. Reliable instruction-following on multi-section output formats and long transcripts. | Higher cost per call than the fast tier. |
claude-haiku-4-5-20251001 |
High-volume batch summarization where cost per meeting is the binding constraint. | Less reliable on deeply nested output structures. |
gemini-3.7-flash |
Long transcripts you would rather not chunk, and summaries that need tool calling to write into downstream systems. | Output formatting drifts more than Claude on strict templates. |
| GPT-5.1 | Meetings that need inference across the whole conversation — connecting a decision in minute 8 to an objection in minute 47. | Slower and more expensive than the flash tier. |
| Opus 5 | The hardest cases: technical design reviews, contract negotiations, anything where a missed nuance is expensive. | Highest cost and latency. Reserve it for meetings that justify it. |
A practical pattern: run qwen3.5-4b-32k-fast over individual chunks or turns for speed, then run one claude-sonnet-4-6 pass to reconcile those partials into the final summary. You get fast-tier economics on the bulk of the tokens and strong reasoning where it matters.
Swapping models is a one-line change to the "model" field. See the chat completions reference for the full parameter list.
Speaker-aware summaries with diarization
A flat transcript tells you what was said. It does not tell you who committed to what — which is most of what a meeting summary is for.
Why diarization quality decides summary quality
Speaker diarization segments audio by speaker. When it drifts, an action item lands on the wrong person and the summary is worse than useless, because it is confidently wrong.
Universal-3.5 Pro is optimized for cpWER (concatenated minimum-permutation word error rate) rather than DER. That choice matters here: DER measures how well speaker boundaries line up in time, while cpWER measures what you actually consume downstream — whether the right words ended up attributed to the right speaker. Lower is better.
| Model | Average cpWER |
|---|---|
| AssemblyAI Universal-3.5 Pro | 30.17 |
| ElevenLabs Scribe v2 | 35.26 |
| Gladia | 36.87 |
| Deepgram Nova-3 English | 37.92 |
That gap compounds across a 60-minute call with six participants. For background on the technique, see what is speaker diarization.
Enable speaker diarization
One parameter. Async standard diarization adds $0.02/hr on top of the model rate.
config = aai.TranscriptionConfig(
speech_models=["universal-3-5-pro"],
speaker_labels=True,
)
transcript = aai.Transcriber(config=config).transcribe(MEETING_URL)
for utterance in transcript.utterances[:5]:
print(f"Speaker {utterance.speaker}: {utterance.text}")
Each utterance carries speaker, text, start and end. If you know the participant count in advance, pass speakers_expected to constrain the clustering.
Format diarized transcripts for the LLM
Models handle speaker attribution far better when the labels are inline in the text than when you describe the structure and hope. Flatten the utterances before you send them:
def format_diarized(transcript) -> str:
return "\n".join(
f"Speaker {u.speaker}: {u.text}" for u in transcript.utterances
)
speaker_transcript = format_diarized(transcript)Then adjust the prompt to use the attribution:
SPEAKER_PROMPT = """You are summarizing a recorded team meeting.
The transcript is labeled by speaker.
Produce:
## Decisions
Each decision, and which speaker made or confirmed it.
## Action items
Owner (by speaker label), task, due date, status.
If the owner was never stated, write "unassigned" — do not guess from context.
## Positions taken
Where speakers disagreed, summarize each position and who held it.
Use only the transcript. Write "Not discussed" for empty sections.
Transcript:
"""
response = requests.post(
GATEWAY_URL,
headers={
"authorization": aai.settings.api_key,
"content-type": "application/json",
},
json={
"model": "claude-sonnet-4-6",
"messages": [
{"role": "user", "content": SPEAKER_PROMPT + speaker_transcript}
],
"max_tokens": 2048,
"temperature": 0.0,
},
timeout=120,
)
response.raise_for_status()
print(response.json()["choices"][0]["message"]["content"])If you have a participant list, map Speaker A to real names before the LLM call. Summaries that say "Priya owns the migration" are acted on; summaries that say "Speaker B owns the migration" are not.
Upload a multi-speaker recording and see speaker-labeled output from Universal-3.5 Pro in the browser — no code required.
Customize your summaries
The transcription call stays the same. Change the prompt and you change the product. These three cover most of what teams actually need.
Focus on action items
ACTION_ITEMS_PROMPT = """Extract every action item from this meeting transcript.
For each, output a bullet with:
- Owner (speaker label or name; "unassigned" if never stated)
- Task, in one sentence
- Due date ("no date" if never stated)
- Status: not started / in progress / blocked
Include only commitments someone actually made. Exclude ideas that were
raised and dropped. If there are no action items, say so.
Transcript:
"""
Generate a technical discussion summary
TECHNICAL_PROMPT = """Summarize the technical content of this engineering meeting.
Cover:
- Technical decisions made, and the reasoning given
- Architecture changes approved
- Dependencies identified, including on other teams
- Technical debt acknowledged
- System constraints discussed (latency, cost, scale, compliance)
Preserve service names, version numbers and technical terms exactly as spoken.
Omit any section with no content.
Transcript:
"""
Create a project status report
STATUS_PROMPT = """Write a project status report from this meeting transcript.
Sections:
- Overall health: on track / at risk / off track, with the evidence for that call
- Completed since last update
- Upcoming deadlines
- Blocking issues and who owns unblocking them
- Resource needs raised
- Risks, with likelihood and impact if stated
Base the health call only on what was said. If there is not enough
information, write "insufficient information".
Transcript:
"""
The same approach extends to sales call summaries, support QA, and compliance review. AssemblyAI also delivers LLM-powered summarization and auto-chapters through Speech Understanding and the LLM Gateway if you would rather not maintain prompts yourself.
Handle very long meetings
A three-hour all-hands can exceed a model's context window, and even when it fits, quality degrades as the transcript grows. The reliable pattern is map-reduce: summarize chunks independently, then summarize the summaries. Chunk on utterance boundaries rather than character counts so you never split a speaker turn in half:
def chunk_utterances(utterances, minutes=15):
"""Group utterances into windows of roughly `minutes` each."""
window_ms = minutes * 60 * 1000
chunks, current, window_start = [], [], None
for u in utterances:
if window_start is None:
window_start = u.start
if u.start - window_start > window_ms and current:
chunks.append(current)
current, window_start = [], u.start
current.append(u)
if current:
chunks.append(current)
return chunks
def summarize(text: str, prompt: str, model: str = "claude-sonnet-4-6") -> str:
response = requests.post(
GATEWAY_URL,
headers={
"authorization": aai.settings.api_key,
"content-type": "application/json",
},
json={
"model": model,
"messages": [{"role": "user", "content": prompt + text}],
"max_tokens": 2048,
"temperature": 0.0,
},
timeout=120,
)
response.raise_for_status()
return response.json()["choices"][0]["message"]["content"]
CHUNK_PROMPT = """Summarize this segment of a longer meeting. Capture decisions,
action items with owners, and unresolved questions. Do not editorialize.
Segment:
"""
REDUCE_PROMPT = """These are sequential summaries of segments of one meeting.
Merge them into a single summary with sections: Decisions, Action items,
Open questions, Risks and blockers. Remove duplicates. Where two segments
conflict, keep the later one and note the change.
Segments:
"""
chunks = chunk_utterances(transcript.utterances, minutes=15)
partials = [
summarize(
"\n".join(f"Speaker {u.speaker}: {u.text}" for u in chunk),
CHUNK_PROMPT,
model="qwen3.5-4b-32k-fast",
)
for chunk in chunks
]
final_summary = summarize("\n\n---\n\n".join(partials), REDUCE_PROMPT)
print(final_summary)This is where the mixed-model pattern pays off. The map step runs many calls over short segments — exactly what qwen3.5-4b-32k-fast is built for — and the reduce step gets a stronger model on a much smaller input.
Handle errors in production
Two failure surfaces sit in this pipeline, and they fail differently. Transcription fails asynchronously and returns a status you have to check; the Gateway fails as an HTTP error. A silent transcription failure is the dangerous one, because transcript.text may be empty rather than raising.
import sys
import requests
import assemblyai as aai
def summarize_meeting(url: str, prompt: str) -> str:
config = aai.TranscriptionConfig(
speech_models=["universal-3-5-pro"],
speaker_labels=True,
)
try:
transcript = aai.Transcriber(config=config).transcribe(url)
except aai.AssemblyAIError as e:
print(f"Transcription request failed: {e}", file=sys.stderr)
sys.exit(1)
# Async jobs can fail without raising — check status explicitly.
if transcript.status == aai.TranscriptStatus.error:
print(f"Transcription failed: {transcript.error}", file=sys.stderr)
sys.exit(1)
if not transcript.utterances:
print("Transcript returned no utterances — check the audio.", file=sys.stderr)
sys.exit(1)
text = "\n".join(
f"Speaker {u.speaker}: {u.text}" for u in transcript.utterances
)
try:
response = requests.post(
GATEWAY_URL,
headers={
"authorization": aai.settings.api_key,
"content-type": "application/json",
},
json={
"model": "claude-sonnet-4-6",
"messages": [{"role": "user", "content": prompt + text}],
"max_tokens": 2048,
"temperature": 0.0,
},
timeout=120,
)
response.raise_for_status()
except requests.exceptions.RequestException as e:
print(f"LLM Gateway request failed: {e}", file=sys.stderr)
sys.exit(1)
return response.json()["choices"][0]["message"]["content"]Three things to add before this goes to production: retry with exponential backoff on 429 and 5xx responses, a fallback chain so a provider outage degrades to a second model instead of failing the job, and persistence of the raw transcript separately from the summary. The transcript is the expensive artifact — if a prompt change means you want a different summary next month, you should not have to re-transcribe.
Keep meeting recordings private
Meeting audio is among the most sensitive data a company holds: compensation discussions, legal strategy, customer names, unannounced roadmaps. Two controls matter when you put that through an LLM.
Zero Data Retention. Customer LLM responses are not stored by the LLM Gateway. Your summaries do not sit in a log waiting to become someone else's breach.
Regional routing. For data that must stay in the EU, point the Gateway at the EU endpoint:
GATEWAY_URL = "https://llm-gateway.eu.assemblyai.com/v1/chat/completions"
Transcription has a matching EU path at api.eu.assemblyai.com, at the same price as US, with data staying in the EU for GDPR purposes. Note that regional routing on the LLM Gateway carries a +10% surcharge — worth building into your cost model up front rather than discovering it at scale. Current rates for every model and add-on are on the pricing page.
Build a voice agent for meeting intelligence
If you want the summary to exist as the meeting ends rather than after it, the Voice Agent API bundles speech-to-text, LLM and text-to-speech behind a single WebSocket at wss://agents.assemblyai.com/v1/ws, billed at a flat $4.50/hr. It supports six languages — English, Spanish, French, German, Italian and Portuguese — with native code-switching across all six, and runs on Universal-3.6 Pro Realtime underneath.
For meeting work it is best suited to structured, transactional interactions: an agent that joins a standup, asks each participant for a status, and files the result. Open-ended note-taking over a long unstructured discussion is still better served by the async pipeline above. See the voice agent docs if you want to go that direction.
What this looks like in production
Granola builds an AI notetaker that turns raw meeting audio into usable notes, which makes it about as direct a test of this pipeline as exists.
"The speed difference is immediately noticeable — our users see their conversations transcribed almost instantaneously. It feels so much more responsive than what we were using before."
— Jonathan Kim, Software Engineer, Granola
The observation worth taking from that: for meeting products, transcription latency is a product feature, not an infrastructure detail. Users notice the wait.
Practical notes on quality
Audio input. Most bad summaries are bad recordings. Individual microphones beat a single room mic by a wide margin, because diarization has real channel separation to work with. Record locally where you can rather than relying on a compressed conference-bridge stream.
Prompt discipline. Version prompts alongside your code and diff summary outputs when you change them. An edit that improves action-item extraction can quietly degrade decision capture, and you will not notice without a before-and-after.
Meeting structure. A verbal recap of decisions before close is the single highest-leverage habit: it hands the model an unambiguous statement of every decision to anchor on.
Wrapping up
The pipeline is short: transcribe with speech_models=["universal-3-5-pro"] and speaker_labels=True, format the utterances with inline speaker labels, and send them to the LLM Gateway with a prompt that specifies the output structure and what to do about missing information. Everything else — model selection, chunking, error handling, regional routing — is tuning around those three steps.
The leverage is at the front. Summarization quality is bounded by transcript quality, and a diarization model that puts the right words next to the right speaker is what separates a summary your team acts on from one they stop reading. If you are building conversation intelligence more broadly, the same foundation carries into call analytics and QA workflows.
Universal-3.5 Pro at $0.21/hr, speaker diarization at $0.02/hr, and the LLM Gateway behind one API key. Free credits included.
Frequently asked questions
How do I debug transcription accuracy issues with poor audio quality?
Start by listening to the source audio before blaming the model — background noise, echo from a room mic, and heavy compression from a conference bridge are the usual causes. Enable speaker_labels=True and read the diarized output rather than the flat transcript, because speaker confusion is easier to spot than word errors and often reveals that two participants shared one microphone. If the meeting uses domain-specific vocabulary, contextual prompting lets you prime the model with product names and terminology before transcription, which resolves a large share of remaining errors.
What are the performance differences between LLM models for meeting summarization?
The tradeoff is reasoning depth against latency and cost. Use claude-sonnet-4-6 as the default for full-meeting summaries with multi-section output, qwen3.5-4b-32k-fast for latency-sensitive per-turn or per-chunk work (it averages 612ms and costs $0.10 per 1M input / $0.50 per 1M output tokens), claude-haiku-4-5-20251001 for high-volume batch jobs, and GPT-5.1 or Opus 5 when a meeting requires connecting arguments made far apart in the conversation. Switching between them is a one-line change to the model field in the LLM Gateway request.
How do I handle very long meetings that exceed LLM token limits?
Use a map-reduce pattern: split the transcript into roughly 15-minute windows on utterance boundaries so speaker turns stay intact, summarize each window independently, then run a final pass that merges the partial summaries and removes duplicates. A cost-efficient version runs the map step on qwen3.5-4b-32k-fast and the reduce step on claude-sonnet-4-6, since the reduce input is much smaller than the original transcript. Chunking on utterance boundaries rather than character counts is what keeps attribution correct across the seams.
What's the best approach for real-time vs batch meeting processing?
Batch processing after the meeting ends is the right default: you get the complete transcript, better diarization because the model has seen every speaker, and no latency constraint on the summarization model. Choose real-time only when the summary has to influence the meeting while it is still happening — live agenda tracking, or an agent participating in the conversation. Real-time work runs through Streaming STT or the Voice Agent API rather than the async pipeline.
How do I optimize costs when processing meetings at scale?
Cache aggressively so the same recording is never transcribed twice, and store transcripts separately from summaries so a prompt change does not force re-transcription. Route the bulk of your token volume to cheaper models and reserve the expensive ones for the final reasoning pass, and trim prompts of anything that does not change the output — prompt tokens are charged on every call. If you use regional routing on the LLM Gateway, budget the +10% surcharge into your per-meeting cost model.
Which speech model should I use for multi-speaker meeting recordings?
Use universal-3-5-pro with speaker_labels=True. It is optimized for cpWER, which measures whether the right words end up attributed to the right speaker, and averages 30.17 against Azure at 30.35, ElevenLabs Scribe v2 at 35.26, Speechmatics at 36.6, Gladia at 36.88 and Deepgram at 37.93. On a long meeting with several participants, that difference is what determines whether action items land on the correct owner.
Does AssemblyAI store my meeting transcripts or LLM responses?
The LLM Gateway operates under Zero Data Retention — customer LLM responses are not stored. For teams with data residency requirements, an EU endpoint is available at https://llm-gateway.eu.assemblyai.com/v1/chat/completions, with a matching EU transcription endpoint at api.eu.assemblyai.com where data stays in the EU for GDPR purposes at the same price as US. Regional routing on the Gateway carries a +10% surcharge.
What languages can I summarize meetings in?
Universal-3.5 Pro handles 18 languages with native code-switching, which means a participant can switch languages mid-sentence without you declaring a language per file — a common situation on international calls. For broader coverage, universal-2 supports 99 languages total at $0.15/hr, with lower tier accuracy than the flagship model. The summarization step is separate: the LLM Gateway models can summarize in a different language than the transcript, so you can transcribe a Spanish meeting and request an English summary in the prompt.