An AI medical scribe is five components in a line: capture the encounter, transcribe it, separate the speakers and pull out the clinical entities, generate a structured note, and hand it to the EHR for signature. Each one is a solved problem. The engineering difficulty is entirely in the seams — what happens when the transcript is 96% right and the 4% is drug names, or when the diarization merges the doctor and the patient into one speaker, or when the note is perfect but contains an unredacted date of birth.
This is the full walkthrough with working code. Not a survey of the category, and not an evaluation checklist — if you want the vendor-selection version, that's the ambient AI scribe evaluation guide, and if you want the argument about which accuracy metric to trust, that's WER vs MER for medical transcription. This post is the build.
Two numbers to anchor on before the code. The transcription layer with clinical accuracy costs $0.36/hr — $0.21/hr for Universal-3.5 Pro plus $0.15/hr for Medical Mode — plus $0.02/hr for diarization, which is about ten cents on a 15-minute encounter. And Medical Mode records a 3.2% Missed Entity Rate, the lowest across the providers on our benchmarks page. Those two facts are why the interesting part of building a scribe in 2026 is the orchestration, not the ASR.
The architecture
Here's the pipeline, and what each stage is actually responsible for.
| Stage | Responsibility | Main failure mode |
|---|---|---|
| 1. Capture | Get clean audio from the room to storage | Mic distance, clipping, dropped segments |
| 2. Transcribe | Audio to text with medical vocabulary intact | Missed or substituted drug and condition names |
| 3. Attribute | Who said what; typed clinical entities | Merged speakers, dropped short turns |
| 4. Structure | Transcript to a specialty-appropriate note | Confident hallucination on a bad transcript |
| 5. Deliver | Redact, persist, push to EHR for signature | Unredacted PHI in retained artifacts |
Stages 4 and 5 get the product attention. Stage 2 sets the ceiling for all of them, which is why it gets the most space here.
Stage 1: capture
Nothing downstream recovers from bad audio, so make two decisions deliberately.
Sample rate and format. 16 kHz mono is enough for speech; anything more is bandwidth you're paying for twice. Uncompressed or losslessly compressed if you can afford the storage — aggressive lossy compression removes exactly the high-frequency consonant detail that separates similar drug names.
Mic placement, and telling the model about it. A badge clip is a different acoustic problem from a desk mic three feet away. For streaming, voice_focus takes near-field or far-field; set it to match the hardware you actually ship, not the laptop you develop on.
Buffer locally and upload with retry. Clinic wifi drops, and losing an encounter is worse than delaying one.
Stage 2: transcription with Medical Mode
This is the whole ballgame. Here's the pre-recorded path, which is what most scribes should use for the note of record.
import assemblyai as aai
aai.settings.api_key = "YOUR_API_KEY"
config = aai.TranscriptionConfig(
speech_models=["universal-3-5-pro"],
domain="medical-v1",
speaker_labels=True,
entity_detection=True,
punctuate=True,
format_text=True,
)
transcriber = aai.Transcriber(config=config)
transcript = transcriber.transcribe("./encounters/enc-40118.wav")
if transcript.status == "error":
raise RuntimeError(transcript.error)
print(transcript.text)
Three things about that config.
speech_models is plural for pre-recorded audio
Pre-recorded requests take speech_models as a list; streaming takes speech_model as a single string. It's a small difference that costs people an afternoon.
Medical Mode is one parameter, not a different model
domain="medical-v1" is the entire activation. You don't switch models, you don't lose diarization or timestamps or entity detection, and you don't inherit a separate feature matrix. It adds $0.15/hr. Against the same base model without it, you get roughly 20% fewer missed medical entities and 87% fewer entity errors — the difference between a note that needs a glance and one that needs a rewrite.
Pass context if you have it
Contextual prompting is the single biggest win most scribe builders skip. Feeding a patient's prior-visit note into the request cut missed medical terms by 31% in an internal healthcare test. Your application has the chart open — it needs it to write the note. Passing it in costs one field and beats maintaining specialty vocabulary lists, which go stale the moment a new drug ships.
Stage 3: speaker attribution and entity extraction
The transcript is prose. What stage 4 needs is who said what, and which clinical concepts appeared.
encounter = {"turns": [], "entities": []}
for utterance in transcript.utterances:
encounter["turns"].append({
"speaker": utterance.speaker,
"start_ms": utterance.start,
"end_ms": utterance.end,
"text": utterance.text,
})
for entity in transcript.entities:
encounter["entities"].append({
"type": entity.entity_type,
"text": entity.text,
"start_ms": entity.start,
})
Mapping speaker labels to roles
Diarization gives you A, B, C — not "clinician" and "patient." You have to resolve roles yourself, and the reliable heuristics are boring and effective: the clinician usually speaks first, asks most of the questions, and accounts for the majority of clinical entities. Combine two or three signals and check them against the label your capture app already has for the logged-in clinician.
Why diarization quality decides your note quality
This is the failure mode teams underestimate. A clinical encounter is short turns, interruptions, and a family member's three words from across the room — exactly what naive diarization drops. And how it's measured matters: Diarization Error Rate scores wrongly-assigned audio time, which rewards long monologues and barely punishes losing a two-word turn. Concatenated minimum-permutation WER (cpWER) scores the words in each speaker's transcript, which is what your note generator consumes.
Universal-3.5 Pro's diarization is optimized for cpWER rather than DER, and it's the most accurate diarization we've shipped, specifically on short turns and overlapped speech. If your note ever attributes a patient's symptom report to the clinician, this is where it came from.
Entity detection gives stage 4 typed inputs
Entity detection covers 50+ types. Handing your note generator a typed list of medications and conditions alongside the prose, rather than prose alone, measurably reduces how often it invents structure that isn't there.
Sign up free and run the code above against one of your own encounter recordings. Medical Mode is one parameter away.
Stage 4: generating the note
Now you turn the structured encounter into a note. The LLM Gateway lets you run this against the transcript without standing up a separate model integration.
import json, requests
prompt = """You are drafting a SOAP note from a clinical encounter transcript.
Rules:
- Use ONLY information present in the transcript. Never infer a diagnosis,
dosage, or finding that was not stated.
- If a required section has no supporting content, write
"Not documented in this encounter."
- Attribute symptoms to the patient and assessments to the clinician,
using the speaker roles provided.
- Quote medication names and dosages exactly as transcribed. Do not
normalize brand to generic.
Detected clinical entities: {entities}
Transcript:
{turns}
"""
payload = prompt.format(
entities=json.dumps(encounter["entities"]),
turns="\n".join(
f'{t["speaker"]}: {t["text"]}' for t in encounter["turns"]
),
)
note = requests.post(
"https://llm-gateway.assemblyai.com/v1/chat/completions",
headers={"authorization": "YOUR_API_KEY"}, json={"model": "claude-sonnet-4-6", "messages": [{"role": "user", "content": payload}]},
)
print(note.json()["choices"][0]["message"]["content"])
The prompt rules are the safety layer
Those four rules aren't stylistic. A language model handed a transcript will fill gaps confidently, and in clinical documentation a plausible invention is worse than an obvious omission — an omission gets caught at review, an invention gets signed. Force explicit "not documented" output, forbid inference, and forbid brand-to-generic normalization. That last one matters because the substitution is often correct and occasionally catastrophically wrong, and you can't tell which from the note.
Structure the output, don't parse prose
Ask for JSON with named sections and validate it against a schema before it goes anywhere near the EHR. A note that fails validation should route to a human, not get patched by another model call.
Stage 5: redaction and delivery
Two things happen here, and teams routinely do only the first.
Redact the transcript and the audio
If you retain the recording, the recording contains a voice saying the patient's name and date of birth. Redacting the transcript alone leaves that in place. Redaction covers both: transcript text at +$0.08/hr and audio at +$0.05/hr, driven by the same 50+ entity types.
redacted_config = aai.TranscriptionConfig(
speech_models=["universal-3-5-pro"],
domain="medical-v1",
speaker_labels=True,
redact_pii=True,
redact_pii_audio=True,
redact_pii_sub="hash",
redact_pii_policies=[
aai.PIIRedactionPolicy.medical_condition,
aai.PIIRedactionPolicy.person_name,
aai.PIIRedactionPolicy.date_of_birth,
aai.PIIRedactionPolicy.phone_number,
aai.PIIRedactionPolicy.email_address,
aai.PIIRedactionPolicy.location,
],
)
Be deliberate about scope. Redacting medical conditions out of the clinical transcript defeats the point — you generally want condition text preserved in the note and redacted in whatever you ship to analytics or long-term storage. That means two transcription passes or a post-processing step, and it's a decision to make on purpose. The full treatment is in our walkthrough on redacting PHI from medical transcripts.
Deliver for signature, not for filing
The note goes to the clinician as a draft. Keep the transcript segment timestamps attached to each note section so a reviewer can jump to the audio behind a claim in one click. That single affordance does more for clinician trust than any accuracy number, because it converts "do I believe this model" into "let me check that line."
Upload an encounter to the playground, turn on Medical Mode, diarization and redaction, and inspect the JSON before you write any orchestration.
The live variant
If the clinician needs a transcript on screen during the encounter, add a streaming path alongside the pre-recorded one. Universal-3.6 Pro Realtime connects at wss://streaming.assemblyai.com/v3/ws for $0.45/hr base, $0.60/hr with Medical Mode.
client = aai.streaming.v3.StreamingClient(
aai.streaming.v3.StreamingClientOptions(
api_key="YOUR_API_KEY",
api_host="streaming.assemblyai.com",
)
)
client.connect(
aai.streaming.v3.StreamingParameters(
sample_rate=16000, speech_model="universal-3-6-pro",
domain="medical-v1",
voice_focus="far-field",
speaker_labels=True,
mode="max_accuracy",
)
)
Singular speech_model here. Use max_accuracy for documentation — you're feeding a display and a note, not a conversational agent, so the latency budget is generous. For clinical audio, raise min_turn_silence to 800ms and max_turn_silence to 3600ms so a clinician pausing mid-sentence does not fragment the turn. Streaming diarization includes revision, correcting earlier speaker assignments as context arrives, for up to 10 speakers.
Running both paths is a legitimate architecture: streaming for the live view, pre-recorded for the record. The pre-recorded pass gets full-context diarization and the best entity accuracy, and it's the cheaper of the two.
What breaks in production
Multilingual encounters. The base model code-switches natively across 18 languages with no configuration. Medical Mode covers English, Spanish, German and French. Those are separate facts — a Spanish-English encounter gets both; a Tagalog-English encounter gets the code-switching without the Medical Mode entity boost. Scope accordingly.
Long encounters. A 90-minute intake produces a transcript that won't fit comfortably in one note-generation call. Chunk on speaker-turn boundaries with overlap, generate section-wise, and reconcile.
Silence and dead air. A recording that starts before the clinician walks in produces minutes of room noise you're paying to transcribe. Trim on the client.
Consent state. Whether the patient consented to recording is application state you have to track, and the pipeline needs to refuse to run without it. This is a product requirement disguised as an infrastructure one.
Compliance, briefly
AssemblyAI signs a Business Associate Addendum (BAA) for customers processing PHI, acting as a business associate under HIPAA, and maintains SOC 2 Type 2. Audio isn't used to train models and isn't shared with third parties. If you're selling into health systems with data residency requirements, ask about EU residency and self-hosted deployment before you architect rather than after. Details are in the BAA FAQ and on the BAA page.
What it costs per encounter
A 15-minute encounter, pre-recorded, with Medical Mode and transcript redaction: roughly twelve cents of speech processing once diarization and entity detection are added. Add the note-generation call and you're still comfortably under fifteen cents. Full rates on the pricing page, and the product surface is on the speech-to-text page and solutions/medical. API reference is in the docs.
Teams shipping clinical documentation on this stack include Commure, Sully AI, Heidi Health, Deepscribe, Knowtex and Magentus Healthcare. Related reading: telehealth speech to text for the virtual care variant.
Conclusion
The five-stage pipeline above is going to get shorter. Stages 3 and 4 — attribution and structuring — are converging, because the same context that improves transcription accuracy also improves note structure, and running them as separate passes throws information away at the boundary. The version of this architecture worth designing toward passes the chart in once and gets back an attributed, entity-typed, section-structured draft in a single round trip. Build the seams loosely enough that you can collapse them when that lands, and keep the timestamp links to source audio no matter what — the model will keep getting better, and the clinician's ability to check it is the part that has to survive every upgrade.
Talk to our team about BAA terms, self-hosted and EU-residency deployment, and volume pricing for scribe workloads.
Frequently asked questions
How do I build an AI medical scribe like Nuance DAX or Abridge?
The architecture is capture, transcribe, attribute speakers and entities, generate a structured note, then redact and deliver for signature. The differentiated work is in stages 4 and 5 — note quality, specialty templates, EHR integration — because the transcription layer is now an API call. Use Universal-3.5 Pro with domain: "medical-v1" for the speech layer and spend your engineering time on the note.
What accuracy do I need before a scribe is usable in a clinic?
Measure Missed Entity Rate rather than word error rate, because entity errors drive the review burden and WER doesn't see them. Universal-3.5 Pro with Medical Mode records 3.2% MER, the lowest across benchmarked providers. That reduces correction time; it doesn't remove the clinician's review and signature, which no accuracy figure does.
Should I use streaming or pre-recorded transcription for a scribe?
Pre-recorded for the note of record — it's cheaper at $0.36/hr with Medical Mode and gets better full-context diarization. Add streaming at $0.60/hr only if the clinician needs a live transcript on screen. Running both is common and reasonable.
Can I use an open-source speech model instead?
You can, and you'll spend the savings on medical vocabulary work, diarization tuning, and the compliance posture — including the fact that a BAA has to come from somewhere. At nine cents per 15-minute encounter, the hosted path is rarely the expensive line item in a scribe product. The exception is a hard data-residency or on-premises requirement, which is worth a direct conversation.
How does speaker diarization separate the doctor from the patient?
Diarization returns anonymous labels — A, B, C — and you map them to roles in your application using signals like who speaks first, who asks the questions, and who produces most of the clinical entities. Universal-3.5 Pro's diarization is optimized for cpWER rather than DER, which is why it holds up on short turns and overlapping speech, and streaming diarization revises earlier assignments as context arrives for up to 10 speakers.
How does AssemblyAI handle HIPAA and PHI?
AssemblyAI signs a Business Associate Addendum (BAA) for customers processing PHI and acts as a business associate under HIPAA. That comes alongside PHI redaction across audio and transcripts, entity detection covering 50+ types, and SOC 2 Type 2. See the BAA FAQ.