New Universal-3.5 Pro is here. Learn more: Async Realtime
LLM Gateway

Run open and frontier LLMs on your voice data

From raw utterance to structured answer—fast open models for the cleanup pass, frontier models for the reasoning pass, one API key for both.

Speech-to-Text

Hi, Marissa. Um, quick recap on where we landed, uh, with the Q3 vendor review. So, uh, we got through 6 of the 8— oh no, 5 of the 8 assessments. Um, and the last 3, they’re waiting on security sign-off. Um, the big thing is that pricing came back higher than, than we modeled, something like, uh, between 12— no, 15% over. So I think we need to revisit The budget line before we commit to anything. Can you pull the original forecast for me? And I’ll put 30 minutes on the calendar.

LLM Gateway

Clean up dictated speech. Remove fillers (um, uh, like, you know), stutters, repetitions, false starts; on self-corrections keep only the final version. Fix punctuation, capitalization, sentence and paragraph breaks. Render spoken commands ("period", "new paragraph") as punctuation.

Final output

Hi Marissa, quick recap on where we landed with the Q3 vendor review. We got through five of the eight assessments. The last three are waiting on security sign-off. The big thing is that pricing came back higher than we modeled, something like 15% over. I think we need to revisit the budget line before we commit to anything. Can you pull the original forecast for me? I'll put 30 minutes on the calendar.

539 ms · 90 tokens · <$0.0001
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
Earmark
ClickUp
HeyGen
Metaview
Dovetail
Granola
Apollo.io
Ashby
Siro
Calabrio
Cluely
Genio
Commure
Retell
CallRail
LiveKit
Earmark
ClickUp
HeyGen
Models

Open and frontier models, one endpoint

37 models from 5 providers, all OpenAI-compatible—switch by changing one string.

Providers0
Regions0
Features0
New models first

Showing 37 of 37 models

LLM Gateway models with context window, input, output and cached token rates per 1 million tokens, capabilities, serving providers and regions.
Capabilities Providers Regions
anthropic/claude-opus-5 New
200K $5.00/M $25.00/M
Read:$0.50/M Write:$6.25/M
Tools Streaming AWS Bedrock USGlobal
google/gemini-3.8-flash New
1M $0.75/M $3.75/M
Read:$0.07/M Write:
JSON Tools Streaming Vertex USEUGlobal
openai/gpt-5.6-sol New
270K $4.00/M $20.00/M
Read:$0.40/M Write:$5.00/M
JSON Tools Streaming Open AI USGlobal
google/gemma-4-31b New
256K $0.14/M $0.40/M
Read: Write:
JSON Tools Streaming AWS Bedrock US
google/gemini-3.7-flash New
1M $0.75/M $3.75/M
Read:$0.07/M Write:
JSON Tools Streaming Vertex USEUGlobal
openai/gpt-5-nano
400K $0.05/M $0.40/M
Read:$0.005/M Write:
JSON Tools Streaming Open AI US
openai/gpt-oss-20b
131K $0.07/M $0.30/M
Read: Write:
Tools AWS Bedrock US
alibaba/qwen3.5-4b-32k-fast
33K $0.10/M $0.50/M
Read: Write:
Streaming AssemblyAI USEU
google/gemini-2.5-flash-lite
1M $0.10/M $0.40/M
Read:$0.01/M Write:
JSON Tools Streaming Vertex USEUGlobal
alibaba/qwen3-32B
200K $0.15/M $0.60/M
Read: Write:
Tools JSON Streaming AWS Bedrock US
alibaba/qwen3-next-80b-a3b
200K $0.15/M $1.20/M
Read: Write:
Tools JSON Streaming AWS Bedrock US
openai/gpt-oss-120b
131K $0.15/M $0.60/M
Read: Write:
JSON Tools AWS Bedrock US
google/gemini-3.1-flash-lite
1M $0.25/M $1.50/M
Read:$0.03/M Write:
JSON Tools Streaming Vertex USGlobal
openai/gpt-5-mini
400K $0.25/M $2.00/M
Read:$0.03/M Write:
JSON Tools Streaming Open AI US
google/gemini-2.5-flash
1M $0.30/M $2.50/M
Read:$0.03/M Write:
JSON Tools Streaming Vertex USEUGlobal
google/gemini-3.5-flash-lite
1M $0.30/M $2.50/M
Read:$0.03/M Write:
JSON Tools Streaming Vertex USGlobal
minimax/minimax-m3
512K $0.30/M $1.20/M
Read:$0.06/M Write:
Tools JSON Streaming Fireworks US
anthropic/claude-haiku-4-5-20251001
200K $1.00/M $5.00/M
Read:$0.10/M Write:$1.25/M
Tools JSON Streaming AWS Bedrock USEUGlobal
openai/gpt-5.6-luna
270K $1.00/M $6.00/M
Read:$0.10/M Write:$1.25/M
JSON Tools Streaming Open AI USGlobal
google/gemini-2.5-pro
200K $1.25/M $10.00/M
Read:$0.13/M Write:
JSON Tools Streaming Vertex USEUGlobal
google/gemini-3.5-flash
1M $1.25/M $9.00/M
Read:$0.13/M Write:
JSON Tools Streaming Vertex USGlobal
openai/gpt-5
400K $1.25/M $10.00/M
Read:$0.13/M Write:
Tools Streaming JSON Open AI US
openai/gpt-5.1
400K $1.25/M $10.00/M
Read:$0.13/M Write:
JSON Tools Streaming Open AI US
google/gemini-3.6-flash
1M $1.50/M $7.50/M
Read:$0.15/M Write:
JSON Tools Streaming Vertex USEUGlobal
openai/gpt-5.2
400K $1.75/M $14.00/M
Read:$0.17/M Write:
JSON Tools Streaming Open AI US
openai/gpt-4.1
1M $2.00/M $8.00/M
Read:$0.50/M Write:
Tools Streaming Open AI US
openai/gpt-5.6-terra
270K $2.50/M $15.00/M
Read:$0.25/M Write:$3.13/M
JSON Tools Streaming Open AI USGlobal
anthropic/claude-sonnet-4-5-20250929
200K $3.00/M $15.00/M
Read:$0.30/M Write:$3.75/M
Tools JSON Streaming AWS Bedrock USEUGlobal
anthropic/claude-sonnet-4-6
200K $3.00/M $15.00/M
Read:$0.30/M Write:$3.75/M
Tools JSON Streaming AWS Bedrock USEUGlobal
anthropic/claude-sonnet-5
200K $3.00/M $15.00/M
Read:$0.30/M Write:$3.75/M
Tools Streaming AWS Bedrock USGlobal
moonshot-ai/kimi-k3
1M $3.00/M $15.00/M
Read:$0.30/M Write:
Tools JSON Streaming Fireworks US
anthropic/claude-opus-4-5-20251101
200K $5.00/M $25.00/M
Read:$0.50/M Write:$6.25/M
Tools JSON Streaming AWS Bedrock USGlobal
anthropic/claude-opus-4-6
200K $5.00/M $25.00/M
Read:$0.50/M Write:$6.25/M
Tools JSON Streaming AWS Bedrock USGlobal
anthropic/claude-opus-4-7
1M $5.00/M $25.00/M
Read:$0.50/M Write:$6.25/M
Tools Streaming AWS Bedrock USGlobal
anthropic/claude-opus-4-8
1M $5.00/M $25.00/M
Read:$0.50/M Write:$6.25/M
Tools Streaming AWS Bedrock US
openai/gpt-5.5
272K $5.00/M $30.00/M
Read:$0.50/M Write:
JSON Tools Streaming Open AI USGlobal
openai/gpt-6-astra
270K $10.00/M $50.00/M
Read:$1.00/M Write:$12.50/M
Tools Streaming Open AIAWS Bedrock US

Prices are for global routing. In-region (US/EU) pricing is 10% higher due to provider cost increases. Names read maker/model; the part after the slash is the model value to send. See the full catalog.

Production grade

Easiest, most reliable way to call multiple LLMs

Ship faster, spend less on tokens, and stop losing users to provider outages.

0% markup

Pay provider rates, not gateway rates. Competing gateways add 5% or more to every call.

Automatic fallbacks

Configure backup models per request. When a provider errors or stalls, your call still goes through.

Security by simplicity

Pay the exact same price as calling the model provider directly. No markup, no hidden fees, no minimum commitment. We make it simple.

OpenAI-compatible

Drop into any OpenAI SDK. Change a base URL and a model string, and everything keeps working.

Voice-native

Your LLM calls run where your transcription does. One less network hop on every turn.

Models worth using

Frontier models from OpenAI, Anthropic, and Google, plus open models we host ourselves. New ones added the day they launch.

Common questions