Voice AI Benchmarks

Compare STT and LLM providers for voice agents

🔄 Updated daily • 17 STT providers • 34 LLM models

Voice agents need Speech-to-Text (STT) that's both fast and accurate. This benchmark measures what matters: semantic accuracy (can the LLM understand the transcript?) and latency (how quickly is the transcript ready?). Tested on 1,000 real-world samples. Read the methodology →

#
Vendor
Model
Transcripts
Perfect
WER Mean
Pooled WER
TTFS Median
TTFS P95
TTFS P99
1speechmaticsSpeechmaticsN/A99.7%83.2%1.40%🏆 1.07%495ms676ms736ms
2azureAzureN/A🏆 100.0%82.9%🏆 1.21%1.18%1016ms1345ms1791ms
3cartesiaCartesiaink-2🏆 100.0%🏆 84.2%1.47%1.25%299ms328ms1584ms
4sonioxSonioxstt-rt-v499.8%84.1%1.25%1.29%249ms281ms310ms
5assemblyaiAssemblyAIu3-rt-pro99.8%83.9%1.74%1.34%335ms534ms613ms
6deepgramDeepgramnova-3-general99.8%76.5%1.71%1.62%247ms298ms326ms
7awsAWSN/A🏆 100.0%77.4%1.68%1.75%1136ms1527ms1897ms
8nvidiaNVIDIANemotron 3.0 ASR (en)🏆 100.0%76.1%1.90%1.95%🏆 221ms🏆 238ms🏆 252ms
9smallestSmallest AIpulse🏆 100.0%72.4%2.30%2.37%398ms533ms1593ms
10googleGooglelatest-long🏆 100.0%69.0%2.84%2.85%878ms1155ms1570ms
11assemblyaiAssemblyAIuniversal-streaming-english99.8%66.8%3.49%3.02%256ms362ms417ms
12openaiOpenAIgpt-4o-transcribe99.3%75.9%3.24%3.06%637ms965ms1655ms
13elevenlabsElevenLabsscribe_v2_realtime99.7%81.3%3.16%3.12%281ms348ms407ms
14gradiumGradiumdefault99.8%65.1%3.56%3.71%570ms596ms622ms
15cartesiaCartesiaink-whisper99.9%60.5%3.92%4.36%266ms364ms898ms
16nvidiaNVIDIANemotron 3.5 ASR (multilingual)99.6%62.0%4.54%4.58%236ms253ms266ms
17mistralMistralvoxtral-mini-transcribe-realtime-260299.3%68.8%4.44%4.97%525ms973ms1913ms