Voice agents need Speech-to-Text (STT) that's both fast and accurate. This benchmark measures what matters: semantic accuracy (can the LLM understand the transcript?) and latency (how quickly is the transcript ready?). Tested on 1,000 real-world samples. Read the methodology →
| # | Vendor⇅ | Model⇅ | Transcripts⇅ | Perfect⇅ | WER Mean⇅ | Pooled WER▲ | TTFS Median⇅ | TTFS P95⇅ | TTFS P99⇅ |
|---|---|---|---|---|---|---|---|---|---|
| 1 | N/A | 99.7% | 83.2% | 1.40% | 🏆 1.07% | 495ms | 676ms | 736ms | |
| 2 | N/A | 🏆 100.0% | 82.9% | 🏆 1.21% | 1.18% | 1016ms | 1345ms | 1791ms | |
| 3 | ink-2 | 🏆 100.0% | 🏆 84.2% | 1.47% | 1.25% | 299ms | 328ms | 1584ms | |
| 4 | stt-rt-v4 | 99.8% | 84.1% | 1.25% | 1.29% | 249ms | 281ms | 310ms | |
| 5 | u3-rt-pro | 99.8% | 83.9% | 1.74% | 1.34% | 335ms | 534ms | 613ms | |
| 6 | nova-3-general | 99.8% | 76.5% | 1.71% | 1.62% | 247ms | 298ms | 326ms | |
| 7 | N/A | 🏆 100.0% | 77.4% | 1.68% | 1.75% | 1136ms | 1527ms | 1897ms | |
| 8 | Nemotron 3.0 ASR (en) | 🏆 100.0% | 76.1% | 1.90% | 1.95% | 🏆 221ms | 🏆 238ms | 🏆 252ms | |
| 9 | pulse | 🏆 100.0% | 72.4% | 2.30% | 2.37% | 398ms | 533ms | 1593ms | |
| 10 | latest-long | 🏆 100.0% | 69.0% | 2.84% | 2.85% | 878ms | 1155ms | 1570ms | |
| 11 | universal-streaming-english | 99.8% | 66.8% | 3.49% | 3.02% | 256ms | 362ms | 417ms | |
| 12 | gpt-4o-transcribe | 99.3% | 75.9% | 3.24% | 3.06% | 637ms | 965ms | 1655ms | |
| 13 | scribe_v2_realtime | 99.7% | 81.3% | 3.16% | 3.12% | 281ms | 348ms | 407ms | |
| 14 | default | 99.8% | 65.1% | 3.56% | 3.71% | 570ms | 596ms | 622ms | |
| 15 | ink-whisper | 99.9% | 60.5% | 3.92% | 4.36% | 266ms | 364ms | 898ms | |
| 16 | Nemotron 3.5 ASR (multilingual) | 99.6% | 62.0% | 4.54% | 4.58% | 236ms | 253ms | 266ms | |
| 17 | voxtral-mini-transcribe-realtime-2602 | 99.3% | 68.8% | 4.44% | 4.97% | 525ms | 973ms | 1913ms |