Whisper Large V3
The most performant Whisper Large V3 implementation, achieving 1800x real-time factor (1 hour of audio transcribed in 2 seconds).
Model details
View repositoryOur live transcription is ideal for use cases like live note-taking applications, content captioning, customer support, and any real-time, voice-powered apps. Learn more about our optimized Whisper transcription pipeline in our launch blog.
Feature highlights
Configurable update cadence for the partial transcriptions delivered
Consistent real-time latency under high volume of concurrent audio streams
Automatic language detection for multilingual transcription
Where accuracy stands
The Pareto plot below shows how this model compares with leading closed-source ASR solutions across accuracy and latency. Results are measured using Pipecat’s open-source STT Benchmark, which evaluates Semantic WER for transcription accuracy and TTFS (time to final segment) for latency.
Accuracy vs latency — semantic word-error-rate vs median time-to-final-segment (single stream), plotted against closed-source providers.Concurrency tradeoffs
The graph below shows how TTFS (time to final segment) and unit economics (cost per audio hour) change as the number of concurrent WebSocket connections per GPU increases. Higher concurrency can significantly reduce cost per audio hour, with a corresponding tradeoff in latency.
Concurrency is configurable and can be tuned to achieve the latency and cost targets that best fit your workload and business requirements. See the concurrency guide for configuration details.
For a balance of latency and throughput, we recommend starting with RTX6000 using concurrency 24, and adjusting based on performance and cost needs based on the chart below if needed.
For workload-specific optimization, you can contact an engineer for further assistance.
Latency vs concurrency — how median latency holds up as concurrent streams scale, by GPU.Recommended setups for different use cases:
Balanced
GPU type: H100 MIG
Concurrency target: <= 18
Highly latency-sensitive
GPU type: H100
Concurrency target: <= 32
Cost-sensitive
GPU type: L4
Concurrency target: <= 12