Benchmarks
Hold us to it.
We compare voice systems using the same spoken prompts and each provider's standard turn-taking behaviour. Response latency runs from the end of user speech to the first audible assistant audio. Every table includes the test date and model version so results can be interpreted and reproduced as services change.
Voice-to-voice response latency single turn, n=16 · 2026-07-07
| Provider / model | p50 | p90 |
|---|---|---|
| Converse | 837 ms | 2,367 ms |
| Gemini Live gemini-3.1-flash-live-preview | 1,269 ms | 1,706 ms |
| OpenAI Realtime gpt-realtime-2 | 1,846 ms | 2,570 ms |
| Gemini Live gemini-2.5-flash-native-audio | 2,962 ms | 3,327 ms |
Barge-in: time to silence when interrupted guided live conversations · 2026-07-14
| Provider | p50 time-to-silence | barges stopped |
|---|---|---|
| Converse | 148 ms | 3/3 |
| OpenAI Realtime | 220 ms | 3/3 |
| Gemini Live | 360 ms | 2/2 |
How to read these numbers
These are deliberately small, published samples rather than universal performance guarantees. Network conditions and service load affect latency. Converse is faster at the median in this response test, while Gemini 3.1 has the tighter p90 result.
Speed is only one part of a natural conversation. We publish response and interruption measurements together because a quick reply that cuts someone off is not a better experience.