24 August 2026
Voice to voice latency, and what the delay costs
Nobody in this category publishes a latency figure. Here is ours, with the method, the sample, the stage breakdown and what it costs an interruption.
8 minutes to read
We read every performance claim on 8 competing products. The number of them publishing a latency figure of any kind, voice to voice or otherwise, was 0. The strongest word available across all of them is realtime.
So here is ours, with the method, the sample, and the part of it that is bad.
Voice to voice, median
1471 ms
One run against Vertex AI in us-central1
MEASURED 2026-08-19, 122 readings
Voice to voice, 95th percentile
1986 ms
One run against Vertex AI in us-central1
MEASURED 2026-08-19, 122 readings
Voice to voice, 99th percentile
2156 ms
One run against Vertex AI in us-central1
MEASURED 2026-08-19, 122 readings
What was measured, and how
Voice to voice, measured from the end of a human utterance to the first sample of persona audio at the client. 122 readings over 200 utterances against Vertex AI in us-central1 and a real SFU. Stage decomposition at p50: endpoint 552, transcript 51, bid 7, generate 812, publish 0.11, in milliseconds.
Voice to voice means from the moment you stop speaking to the first sample of a persona's audio arriving at your machine. It is the measurement a listener actually experiences, and it is deliberately not the one that flatters us. A time to first token figure would be a great deal smaller and would describe none of the silence you sit in.
The run was taken on 19 August 2026 and it is one run. It has not been repeated, it was not taken across regions, and it was not taken with a human in the room. Every one of those is a reason to treat the figure as provisional, and none of them is a reason to withhold it.
Where the time goes
The interesting part of a latency figure is never the total. It is which stage owns it, because that is what tells you whether the number can move.
| Stage | What happens | milliseconds at p50 |
|---|---|---|
| Endpoint | Deciding the human has stopped speaking | 552 |
| Transcript | Turning the utterance into text | 51 |
| Bid | The room deciding which persona takes the floor | 7 |
| Generate | The model producing the reply as speech | 812 |
| Publish | Putting the first audio sample on the wire | 0.11 |
Two things fall out of that table. The first is that more than half of the median is the model generating speech, which is a vendor's number and not ours to improve by writing better code. The second is that the room deciding who speaks costs almost nothing. The floor arbitration that people assume is the expensive part of a multi speaker system is the cheapest thing in the list by a wide margin.
The stage that is ours, and that is large, is deciding you have finished speaking. Cutting it means guessing earlier, and guessing earlier means interrupting people mid sentence. That trade is a product decision rather than an optimisation, and we have made it in the conservative direction on purpose.
What that delay feels like
The gap between turns in ordinary speech is short enough that nobody counts it. Ours is long enough that you notice, and a listener who notices a pause starts attributing something to it: reluctance, confusion, or disagreement about who should speak.
In a room of advisers that attribution works in our favour, and we did not design for it. A person who has just made an uncomfortable claim and hears a beat of silence before the first objection reads the silence as the room thinking. The same delay from a single assistant reads as software being slow. We would rather have the shorter number, and we are telling you what the longer one buys.
Where it does hurt is the interruption. You speak over a persona, it stops, and then there is the same delay again before anybody responds to what you said. That gap is the one that costs the format something, because the whole promise of speaking is that it is faster than typing.
The figure we withdrew
An earlier target existed and it was much better than the measurement. It was derived from vendor documentation, it was never tested against a running room, and it was wrong. It is gone from every surface on this site and there is no version of this page that still quotes it.
We are saying so here rather than deleting it quietly, because a target read out as though it were a measurement is precisely the failure this site is meant to be the opposite of, and it was ours before it was anybody else's.
What is not measured
The figures above were taken against a real media server with pre-rendered audio standing in for a person. Zero real human sessions have taken place, so nothing here tells you what the delay feels like to somebody who is actually arguing. The barge-in path, from your voice reaching the microphone to the persona going quiet, has not been measured end to end at all, and the evidence page says so in the row where the number would be.
When it is measured it will be published on the same page as this one, with whatever it turns out to be.
Not yet produced
A bar splitting the voice to voice figure into its stages. The figures are measured and dated, so this one is producible today and has not been drawn. Every stage is published as text on the evidence page.