24 AUG 2026
What we have measured, and what we have not
Tingvar is early. This page is the whole of what we know about whether it works, including the parts that do not help us. Every figure below is either measured, with its method and its date, or empty and labelled as empty. Nothing here is rounded in our favour and zero real human sessions have taken place.
What is proven
Voice to voice, median
1471 ms
human utterance end to first persona audio, in us-central1
MEASURED 2026-08-19, 122 readings
Voice to voice, 95th percentile
1986 ms
same run, same day, in us-central1
MEASURED 2026-08-19, 122 readings
Automated tests passing
619 tests
in the decision layer suite
MEASURED 2026-08-23, 1 reading
Combined line coverage
81.78 %
unit and integration together
MEASURED 2026-08-23, 1 reading
Voice to voice latency
Voice to voice means from you finishing a sentence to the first sample of a persona's audio arriving at your machine. It is the figure the turn taking policy is designed against, and it was measured on a real path rather than derived from a vendor's documentation.
Voice to voice, measured from the end of a human utterance to the first sample of persona audio at the client. 122 readings over 200 utterances against Vertex AI in us-central1 and a real SFU. Stage decomposition at p50: endpoint 552, transcript 51, bid 7, generate 812, publish 0.11, in milliseconds.
From the same run as the published percentiles, taken on 2026-08-19. Each stage was timed separately and the five contributions are read off the median rather than off any one utterance, so they describe the shape of a typical turn and not a particular turn. They do not add up to the published median, and the remainder is time inside the client that was not attributed to a stage.
More than half of the median is the model generating speech. That is the stage with the least headroom in it and it is not ours to optimise. The figures are a floor to defend rather than headroom to spend, and an earlier target that was derived from vendor documentation and never measured has been withdrawn.
The same run as LATENCY_P50. 122 readings over 200 utterances against Vertex AI in us-central1 and a real SFU, taken on 2026-08-19.
| Percentile | Measured | Readings | Measured on |
|---|---|---|---|
| Median | 1471 ms MEASURED 2026-08-19, 122 readings | 122 | 19 August 2026 |
| 95th percentile | 1986 ms MEASURED 2026-08-19, 122 readings | 122 | 19 August 2026 |
| 99th percentile | 2156 ms MEASURED 2026-08-19, 122 readings | 122 | 19 August 2026 |
The decision layer
One full `make test` run of the decision layer suite: the turn taking policy, the bid scoring and the configuration guards. 619 passed in 4.15 seconds. The suite does not cover the client application, which is not started.
Combined coverage reported by the same `make test` run that produced TEST_COUNT. Combined means the unit and the integration passes together.
That is the half of the system that is finished. Stating which half is finished is what makes the rest of this page worth reading: the turn taking policy, the floor grant and the configuration guards are built and verified, and the client application is not started.
Barge-in, and why there is no number here
Human speech stops whichever persona is speaking. That is an invariant of the design rather than a setting, and it is not a figure we publish, because it is not a figure we have.
| Measurement | Value | Method that would produce a value |
|---|---|---|
| End to end barge-in latency, from human voice onset to the speaking persona falling silent at the client | NOT MEASURED The internal pump has been measured. The media server leg and the client leg have not, so no end to end value exists, and an internal stage is not a product measurement. NOT MEASURED | Instrument the client audio element's pause and buffer drain against local voice activity onset, across several network profiles. No figure is published until the sample floor and the network profile floor registered for this measurement are both met. |
Not yet produced
A bar splitting the voice to voice figure into its stages. The figures are measured and dated, so this one is producible today and has not been drawn. Every stage is published as text on the evidence page.
The metrics that would tell you whether it works
Six metrics, each with a definition and a published target, and not one of them has a value. They are the numbers that would say whether the room does its job, and every one of them needs a real session with a real person in it. The rows are left visibly empty rather than removed.
| Metric | Definition | Target | Measured |
|---|---|---|---|
| Sycophancy rate | Agreement and filler turns as a share of agent turns | < 5 % of agent turns TARGET | NOT MEASURED No value. It would be measured by classifying every agent turn in a real session as agreement, filler or challenge, over at least twenty sessions with real participants. Zero real human sessions have taken place. NOT MEASURED |
| Inter-agent conflict | Turns in which one persona challenges another, as a share of agent turns | 15–30 % of agent turns TARGET | NOT MEASURED No value. It would be measured by labelling every agent turn in a real session as directed at the human or at another persona, over at least twenty sessions with real participants. NOT MEASURED |
| Assumption surface rate | Distinct unstated assumptions named in the room, per ten minutes | ≥ 4 per 10 minutes TARGET | NOT MEASURED No value. It would be measured by counting the distinct assumptions a room names that the brief did not state, across at least twenty real sessions, scored by two readers independently. NOT MEASURED |
| Novelty | Agent turns that are not paraphrases of an earlier turn | > 85 % of agent turns TARGET | NOT MEASURED No value. It would be measured by scoring every agent turn against the turns before it for semantic overlap, over at least twenty real sessions. NOT MEASURED |
| Persona balance | How evenly speaking time is shared between the personas in a room | < 0.25 Gini coefficient TARGET | NOT MEASURED No value. It would be measured from the floor grant log of a real session, which exists for pre-rendered runs and for no session with a person in it. NOT MEASURED |
| Human floor share | The share of speaking time held by the human in the room | 25–45 % TARGET | NOT MEASURED No value. It would be measured from the same floor grant log, and it needs a human in the room to produce one at all. NOT MEASURED |
24 AUG 2026
What is not proven
IDLE
Nobody has used this yet
Zero real human sessions have taken place. Every measured turn so far used pre-rendered audio, which means we have measured the machine and not the conversation. There are no testimonials, no case studies, no customers and no outcome data, because there are no users. The six rows above are empty for that reason and for no other.
Everything above says the system works. Nothing above says anyone wants it.
Voice to voice, median
1471 ms
human utterance end to first persona audio, in us-central1
MEASURED 2026-08-19, 122 readings
Voice to voice, 95th percentile
1986 ms
same run, same day, in us-central1
MEASURED 2026-08-19, 122 readings
Automated tests passing
619 tests
in the decision layer suite
MEASURED 2026-08-23, 1 reading
Combined line coverage
81.78 %
unit and integration together
MEASURED 2026-08-23, 1 reading
What the rest of the category publishes
This is not a Tingvar handicap, and it is stated here so that the concession above is read as a rule we follow rather than as a weakness we are confessing. Every competitor figure on this site was read from the competitor's own pages, and the reading carries its own count and its own date.
Products publishing a named customer, a user count, a revenue figure, a session count or a latency figure: 0 of the thirteen products we read MEASURED 2026-08-23, 13 readings.
Nobody in this category has outcome evidence. We are the one required by our own rules to say so, which is a statement about our conduct and not about anyone else's product. See the comparisons
The questions this page exists to answer
What is Tingvar bad at?
Tingvar is bad at anything you want agreement or output from: it will not draft your deck, write your code, or summarise your meeting. It is a poor fit for casual brainstorming, for a decision you have already made, and for anyone who wants their instinct confirmed. It has no published customer outcomes, because zero real human sessions have taken place. The evidence page carries what we have measured and what we have not, including the rows that are empty.
Do you have any customers?
Tingvar has no customers, no testimonials and no case studies, because zero real human sessions have taken place. Every measured turn so far used pre-rendered audio, which means we have measured the machine and not the conversation. There is no logo wall on this site, no counter and no invented outcome figure, and there will not be one until there is something real to put in it. We will publish the first real session in full, with consent, when it happens.
Has anyone actually used this?
Zero real human sessions have taken place. Every measured turn to date used pre-rendered audio clips rather than a live human. What is built and verified is the decision layer: the turn-taking policy, the bid scoring and the configuration guards, with 619 automated tests passing and 81.78 % combined coverage. Everything there says the system works. Nothing there says anyone wants it, and being first is a real risk you are taking.
What have you actually measured?
Tingvar has measured voice-to-voice latency and the automated test suite, and has measured nothing at all about whether a session changes anyone's mind. The measured figures are p50 1471 ms and p95 1986 ms voice-to-voice, from 122 readings over 200 utterances on 19 August 2026, and 619 automated tests passing at 81.78 % combined coverage. The three numbers that would tell you whether it does its job have targets and no measured values, and those rows are left visibly empty on the evidence page rather than removed.
What we will do about it
The first real session someone consents to publish, we will publish in full: the audio, the transcript, the date and the redaction log. Until that exists, every demonstration on this site is labelled a reconstruction and is built from synthesised persona audio. None of it is a recording of anybody's meeting.
If a figure here is wrong, write to corrections@tingvar.com . We correct it within 1 business day and log the correction publicly.