VICTOR VANCE 00:00
Strong is not a number. Give me the churn figure and who measured it.
01 VICTOR VANCE
02 DR ELENA ROSTOVA
03 MARCUS THORNE
04 PRIYA RAMAN
05 YOU
The third lane stops mid word. That vertical edge is the whole mechanism in one mark: the human started speaking and the persona lost the floor inside the sentence rather than at the end of it. The transcript below marks the same event in text, because position and colour are not available to every reader.
RECONSTRUCTION composed arrangement
VICTOR VANCE 00:00
Strong is not a number. Give me the churn figure and who measured it.
DR ELENA ROSTOVA 00:03
It holds until the first bad Tuesday. Then you pay for the outage and the rewrite.
MARCUS THORNE 00:05
There is no precedent, so assume no demand. What does it cost a buyer to
Cut off here
YOU 00:07
Stop. Price the churn first and come back to acquisition after.
You spoke here
PRIYA RAMAN 00:09
Name who ships it. If it is this hard to schedule, it is probably the wrong shape.
Not yet produced
A whole session end to end, with chapters, so you can hear the argument change its mind rather than read that it did. No session has been recorded, and the 40 second reconstruction is not one.
Voice to voice
1471 ms
Median, 1986 ms at the 95th percentile, on the same run, in us-central1
MEASURED 2026-08-19, 122 readings
Real human sessions
0 sessions
Have taken place. Every measured turn so far used pre-rendered audio
MEASURED 2026-08-23, 1 reading
Not yet produced
A bar splitting the voice to voice figure into its stages. The figures are measured and dated, so this one is producible today and has not been drawn. Every stage is published as text on the evidence page.
Running three characters inside one chat window looks like this from the outside and behaves nothing like it. The three cannot interrupt each other, you cannot interrupt them, and nothing outside them decides who answers. The table states the differences as behaviour rather than as architecture, because behaviour is what a reader can check.
| Can it do this? | One assistant | Three prompts in one window | An arbitrated room |
|---|---|---|---|
| One party can interrupt another mid sentence | No | No | Yes |
| You can interrupt it with your voice | No | No | Yes |
| Something outside the speakers decides who speaks | No | No | Yes |
| The parties hold opposed objectives by design | No | No | Yes |
Not yet produced
A static figure putting the three arrangements side by side. It is producible today and it has not been drawn. Both pages carry the same three states as a table, so the gap costs a reader the picture.
01
A brief is the written proposal you supply before the room opens: the plan you want put under pressure. It is not a prompt. What you are proposing, what it costs, what you are assuming, and what would make it fail. The personas have read it before they arrive, so they arrive already hostile to the right things. Sessions that start from a written proposal are markedly better than cold starts. That is our own internal observation rather than a measurement, and the ledger above says why: no real human session has taken place to measure it against.
02
They differ in what they optimise for rather than in what they know, which is why the financial one and the growth one reach opposite conclusions from the same page. Each carries a deliberate blind spot the others can attack.
03
Three invariants, in plain English, with no implementation in them.
04
You will not have to wait for a gap. Start talking and whoever is speaking stops. This is the part two years of turn based assistants have trained people to disbelieve, which is why the strip above draws the interruption rather than describing it.
The client leg of an interruption is not yet measured. The method that would measure it is written down: instrument the client audio element's pause and buffer drain against local voice activity onset, over at least 100 barge-in events and at least 3 network profiles.
05
Voice to voice
1471 ms
Median, 1986 ms at the 95th percentile, on the same run, in us-central1
MEASURED 2026-08-19, 122 readings
Voice to voice, measured from the end of a human utterance to the first sample of persona audio at the client. 122 readings over 200 utterances against Vertex AI in us-central1 and a real SFU. Stage decomposition at p50: endpoint 552, transcript 51, bid 7, generate 812, publish 0.11, in milliseconds.
More than half of that is the model generating speech. It is not instant, and a figure with its method attached is worth more than an adjective.
06
Each persona has a distinct synthesised voice and a fixed position in the stereo field, and the live transcript labels every line with who said it. A mono mode exists, because a single earbud or a laptop speaker makes panning useless and you should not be punished for either. Four personas is the practical ceiling for the same reason.
07
A transcript, every line attributed to whoever said it, and one question you answer yourself: name one thing you now believe about this proposal that you did not believe an hour ago. An empty answer is a failed session regardless of what the metrics say.
The transcript is exportable on every tier. A written post-session brief is a v1.5 item and does not ship today.
Not yet produced
A picture of the room in use. There is no room to photograph. A design render labelled as one is allowed on this page and has not been made, and it would never be allowed in a store listing.
Not yet produced
A picture of what you keep after a session: the transcript as the product renders it. There is no screen to photograph, and a design render dressed as a screenshot is the first thing we asked you not to trust us about.
| Lane | Name | Role | Optimises for | Blind spot |
|---|---|---|---|---|
| 01 | VICTOR VANCE | Chief financial officer and private equity operator | Cash survival and the unit economics underneath every claim | Undervalues a bet whose payoff arrives after the current cycle |
| 02 | DR ELENA ROSTOVA | Principal engineer, distributed systems | What the system does on the day the load is not the happy path | Builds for a failure mode that the traffic has not yet produced |
| 03 | MARCUS THORNE | Head of go to market | Whether a buyer moves, and what it costs that buyer to leave | Reads the absence of precedent as the absence of any demand |
| 04 | PRIYA RAMAN OPTIONAL FOURTH | Delivery lead | Who ships this, on what week, with the people already here | Treats a plan that is hard to schedule as a plan that is wrong |
We publish that because a room whose value is pressure testing your numbers, and which invents numbers, is worse than no room at all. What we can tell you is what the personas argue from, which is your brief and what you say out loud, and what they cannot see, which is your accounts and anything you have not told them. What we cannot yet tell you is a rate, because that needs real sessions and there have been none.
If a number sounds wrong, say so out loud. Speech is how you steer, and correcting a misheard figure is the same control as changing the subject.
Tingvar uses voice because interruption is the primary steering control, and interruption is native only in speech. In text you wait for a paragraph to finish. In a room you cut in the moment a persona goes wrong, and it stops. Spoken debate is also faster than text and forces brevity on the personas. You still get the record: a live transcript labels every line with the name of whoever said it, and it is exportable.
A Tingvar session ends with a full transcript, every line attributed to the persona who said it, and one question you have to answer yourself. The question is: name one thing you now believe about this proposal that you did not believe an hour ago. The transcript is exportable on every tier. A written post-session brief is a v1.5 item and does not ship today. There is no verdict and no decision memo, because the argument is the output.
Yes, speaking is how you steer a Tingvar session, and your voice stops whichever persona is talking rather than queueing behind it. Human speech preempts everything in the room, and that is an invariant of the design rather than a setting you switch on. Speaking is also how you change the subject, ask a persona to defend a claim, or stop a line of argument that is going nowhere. The client leg of barge-in is not yet measured, and we publish measurements rather than targets.
Exactly one persona holds the floor in a Tingvar room at any moment, and the room decides who holds it rather than each persona deciding for itself. Server-side voice activity detection is deliberately switched off, so no persona ever decides on its own to speak. Without that, three personas answer every human utterance at once and the room is noise. The floor is held for a bounded period and released, and it can be taken from whoever has it.
Tingvar's measured voice-to-voice latency is p50 1471 ms and p95 1986 ms, from 122 readings over 200 utterances taken on 19 August 2026. p99 is 2156 ms. Voice-to-voice means from you finishing a sentence to the first persona's audio starting. The readings were taken against Vertex AI in us-central1 with a real media server. More than half of that time is the model generating speech, and the stage by stage breakdown is on the evidence page. It is not instant.
Each Tingvar persona has a distinct synthesised voice and a fixed position in the stereo field, and a live transcript labels every line with who said it. Separating voices across the stereo field makes rapid speaker changes easier to follow and easier to attribute afterwards. A mono fallback exists, because a single earbud or a laptop speaker makes panning useless. Four personas is the ceiling for the same reason: past four, a listener stops being able to attribute what was just said.
Yes, Tingvar asks for a written brief before the room opens, and sessions that start from a written proposal are markedly better than cold starts. The brief is the proposal you want put under pressure: what you are proposing, what it costs, what you are assuming and what would make it fail. It is a document the personas have already read when they arrive. The cold start comparison is our own observation from internal testing, not a measurement.
Tingvar needs headphones and a room where you can talk out loud for 45 minutes, because it is a spoken conversation and you will be interrupting people. Open plan is the honest problem here, and we do not have a good answer for it beyond a closed room or a quiet hour. A mono fallback exists for a single earbud, though you lose the stereo separation that makes speakers easier to tell apart. This is a hypothesis rather than something we have watched people do.
Where the audio goes
To Google for speech recognition and synthesis, and through a media server we run. It is processed while the room is live, it is not retained, and it does not stay on your device.
What we keep
The brief and the transcript. Audio is discarded once the transcript is written.
Who can read it
You. Nobody at Tingvar reads, inspects or processes a transcript without your express written request on an active support case. Every authorised administrative access is time bounded, restricted to the session you named, and recorded in a tamper evident audit log.
How you destroy it
One control on the account page. It removes the brief, the transcript and the account together, and backups are purged within 30 days.
A session is a complete room rather than a taster, and your first one carries the guarantee. The three commitments below are the ones you can check.
Refund
End your first session inside 5 minutes and we refund it in full.
Cancellation
3 clicks or fewer from the account page, with no retention screen in the way.
Export
Every transcript, on every tier.