---
url: https://tingvar.com/how-it-works
title: "How it works · Tingvar"
updated: 2026-09-12
---

# The room decides who speaks. No persona decides for itself.

Exactly one persona holds the floor at any moment. The room decides who gets it, not the personas. Your voice takes it from all of them. This page is the mechanism behind those three sentences, with the figures we have measured and the ones we have not.

Start a session

## Twelve seconds of a room, written out

01 VICTOR VANCE

02 DR ELENA ROSTOVA

03 MARCUS THORNE

04 PRIYA RAMAN

05 YOU

Scroll to see the rest

RECONSTRUCTION Four agent lanes over twelve seconds, above a taller human lane. Each lane's envelope is measured from that persona's own voice sample. The third lane is cut off by a hard vertical edge at the moment the human lane begins. The arrangement of turns is composed and is not a recording of a session. Each envelope is measured from that persona's own voice sample. The voices are real and the words are written. The arrangement of turns in the strip is composed, is labelled a reconstruction, and is not a recording of a session.

The third lane stops mid word. That vertical edge is the whole mechanism in one mark: the human started speaking and the persona lost the floor inside the sentence rather than at the end of it. The transcript below marks the same event in text, because position and colour are not available to every reader.

RECONSTRUCTION composed arrangement

VICTOR VANCE 00:00

> Strong is not a number. Give me the churn figure and who measured it.

DR ELENA ROSTOVA 00:03

> It holds until the first bad Tuesday. Then you pay for the outage and the rewrite.

MARCUS THORNE 00:05

> There is no precedent, so assume no demand. What does it cost a buyer to

Cut off here

YOU 00:07

> Stop. Price the churn first and come back to acquisition after.

You spoke here

PRIYA RAMAN 00:09

> Name who ships it. If it is this hard to schedule, it is probably the wrong shape.

Not yet produced

A whole session end to end, with chapters, so you can hear the argument change its mind rather than read that it did. No session has been recorded, and the 40 second reconstruction is not one.

Blocked on one real recorded session 2026-08-26

## What we measured before publishing any of this

Voice to voice

1471 ms

Median, 1986 ms at the 95th percentile, on the same run, in us-central1

MEASURED 2026-08-19, 122 readings

How we measured it

Real human sessions

0 sessions

Have taken place. Every measured turn so far used pre-rendered audio

MEASURED 2026-08-23, 1 reading

What happens next

[See the evidence](/evidence)

Not yet produced

A bar splitting the voice to voice figure into its stages. The figures are measured and dated, so this one is producible today and has not been drawn. Every stage is published as text on the evidence page.

Blocked on design production 2026-08-26

## Three prompts in one window is a different machine

Running three characters inside one chat window looks like this from the outside and behaves nothing like it. The three cannot interrupt each other, you cannot interrupt them, and nothing outside them decides who answers. The table states the differences as behaviour rather than as architecture, because behaviour is what a reader can check.

| Can it do this? | One assistant | Three prompts in one window | An arbitrated room |
| --- | --- | --- | --- |
| One party can interrupt another mid sentence | No | No | Yes |
| You can interrupt it with your voice | No | No | Yes |
| Something outside the speakers decides who speaks | No | No | Yes |
| The parties hold opposed objectives by design | No | No | Yes |

Three arrangements of the same models, and what each one can and cannot do

Scroll to see the rest

[The full comparison](/compare/chatgpt)

Not yet produced

A static figure putting the three arrangements side by side. It is producible today and it has not been drawn. Both pages carry the same three states as a table, so the gap costs a reader the picture.

Blocked on design production 2026-08-26

## What happens, in order

01

### You write a brief

A brief is the written proposal you supply before the room opens: the plan you want put under pressure. It is not a prompt. What you are proposing, what it costs, what you are assuming, and what would make it fail. The personas have read it before they arrive, so they arrive already hostile to the right things. Sessions that start from a written proposal are markedly better than cold starts. That is our own internal observation rather than a measurement, and the ledger above says why: no real human session has taken place to measure it against.

[See what a brief asks for](/brief)

02

### The personas arrive with opposed objectives

They differ in what they optimise for rather than in what they know, which is why the financial one and the growth one reach opposite conclusions from the same page. Each carries a deliberate blind spot the others can attack.

03

### One speaker holds the floor, and the room decides who

Three invariants, in plain English, with no implementation in them.

1. Exactly one persona holds the floor at any moment.
2. No persona decides on its own to speak. Without that rule every persona answers every human utterance at once and the room is noise.
3. Human speech preempts everything. The floor is held for a bounded period, it is released, and it can be taken from whoever has it.

04

### You cut in

You will not have to wait for a gap. Start talking and whoever is speaking stops. This is the part two years of turn based assistants have trained people to disbelieve, which is why the strip above draws the interruption rather than describing it.

The client leg of an interruption is not yet measured. The method that would measure it is written down: instrument the client audio element's pause and buffer drain against local voice activity onset, over at least 100 barge-in events and at least 3 network profiles.

05

### It answers, and here is how fast

Voice to voice

1471 ms

Median, 1986 ms at the 95th percentile, on the same run, in us-central1

MEASURED 2026-08-19, 122 readings

Voice to voice, measured from the end of a human utterance to the first sample of persona audio at the client. 122 readings over 200 utterances against Vertex AI in us-central1 and a real SFU. Stage decomposition at p50: endpoint 552, transcript 51, bid 7, generate 812, publish 0.11, in milliseconds.

More than half of that is the model generating speech. It is not instant, and a figure with its method attached is worth more than an adjective.

[What that actually feels like](/blog/what-latency-feels-like)

[Where that median goes, stage by stage](/evidence)

06

### You follow every voice in the room

Each persona has a distinct synthesised voice and a fixed position in the stereo field, and the live transcript labels every line with who said it. A mono mode exists, because a single earbud or a laptop speaker makes panning useless and you should not be punished for either. Four personas is the practical ceiling for the same reason.

07

### You leave with a transcript and one question

A transcript, every line attributed to whoever said it, and one question you answer yourself: name one thing you now believe about this proposal that you did not believe an hour ago. An empty answer is a failed session regardless of what the metrics say.

The transcript is exportable on every tier. A written post-session brief is a v1.5 item and does not ship today.

Not yet produced

A picture of the room in use. There is no room to photograph. A design render labelled as one is allowed on this page and has not been made, and it would never be allowed in a store listing.

Blocked on the client application, which has not been started 2026-08-26

Not yet produced

A picture of what you keep after a session: the transcript as the product renders it. There is no screen to photograph, and a design render dressed as a screenshot is the first thing we asked you not to trust us about.

Blocked on the client application, which has not been started 2026-08-26

## Who is in the room

| Lane | Name | Role | Optimises for | Blind spot |
| --- | --- | --- | --- | --- |
| 01 | VICTOR VANCE | Chief financial officer and private equity operator | Cash survival and the unit economics underneath every claim | Undervalues a bet whose payoff arrives after the current cycle |
| 02 | DR ELENA ROSTOVA | Principal engineer, distributed systems | What the system does on the day the load is not the happy path | Builds for a failure mode that the traffic has not yet produced |
| 03 | MARCUS THORNE | Head of go to market | Whether a buyer moves, and what it costs that buyer to leave | Reads the absence of precedent as the absence of any demand |
| 04 | PRIYA RAMAN OPTIONAL FOURTH | Delivery lead | Who ships this, on what week, with the people already here | Treats a plan that is hard to schedule as a plan that is wrong |

Four personas, their lane position, what each optimises for and what each misses. The blind spot is published because it was designed, and the others are meant to find it.

01

VICTOR VANCE

Role Chief financial officer and private equity operator

Optimises for Cash survival and the unit economics underneath every claim

Blind spot Undervalues a bet whose payoff arrives after the current cycle

Voice

02

DR ELENA ROSTOVA

Role Principal engineer, distributed systems

Optimises for What the system does on the day the load is not the happy path

Blind spot Builds for a failure mode that the traffic has not yet produced

Voice

03

MARCUS THORNE

Role Head of go to market

Optimises for Whether a buyer moves, and what it costs that buyer to leave

Blind spot Reads the absence of precedent as the absence of any demand

Voice

04

PRIYA RAMAN

OPTIONAL FOURTH

Role Delivery lead

Optimises for Who ships this, on what week, with the people already here

Blind spot Treats a plan that is hard to schedule as a plan that is wrong

Voice

[Meet the four](/personas)

## When a persona invents a number

Caution Will it invent numbers about my business?

A Tingvar persona once turned the misheard phrase "turns in" into a confident objection about 15 % monthly churn that had never been said. We publish this because a room whose value is pressure-testing your numbers, and which invents numbers, is worse than no room at all. The personas argue from the brief you write and from what you say, and they cannot see your accounts or anything you have not told them. If a number sounds wrong, say so out loud.

We publish that because a room whose value is pressure testing your numbers, and which invents numbers, is worse than no room at all. What we can tell you is what the personas argue from, which is your brief and what you say out loud, and what they cannot see, which is your accounts and anything you have not told them. What we cannot yet tell you is a rate, because that needs real sessions and there have been none.

If a number sounds wrong, say so out loud. Speech is how you steer, and correcting a misheard figure is the same control as changing the subject.

[AI disclosure](/ai-disclosure)

## The questions this page gets asked

Why voice, and wouldn't text be easier to read?

Tingvar uses voice because interruption is the primary steering control, and interruption is native only in speech. In text you wait for a paragraph to finish. In a room you cut in the moment a persona goes wrong, and it stops. Spoken debate is also faster than text and forces brevity on the personas. You still get the record: a live transcript labels every line with the name of whoever said it, and it is exportable.

What do I actually get at the end?

A Tingvar session ends with a full transcript, every line attributed to the persona who said it, and one question you have to answer yourself. The question is: name one thing you now believe about this proposal that you did not believe an hour ago. The transcript is exportable on every tier. A written post-session brief is a v1.5 item and does not ship today. There is no verdict and no decision memo, because the argument is the output.

Can I interrupt it?

Yes, speaking is how you steer a Tingvar session, and your voice stops whichever persona is talking rather than queueing behind it. Human speech preempts everything in the room, and that is an invariant of the design rather than a setting you switch on. Speaking is also how you change the subject, ask a persona to defend a claim, or stop a line of argument that is going nowhere. The client leg of barge-in is not yet measured, and we publish measurements rather than targets.

Will they talk over each other?

Exactly one persona holds the floor in a Tingvar room at any moment, and the room decides who holds it rather than each persona deciding for itself. Server-side voice activity detection is deliberately switched off, so no persona ever decides on its own to speak. Without that, three personas answer every human utterance at once and the room is noise. The floor is held for a bounded period and released, and it can be taken from whoever has it.

How fast does it respond?

Tingvar's measured voice-to-voice latency is p50 1471 ms and p95 1986 ms, from 122 readings over 200 utterances taken on 19 August 2026. p99 is 2156 ms. Voice-to-voice means from you finishing a sentence to the first persona's audio starting. The readings were taken against Vertex AI in us-central1 with a real media server. More than half of that time is the model generating speech, and the stage by stage breakdown is on the evidence page. It is not instant.

How will I know who is speaking?

Each Tingvar persona has a distinct synthesised voice and a fixed position in the stereo field, and a live transcript labels every line with who said it. Separating voices across the stereo field makes rapid speaker changes easier to follow and easier to attribute afterwards. A mono fallback exists, because a single earbud or a laptop speaker makes panning useless. Four personas is the ceiling for the same reason: past four, a listener stops being able to attribute what was just said.

Do I have to prepare anything?

Yes, Tingvar asks for a written brief before the room opens, and sessions that start from a written proposal are markedly better than cold starts. The brief is the proposal you want put under pressure: what you are proposing, what it costs, what you are assuming and what would make it fail. It is a document the personas have already read when they arrive. The cold start comparison is our own observation from internal testing, not a measurement.

Do I need headphones, and where am I supposed to do this?

Tingvar needs headphones and a room where you can talk out loud for 45 minutes, because it is a spoken conversation and you will be interrupting people. Open plan is the honest problem here, and we do not have a good answer for it beyond a closed room or a quiet hour. A mono fallback exists for a single earbud, though you lose the stereo separation that makes speakers easier to tell apart. This is a hypothesis rather than something we have watched people do.

[FAQ](/faq)

## Where the recording goes

Where the audio goes

To Google for speech recognition and synthesis, and through a media server we run. It is processed while the room is live, it is not retained, and it does not stay on your device.

What we keep

The brief and the transcript. Audio is discarded once the transcript is written.

Who can read it

You. Nobody at Tingvar reads, inspects or processes a transcript without your express written request on an active support case. Every authorised administrative access is time bounded, restricted to the session you named, and recorded in a tamper evident audit log.

How you destroy it

One control on the account page. It removes the brief, the transcript and the account together, and backups are purged within 30 days.

[Your data](/trust/your-data)

## Write a brief. Join the room. Cut in whenever you like.

A session is a complete room rather than a taster, and your first one carries the guarantee. The three commitments below are the ones you can check.

Refund

End your first session inside 5 minutes and we refund it in full.

Cancellation

3 clicks or fewer from the account page, with no retention screen in the way.

Export

Every transcript, on every tier.

Start a session

Book a demo
