24 AUG 2026
Tingvar vs building it yourself
You are right about the easy part. Multi agent debate in text is an afternoon. The floor is the part that is not.
Where they are ahead of us
You can wire three agents with three system prompts to a model in an afternoon, run them over a document, and get something genuinely useful out of it. It will cost you nothing per seat, it will run inside your own network, you can change a prompt in a second, and you can point it at whatever data you like. None of that is available from us.
Take that list one item at a time, because each one is a real advantage and none of them stops being true later. Nothing per seat means the marginal cost of another reader is the tokens that reader spends, and there is no invoice from us in the way of it. Inside your own network means the document never leaves it, which is a direct answer to the procurement question we answer with a data processing agreement and a sub-processor list instead. A prompt you can change in a second means the cast is yours: you can replace one agent between two runs on a Tuesday, where our personas stay as they are until we ship different ones. Pointing it at whatever data you like means your own repositories, your own tickets and your own warehouse, none of which we read, because what we take is a written brief and nothing else. The afternoon is not an underestimate either. A loop over three prompts with a shared document is a small amount of code, the libraries for it are already on your machine, and none of it gets harder when you add a fourth agent, because text does not overlap and nobody is kept waiting on it.
There is a second reader this page is not written for. If your team already runs a speech stack in production, the part we are about to call hard is a part you have paid for once already. Streaming audio, endpointing, a media server and a voice that stops when somebody talks over it are engineering you have done, and the argument below reads differently to you than it does to somebody starting from a text harness. We are not going to claim the floor is hard for everybody on the grounds that it was hard for us.
If a written debate over a written plan is what you want, the honest answer is to build it. This page is not about that version. It is about what happens when you try to make the same thing speak.
Turn-taking is trivial in text and is the whole problem in speech
In text, turn-taking is trivial: nothing overlaps and nobody is kept waiting. In speech it is the hard part, and it is where the latency budget is spent. Agent level turn-taking, where each agent decides for itself whether to speak, stops working at three agents, because three personas then answer every human utterance at once and the room is noise. Room level turn-taking can see every persona's intent before it grants the floor to one of them. It is the core of what we build, and it is the reason a room sounds like a conversation.
The budget is what makes it hard rather than merely fiddly. Voice to voice, from a person finishing a sentence to the first sample of agent audio, has a median of 1471 ms MEASURED 2026-08-19, 122 readings in our own measured run in us-central1. More than half of that is the model generating speech. Whatever decides who speaks next has single digit milliseconds to work in, which rules out asking a model to decide, and asking a model to decide is the design every orchestration library reaches for first.
The arbiter is the room level policy that decides which persona takes the floor next. The arbiter is not a persona. It has no opinion about your plan, it never speaks, and it is not a fifth voice settling the argument on your behalf. It is pure policy over events, with no network access and no audio, which is why it can be run and tested offline. Each persona bids to speak, the bids are scored locally against sentence embeddings rather than by a model call, and the arbiter grants the floor to one of them. Agent level turn-taking, where each agent decides for itself whether to speak, stops working at three agents.
Agents differentiated by role converge
This is the failure that costs the most and shows up last. Give three agents three job titles and one document and they will agree with each other, politely and at length, because all three are optimising for being a good expert about the same evidence.
It has been published. Crucible, by Roundtable Labs, differentiated its agents by domain expertise and released two complete debate records through a public interface. Sixteen agents across those two debates produced sixteen supporting positions and no opposing ones. Its own advanced usage guide shipped a dissent quota, described as ensuring a minimum level of disagreement, and told the customer to add an agent that would challenge assumptions.
A harness that needs a quota on disagreement is one in which disagreement does not arise on its own. The fix is not a quota. It is giving each persona something different to protect, and a blind spot the others can attack. Read the thesis
Row by row
| A text harness you write | A spoken room | |
|---|---|---|
| Time to a working version | An afternoon, and it will be genuinely useful | The audio path alone is weeks before anyone hears anything |
| Who speaks next | A loop over a list, or a planner call | A room level decision made again after every utterance, from every agent's bid |
| What happens when three agents all want to speak | Nothing. Text does not overlap and nobody is kept waiting | They talk over each other, and the room is noise |
| Cost of a wrong turn-taking decision | A paragraph in the wrong order, which a reader skips | A collision a listener cannot unhear, in real time |
| Budget for the decision | No budget. The reader is not waiting on it | Milliseconds, inside a voice to voice budget that is already mostly the model |
| Interruption | Not applicable. There is nothing in progress | Human speech preempts everything, and it is an invariant rather than a setting |
| Two agents believing they hold the floor | Cannot happen | Prevented. Exactly one persona holds the floor, and it is an invariant rather than a setting |
| What differentiates the agents | Usually a role and a system prompt | An objective function and a published blind spot |
| What happens when they agree anyway | You add a devil's advocate, or a quota on disagreement | It is a measured failure, against a published target with no value yet |
| Evaluation | You read the output and judge it | Six metrics defined, six targets published, no measured values |
When to build it, and when not to
| Build it | Do not build it | |
|---|---|---|
| Whether anyone is present while it runs | Nobody needs to be in the room while it happens | You want it spoken, and you want to cut in |
| What comes out of it | Written output over documents you cannot send anywhere | Talk, heard while it happens, rather than a document read afterwards |
| How often the cast changes | You want to change the personas every week | You do not need to reshape the cast every week |
| Where it runs | Inside your own network, with your own model keys | A hosted service is acceptable and running it yourself is not the point |
| Deciding who speaks next | A loop or a planner call, because nothing overlaps in text | You would have to solve floor arbitration inside a millisecond budget |
| Where disagreement comes from | A devil's advocate or a quota is enough for what you are doing | You want the disagreement to arise rather than be mandated |
Questions engineers ask first
How is this different from three ChatGPT tabs?
Three ChatGPT tabs give you three answers, and a Tingvar room gives you an argument, because separate tabs cannot hear each other. Nothing in tab two knows what tab one said unless you paste it there yourself, and by then you have chosen which criticism survives. Inside one chat window the instruction to be critical decays after two or three turns. In a room, one persona attacks another persona's reasoning while you listen, and you can cut in the moment either of them goes wrong.
Could we build this ourselves with CrewAI or LangGraph?
You could build multi-agent text debate in an afternoon, and the hard part of Tingvar is none of that. The hard part is deciding which of three simultaneously willing speakers gets the floor, fast enough that the room still sounds like a conversation, without any of them talking over the human. Agent level turn-taking, where each agent decides for itself whether to speak, stops working at three agents. Our arbiter is a room level policy, and it is where the latency budget is spent.
How we compare
Every claim about another product on this site is a fact that product publishes about itself, read from its own pages on the date below. Where a product publishes nothing on a point, the cell reads not published, because an absent published fact is not an absent capability. Our own weaknesses are in the same table, in the same column style.
Compared on 23 August 2026
Sources
| Claim | Read on their own pages at | Read on |
|---|---|---|
| Median voice to voice latency of 1471 ms from 122 readings over 200 utterances in us-central1 | Our own measurement run, published in full with its method on /evidence | 19 August 2026 |
| Sixteen agents produced sixteen supporting positions and no opposing ones across two published debates | The Crucible public session API at crucible.roundtablelabs.ai | 23 August 2026 |
| A dissent quota shipped to force a minimum level of disagreement, and advice to add an agent that challenges assumptions | The archived Crucible advanced usage guide | 23 August 2026 |
| Agent level turn-taking stops working at three agents | Our own decision layer work, described on /how-it-works and in the glossary entry for turn-taking | 24 August 2026 |
Right of reply
If you work at one of these companies and something here is wrong, email corrections@tingvar.com. We correct it within 1 business day and log the correction publicly.
Correction log
No correction has been requested or issued. This log is dated and it stays here whether or not it has entries.
The mechanism, in one page
Three invariants, in plain English, with no vendor names and no architecture diagram. Read how it works