---
url: https://tingvar.com/why-disagreement
title: "Why disagreement · Tingvar"
updated: 2026-09-12
---

24 AUG 2026

# Why disagreement

Disagreement between agents is more informative than criticism from an agent. The load-bearing assumption is the one two personas price differently.

You have had this conversation. You describe a plan to an assistant. It tells you the instinct is right, picks out the sharpest thing you said and agrees with it, then gets to work on the problem exactly as you framed it. You come away more confident. Nothing about the plan changed.

That is not a fault in your reading of it. It is what the model was tuned to do. Post-training optimises assistants for helpfulness and for user satisfaction, and the reliable way to satisfy someone is to agree with them. Agreement from something that sounds competent reads as evidence. It is not evidence.

## Three things go wrong when you tell it to be critical

1. It stops being critical Be critical holds for two or three turns and then relaxes back toward agreement.
2. Nothing argues with the critic A lone critic has no counterparty, so its weakest objection and its strongest arrive in the same tone.
3. It has one yardstick One critic reasons from one implicit idea of what matters, and prices every assumption against it.

01

### It stops being critical

The instruction is in the context and the training is in the weights. Ask an assistant to attack a plan and it will, for a turn or two, and then the register softens: the objections get shorter, the qualifications get longer, and by the fifth exchange it is helping you build the thing it was supposed to be dismantling.

Nothing in that conversation tells you it has happened. The tone at turn eight is the same tone as at turn two, so your one signal is that the criticism stopped arriving, which reads as the plan having survived it.

02

### Nothing argues with the critic

A lone critic has no counterparty. Its weakest objection and its strongest objection arrive in the same tone, with the same confidence, and nothing in the exchange separates them. So you have to evaluate the critique yourself, which is the work you were trying to get help with in the first place.

Put a second critic in the room with an opposed interest and the separation happens on its own. An objection that survives being attacked by someone who wanted it to fail has been through something an objection nobody contested has not.

03

### It has one yardstick

One critic reasons from one implicit idea of what matters. Ask it about a plan and you get that plan assessed against a single value function, consistently and often well. What it cannot give you is the thing you actually need, which is the assumption two competent people would price differently.

A finance reading and a distribution reading of the same paragraph disagree about which sentence is the risk. That disagreement is a coordinate. One reading, however hostile, is a score.

Not yet produced

A figure showing how far an instruction travels before a model stops applying it. It is producible today from the written argument and it has not been drawn. The pages carry the argument as prose, so the gap costs a reader the picture and not the point.

Blocked on design production 2026-08-26

## The thesis

Disagreement between agents is more informative than criticism from an agent. The load-bearing assumption is the one two personas price differently. So when a persona optimising for cash survival and a persona optimising for distribution reach opposite conclusions from the same proposal, the sentence the two of them read differently is the one to go back to. That is a location, not an opinion, and a single critic cannot give you it however hostile it is, because a location needs two readings.

Not yet produced

A figure showing two defensible conclusions drawn from one brief. It is producible today and it has not been drawn. The page argues the same point in prose above.

Blocked on design production 2026-08-26

## Why it has to be spoken

Voice is not a feature here. It is how interruption becomes possible. In text you wait for a paragraph to finish and then correct it. In a room you cut in at the word where it went wrong, which is the moment the correction is worth most and the one steering control that operates at the speed of the mistake.

In text, turn-taking is trivial: nothing overlaps and nobody is kept waiting. In speech it is the hard part, and it is where the latency budget is spent. Agent level turn-taking, where each agent decides for itself whether to speak, stops working at three agents, because three personas then answer every human utterance at once and the room is noise. Room level turn-taking can see every persona's intent before it grants the floor to one of them. It is the core of what we build, and it is the reason a room sounds like a conversation.

The arbiter is the room level policy that decides which persona takes the floor next. The arbiter is not a persona. It has no opinion about your plan, it never speaks, and it is not a fifth voice settling the argument on your behalf. It is pure policy over events, with no network access and no audio, which is why it can be run and tested offline. Each persona bids to speak, the bids are scored locally against sentence embeddings rather than by a model call, and the arbiter grants the floor to one of them. Agent level turn-taking, where each agent decides for itself whether to speak, stops working at three agents.

[Read how it works](/how-it-works)

## A competitor published the counter-example

Crucible, by Roundtable Labs, ran a text debate harness whose personas were differentiated by domain expertise and by which model vendor ran them. It published two complete debate records through a public API, which is more primary evidence of its own behaviour than most products in this category publish about theirs.

Across those two debates, sixteen agents produced sixteen supporting positions and no opposing ones. Its advanced usage guide shipped a Dissent Quota, described as ensuring a minimum level of disagreement, and told the customer to include at least one agent that would challenge assumptions. A harness that needs a quota on disagreement is one in which disagreement does not arise on its own.

That is a design finding with a company attached to it, and it is the strongest external evidence that exists for the argument on this page. It says nothing about voice, and we do not claim it does.

| Claim | Read on their own pages at | Read on |
| --- | --- | --- |
| Sixteen agents produced sixteen supporting positions and no opposing ones across two published debates | The Crucible public session API at crucible.roundtablelabs.ai | 23 August 2026 |
| A dissent quota shipped to force a minimum level of disagreement, and advice to add an agent that challenges assumptions | The archived Crucible advanced usage guide | 23 August 2026 |
| Personas defined by name, role, prompt and domains | The same public session API | 23 August 2026 |

Where each Crucible statement above was read, and when

Scroll to see the rest

## What this page does not claim

It does not claim the disagreement is useful, or correct, or that it happens at any measured rate in a real session. The internal target for turns in which one persona challenges another is published on the evidence page with an empty measured column, because zero real human sessions have taken place and every measured turn so far used pre-rendered audio.

It also does not borrow a literature. There are papers about multi agent debate and we cite none of them here, because a paper about someone else's system is not evidence about ours. [See what we have measured](/evidence)

## Questions about the thesis

### Does this just agree with me the way ChatGPT does?

No, Tingvar is built to attack the framing of your plan rather than to be helpful inside it, which is the failure a single assistant has by design. Post-training alignment optimises assistants for helpfulness and user satisfaction, so a model validates your framing and then solves the problem inside it. Three personas with opposed objective functions disagree with each other, and that disagreement is the signal. Our target for agreement-only turns is under 5 % of agent turns. It is a target, not a measurement.

### Why would I want an AI to argue with me?

Tingvar's personas attack the proposal and never the proposer, and the useful signal is their disagreement with each other rather than their criticism of you. If a finance persona and a growth persona reach opposite conclusions from the same plan, the assumption they price differently is the one carrying the plan's weight. Every session ends with one question you answer yourself: name one thing you now believe about this proposal that you did not believe an hour ago. An empty answer is a failed session.

### Why voice, and wouldn't text be easier to read?

Tingvar uses voice because interruption is the primary steering control, and interruption is native only in speech. In text you wait for a paragraph to finish. In a room you cut in the moment a persona goes wrong, and it stops. Spoken debate is also faster than text and forces brevity on the personas. You still get the record: a live transcript labels every line with the name of whoever said it, and it is exportable.
