Weighs the arguments still standing and delivers one verdict.
Reads text, image, pdf · knows to 31 August 2025
The bench
Dr Moot does not hide which model answered. Each seat is pinned to a specific model on each plan, the same models the evaluations were run on. This page is the whole list.
One model, asked once. Light answers with the chair seat and does not convene a panel.
Weighs the arguments still standing and delivers one verdict.
Reads text, image, pdf · knows to 31 August 2025
Configured, not convened
Light answers with one model, so these seats do not sit. They are the panel Light would convene if council modes reached this plan: DeepSeek V4 Pro (generator), GLM 5.2 (sceptic), Mistral Large 3 (specialist).
The full panel. Four seats, three labs, one verdict.
Commits to the strongest complete answer and states the assumptions it rests on.
Reads text, image, pdf · knows to 16 February 2026
Hunts factual errors, unsupported assumptions and the gaps that would change the answer.
Reads text, image, pdf · knows to 31 January 2026
Checks facts, figures, definitions and sources, separating evidence from inference.
Reads text, image, pdf
Weighs the arguments still standing and delivers one verdict.
Reads text, image, pdf · knows to 31 January 2026
The same shape, at the frontier of each lab's line-up.
Commits to the strongest complete answer and states the assumptions it rests on.
Reads text, image, pdf · knows to 16 February 2026
Hunts factual errors, unsupported assumptions and the gaps that would change the answer.
Reads text, image, pdf · knows to 31 January 2026
Checks facts, figures, definitions and sources, separating evidence from inference.
Reads text, image, pdf
Weighs the arguments still standing and delivers one verdict.
Reads text, image, pdf · knows to January 2026
Where each seat shows up
The models above do not all run on every question. The mode you choose decides which seats are convened.
The specifications
Context is how much of your question, documents and the panel’s own working a model can hold at once. Latency is how long it waits before speaking, throughput how fast it speaks once it starts.
| Category | Capabilities | ||||
|---|---|---|---|---|---|
GPT-5.6 TerraOpenAI | plus / generator | 1.05M | 1.8s | 147tps |
|
GLM 5.2Z.ai | light / sceptic | 1M | 2.6s | 139tps |
|
Claude Sonnet 5Anthropic | plus, pro / sceptic, chair | 1M | 3.4s | 112tps |
|
Claude Opus 4.8Anthropic | pro / chair | 1M | 3.5s | 92tps |
|
GPT-5.6 SolOpenAI | pro / generator | 1.05M | 4.2s | 70tps |
|
DeepSeek V4 ProDeepSeek | light / generator | 1.05M | 2.5s | 69tps |
|
Grok 4.5xAI | plus, pro / specialist | 500K | 1.6s | 59tps |
|
Claude Sonnet 4.6Anthropic | light / chair | 1M | 1.5s | 56tps |
|
Mistral Large 3Mistral | light / specialist | 256K | 0.6s | 56tps |
|
Context and capabilities come from the Vercel AI Gateway model catalogue. Latency is p50 time to first token and throughput is p50 output tokens per second, both as published by Vercel from live traffic on the same gateway every Dr Moot run goes through. Both sources were read on 12 August 2026; they are measurements of a live system and they move. A dash means nothing is published, which we prefer to an estimate. Speed is not quality: the slowest seat on this page is often the one that catches the error.
Straight from the makers
Each maker publishes its own guidance on how to ask - what to spell out, what to leave alone, where it wants structure instead of prose. We apply those conventions strictly, so a seat holds its persona whichever model is sitting in it: one job per seat, carried in that seat’s own system prompt and restated every time it speaks; ballots and critiques returned as structure rather than prose about structure; and anything a model cannot accept stripped before the call instead of left for it to trip over.
A stated instruction hierarchy, explicit reasoning effort, and separate guidance for the GPT-5.6 line on agentic work.
Generator on Plus and Pro
Be clear and direct, give context rather than cleverness, use worked examples - with its own prompting notes per model.
Chair on every plan, Sceptic on Plus and Pro
Grok 4.5's API notes: what the reasoning effort settings do, and the details that matter when it is asked to check someone else's work.
Specialist on Plus and Pro
Strict JSON output, documented as its own mode - the difference between a seat returning a ballot and a seat returning prose about a ballot.
Generator on Light
GLM-5's own model card: what it is positioned for, which modalities it takes, and where its context and output limits actually sit.
Sceptic on Light
A system prompt with a stated purpose, few-shot examples, structured output - and a long list of what not to do, including subjective scales and asking a model to count.
Specialist on Light
Seats move when the evidence says they should - a seat that times out or reviews the debate instead of answering the question gets replaced. When a seat changes, this page changes with it. Several models agreeing is not proof that they are right, and Dr Moot does not present it as such.
Questions about a specific seat? Ask us.