Same thesis. Same judge. Three research setups.
Imagine you have an investment claim (“this copper thesis is right”) and a fixed folder of documents the AI is allowed to use. You can ask a judge model for a probability that the claim is correct. The interesting question is what you put in front of that judge.
I ran three setups on every claim:
The judge sees only the thesis text. No extra analysis. This is the “just ask the model” baseline.
One researcher model must argue for and against inside a single memo, capped at about 5,000 tokens. This controls for “the judge just read more text.”
Two writers work independently (about 2,500 tokens each): one pushes the upside, one attacks the thesis. Their outputs are combined and handed to the same judge. Total budget matches the balanced memo.
Everything else stays locked: the thesis wording, the document pack, the judge instructions, and the output form (a probability from 0 to 1, a confidence label, and a named “this would kill the thesis” condition). The research writers use a local 8B model. The judge is a different model family so the writers are not grading their own style.
Research helps. Split debate does not beat a balanced memo. Tone does.
Average judge probability
On the 0–1 probability the judge assigns to “this thesis is correct,” the naive setup sits lowest. Both research setups sit higher and close to each other. The balanced memo is slightly above the split debate.
Paired differences (same thesis, same settings, same repeat)
The cleaner comparison matches each run to its siblings: same thesis, same intensity setting, same repeat number. That produces three shifts:
- Balanced memo vs naive: about +0.09 (research matters).
- Split debate vs naive: about +0.08 (also matters, same ballpark).
- Split debate vs balanced memo: about −0.01 (no win for the fancy structure).
The intensity dial
Inside the split-debate setup I also changed how aggressive the writers sound (neutral analyst → advocate/skeptic → promoter/short seller), and ran “elasticity” cells that hold the bull fixed while only the bear gets harsher.
That single dial moves the judge by about 14 points (mean probability 0.51 with a mild bear, 0.37 with a short-seller bear). Structure was almost flat. Prompt authorship was not.
The judge always names a kill-condition. Only ~two-thirds are settleable.
Every verdict must include a primary falsifier: one concrete condition that would make the thesis wrong (for example, “gross margin stays below 41% for two quarters,” not “risks remain”).
I scored a 50-run sample on three yes/no questions:
- Specific: Is it a real threshold or event, not mush?
- Grounded in the given docs: Could you justify it from the frozen folder, not invented outside facts?
- Publicly checkable in form: Could a human later mark it true/false from prices, filings-style lines, or announced events?
Results after a first pass plus a conservative automated second pass: 100% specific, 94% grounded, 68% checkable, 64% all three.
Common checkability fails: product unit counts, CAC payback, “failed because of X” causal wording, and “WoodMac or equivalent” escape hatches. The split-debate setup did not produce cleaner falsifiers than the naive setup on this sample.
What a careful reader should not over-claim
- No accuracy. Demo theses; no resolved outcomes.
- No claim that debate is useless forever. This version is parallel/blind only. No live rebuttal round, no web tools, no real investable book.
- Three theses. Patterns are consistent across them, but this is not N=40 real names.
- Falsifier audit is small and partly automated. Good enough for a draft methods footnote; not dual-human IRR.
Takeaway: If you want the judge to move off a cold first answer, give it structured research. If you want the “adversarial” label, budget the same tokens for a single both-sides memo before you build a bull/bear orchestra. And if the number moves when you only rewrite the bear’s personality, say that out loud.
August 2026 · Alex Behm · Project Mantis · Not investment advice · Synthetic demo theses only.