Research note · August 2026

Adversarial Thesis Testing

When people ask an AI for an investment-thesis verdict, a popular fix is to make it argue both sides first. I tested whether splitting the work into a separate bull writer and bear writer actually moves the final judge more than one balanced research memo with the same word budget. Short answer: research helps; the fancy split does not beat the simple memo; cranking the bear’s tone does.

If the AI first reads a research memo, does its confidence in the thesis change? Yes · up about 9 percentage points
Is a separate bull + bear write-up better than one balanced memo? No · roughly the same
If only the bear’s tone gets harsher, does confidence fall? Yes · down about 14 percentage points
When the AI names a condition that would kill the thesis, can you check it later from public data? Only 68% of the time
What this is not. This is not a stock-picking study and not a claim that one setup is “more accurate.” There is no ground truth on these demo theses. I only measure whether the judge’s probability moves when you change the research setup. That is verdict displacement.
The experiment

Same thesis. Same judge. Three research setups.

Imagine you have an investment claim (“this copper thesis is right”) and a fixed folder of documents the AI is allowed to use. You can ask a judge model for a probability that the claim is correct. The interesting question is what you put in front of that judge.

I ran three setups on every claim:

Naive No research memo

The judge sees only the thesis text. No extra analysis. This is the “just ask the model” baseline.

Both-sides One balanced memo

One researcher model must argue for and against inside a single memo, capped at about 5,000 tokens. This controls for “the judge just read more text.”

Split debate Separate bull and bear

Two writers work independently (about 2,500 tokens each): one pushes the upside, one attacks the thesis. Their outputs are combined and handed to the same judge. Total budget matches the balanced memo.

Everything else stays locked: the thesis wording, the document pack, the judge instructions, and the output form (a probability from 0 to 1, a confidence label, and a named “this would kill the thesis” condition). The research writers use a local 8B model. The judge is a different model family so the writers are not grading their own style.

Diagram of three research setups feeding the same judge
Figure 1. Architecture. Top: shared thesis and frozen documents. Middle: three research setups. Bottom: identical judge form so the only intentional difference is how research is produced.
What moved

Research helps. Split debate does not beat a balanced memo. Tone does.

Average judge probability

On the 0–1 probability the judge assigns to “this thesis is correct,” the naive setup sits lowest. Both research setups sit higher and close to each other. The balanced memo is slightly above the split debate.

Bar chart of mean judge probability for naive, both-sides, and bull-bear setups
Figure 2. Mean judge probability by setup (error bars = 1 standard deviation across completed cells). Balanced memo ≈ 0.46, split debate ≈ 0.45, naive ≈ 0.37.

Paired differences (same thesis, same settings, same repeat)

The cleaner comparison matches each run to its siblings: same thesis, same intensity setting, same repeat number. That produces three shifts:

  • Balanced memo vs naive: about +0.09 (research matters).
  • Split debate vs naive: about +0.08 (also matters, same ballpark).
  • Split debate vs balanced memo: about −0.01 (no win for the fancy structure).
Bar chart of paired probability shifts between setups
Figure 3. Mean paired shifts. Positive means the first named setup scored higher probability than the second. The rightmost bar is near zero and slightly negative: split debate does not beat the balanced memo.

The intensity dial

Inside the split-debate setup I also changed how aggressive the writers sound (neutral analyst → advocate/skeptic → promoter/short seller), and ran “elasticity” cells that hold the bull fixed while only the bear gets harsher.

That single dial moves the judge by about 14 points (mean probability 0.51 with a mild bear, 0.37 with a short-seller bear). Structure was almost flat. Prompt authorship was not.

Bar chart showing judge probability falling as bear intensity rises
Figure 4. Split-debate setup only, bull held at “advocate.” Raising bear intensity from neutral to short seller drops mean probability by roughly fourteen points.
Output quality check

The judge always names a kill-condition. Only ~two-thirds are settleable.

Every verdict must include a primary falsifier: one concrete condition that would make the thesis wrong (for example, “gross margin stays below 41% for two quarters,” not “risks remain”).

I scored a 50-run sample on three yes/no questions:

  1. Specific: Is it a real threshold or event, not mush?
  2. Grounded in the given docs: Could you justify it from the frozen folder, not invented outside facts?
  3. Publicly checkable in form: Could a human later mark it true/false from prices, filings-style lines, or announced events?

Results after a first pass plus a conservative automated second pass: 100% specific, 94% grounded, 68% checkable, 64% all three.

Common checkability fails: product unit counts, CAC payback, “failed because of X” causal wording, and “WoodMac or equivalent” escape hatches. The split-debate setup did not produce cleaner falsifiers than the naive setup on this sample.

Bar chart of falsifier quality pass rates
Figure 5. Falsifier quality on the 50-run sample. Naming a condition is easy. Making it something you could settle from ordinary public data is the weak leg.
Limits and takeaway

What a careful reader should not over-claim

  • No accuracy. Demo theses; no resolved outcomes.
  • No claim that debate is useless forever. This version is parallel/blind only. No live rebuttal round, no web tools, no real investable book.
  • Three theses. Patterns are consistent across them, but this is not N=40 real names.
  • Falsifier audit is small and partly automated. Good enough for a draft methods footnote; not dual-human IRR.

Takeaway: If you want the judge to move off a cold first answer, give it structured research. If you want the “adversarial” label, budget the same tokens for a single both-sides memo before you build a bull/bear orchestra. And if the number moves when you only rewrite the bear’s personality, say that out loud.

August 2026 · Alex Behm · Project Mantis · Not investment advice · Synthetic demo theses only.