Adversarial Thesis Testing
Bull/bear structure vs equal-budget both-sides research. 450 cells: intensity moves verdicts more than topology. Displacement only.
Open study →Working write-ups from applied AI projects: training small models for specific tasks, designing tests that measure what actually matters, combining structured databases with AI search, and measuring whether expensive cloud models are worth it. I test in finance because that's my day job.
When you ask an AI for a thesis verdict, does a bull/bear split beat one balanced research memo with the same length budget? 450 runs: research helps; the split does not beat the memo; making the bear meaner moves the number a lot. Displacement only. Not accuracy.
Same thesis text, same frozen documents, same judge instructions. Three research setups: no memo, one both-sides memo (~5k tokens), or separate bull and bear writers (~2.5k each). Writers are a local 8B model; the judge is a different model family.
1. Naive: judge sees the thesis only
2. Both-sides: one researcher must cover for and against
3. Split: independent bull and bear writers, then the same judge
4. Lock length budget between both-sides and split
5. Also turn the bear’s tone up and down (intensity)
| Comparison | Mean shift | n |
|---|---|---|
| Both-sides − Naive | +0.087 | 149 |
| Split − Naive | +0.075 | 145 |
| Split − Both-sides | −0.012 | 145 |
Without the both-sides control, any “debate” effect could just be longer context.
Each note answers a concrete question with a fixed setup and stated limits.
Bull/bear structure vs equal-budget both-sides research. 450 cells: intensity moves verdicts more than topology. Displacement only.
Open study →Extracting executive careers from real SEC filings. A graph database keeps facts tied to real people; a locally-run AI model writes memos with source citations.
Open study →Rule-based bias detection from trade records. A task-trained AI model narrates the findings.
Open study →8-billion-parameter earnings-call extractor vs GPT-5.5 and Claude Fable 5. Same pipeline, measured on cost, speed, reliability, and how many evasions each model surfaces.
Open study →I publish these notes to keep a public record of applied AI experiments. I am learning where task-specific training, simulated data, combining graphs with search, and multi-step workflows help, and where a simpler pipeline is enough. The goal is practical: measure quality, cost, and failure modes before treating any of this as production infrastructure.
An open workspace for independent experiments. Each project names a specific question up front, freezes the setup, and defines a pass/fail test before the results come in. If the work misses that bar, the note still ships and says so. That is the point of writing the test first.
Occasional email when a new study ships. No cadence pressure, no list spam.
Happy to share more detail on methods, data, or figures.