The problem in plain language
Every U.S. public company files an annual report called a 10-K with the SEC. Inside that report are two kinds of numbers:
- Tagged numbers — wrapped in machine-readable iXBRL codes so software can identify them without human reading. Example: the computer knows exactly which number is "Total Revenue" because the tag says so.
- Untagged numbers — buried in paragraphs of management commentary, footnotes, and explanations. The MD&A (Management's Discussion and Analysis) section is where executives explain results in their own words. These numbers have no machine-readable labels.
Tagged coverage is incomplete — it captures the headline statement figures but misses segment breakdowns, non-standard metrics, and forward-looking commentary. Most systems pick one lane: either deterministic tag extraction (accurate but incomplete) or AI reading of prose (comprehensive but unverified). I wanted both, with every number traced back to its exact source.
What "tracing back to source" means
When I say "provenance," I mean a four-part audit trail for every number — stored in a DAG (Directed Acyclic Graph), which is a type of database that links facts together without circular references:
- Which document — the exact SEC filing (Apple 10-K for fiscal year ending September 28, 2024)
- Which section — Item 8 (financial statements) vs Item 7 (MD&A prose)
- How it was found — machine tag, cross-checked tag, text quote, or AI extraction
- Confidence level — did the machine tag and the AI prose extraction agree?
A concrete example: Apple total revenue
Here is how the same number lives in two places inside the same filing:
This is not cherry-picking. It is the first and only headline revenue figure in the MD&A Results of Operations section. The AI found it without being told where to look.
Results
| Metric | Phase 2 (raw) | Phase 3 (aliases + computed) |
|---|---|---|
| AAPL FY23 coverage | 74.3% | 88.6% |
| AAPL FY24 coverage | 80.0% | 88.6% |
| MSFT FY24 coverage | 74.3% | 88.6% |
| GOOGL FY23 coverage | 80.0% | 91.4% |
| MMM FY23 coverage | 68.6% | 88.6% |
| Kill threshold (60%) | All 5 filings cleared | |
Phase 4 LLM extraction — ground truth validated: The model extracted "Total net sales: $391.0 billion" from Apple FY24 MD&A prose. The iXBRL annual total (context c-1) for the same period is $391.035 billion. Error: 0.0%. This is not cherry-picking — it is the first and only headline revenue figure in the MD&A Results of Operations section.
Pass/fail gates: what we feared, what we tested, what we found
Each gate was written as a kill criterion — a test that, if failed, means the approach is fundamentally broken and should stop. These are not after-the-fact rationalizations; they were defined before any code was written.
Gate 1: Can we extract enough concepts from machine-readable tags alone?
The fear: If raw tag extraction misses more than 40% of the concepts we need, the deterministic foundation is too weak to build on. We would need to rely entirely on AI for everything.
The test: Measure coverage of 35 canonical financial concepts across 5 company filings without any label normalization. Pass threshold: >60%.
What we found: Raw coverage ranged from 68.6% to 80.0%. After adding alias mappings and computed totals, coverage rose to 88.6–91.4%. Result: PASS. The deterministic foundation is solid enough to build on.
Gate 2: Can the data model handle a restatement cleanly?
The fear: If Apple restates FY2023 revenue in a future filing, most systems overwrite the old number. We would lose the original value and the reason it changed.
The test: The schema must support two versions of the same concept (as originally reported vs as recast), with flags for why it changed — amendment, accounting standard adoption (e.g., ASC 606, the accounting standard for revenue recognition), segment reorganization, or simple error correction.
What we found: The RestatementLink object supports all four cases. Result: PASS. A restatement creates a new linked node; the old value remains queryable.
Gate 3: Does AI extraction from prose hallucinate headline numbers?
The fear: A small local AI might invent numbers, confuse fiscal years, or extract a quarterly figure instead of the annual total. If error exceeds 50%, prose extraction is unreliable.
The test: Extract revenue and segment figures from Apple FY24 MD&A. Compare against the iXBRL annual totals (context-filtered, not first-match). Pass threshold: <5% error.
What we found: Total revenue matched at 0.0% error ($391.0B vs $391.035B). All five geographic segments matched. Result: PASS.
Note: "Context-filtered" means filtering the iXBRL tags to only the annual (not quarterly or segment-specific) values. An unfiltered first-match would return $294.9B — a segment total, not the company-wide figure. This is exactly why careful prose extraction matters: the AI correctly identified the headline annual number.
Gate 4: Is a competitor already doing this cheaper and better?
The fear: If Calcbench, idaciti, or another incumbent already solves this problem at scale, a solo build is futile.
The test: Honestly assess competitive position. This is a portfolio artifact, not a product pitch.
What we found: Calcbench ($50M+ funding, 10+ years) and idaciti have massive head starts. This project is a proof of concept for a hiring committee, not a commercial play. Result: NOT APPLICABLE. The artifact still demonstrates systems thinking, data architecture, and validation discipline.
Limits
- Only 5 companies tested. Not a market-wide benchmark.
- The remaining ~10% coverage gap is conditional absence (M&A activity, zero foreign-exchange effects, segment reorganizations) — not extraction failure.
- AI validation is single-company (Apple FY2024). Multi-company validation pending.
- Segment discussions extracted drivers and risks but not quantitative segment-level profit margins.
- No user interface. All output is JSON, typed Python objects, and test assertions.
- Not investment advice. Not a production investor-relations product.
Related notes
Executive Track Record Engine. Knowledge graph + vector search over SEC proxy filings.
Earnings Call Interrogator. Fine-tuned Qwen3-8B vs frontier APIs on earnings-call extraction.
Adversarial Thesis Testing. Bull/bear split vs equal-budget memo. 450 runs.