Research note · August 2026

SEC Filing Provenance Engine

Public companies file annual reports (called 10-Ks) with the U.S. Securities and Exchange Commission (SEC). Since 2009, those filings include machine-readable tags called iXBRL (Inline eXtensible Business Reporting Language). Think of iXBRL as barcodes embedded in every figure — the computer knows "this is Total Revenue" without a human reading the page. But these tags only cover about 80% of what an analyst needs. The rest — management commentary, segment breakdowns, risks — lives in prose. I built a system that reads both the tags and the prose, traces every extracted number back to its exact source, and validates AI-extracted figures against the machine-readable originals.

Tag coverage (5 filings)88.6–91.4% · 35 concepts
Revenue extraction error0.0% · AI vs tags
Pipeline errors0 · 2 chunks tested
AI model cost$0 · local inference

The problem in plain language

Every U.S. public company files an annual report called a 10-K with the SEC. Inside that report are two kinds of numbers:

  • Tagged numbers — wrapped in machine-readable iXBRL codes so software can identify them without human reading. Example: the computer knows exactly which number is "Total Revenue" because the tag says so.
  • Untagged numbers — buried in paragraphs of management commentary, footnotes, and explanations. The MD&A (Management's Discussion and Analysis) section is where executives explain results in their own words. These numbers have no machine-readable labels.

Tagged coverage is incomplete — it captures the headline statement figures but misses segment breakdowns, non-standard metrics, and forward-looking commentary. Most systems pick one lane: either deterministic tag extraction (accurate but incomplete) or AI reading of prose (comprehensive but unverified). I wanted both, with every number traced back to its exact source.

What "tracing back to source" means

When I say "provenance," I mean a four-part audit trail for every number — stored in a DAG (Directed Acyclic Graph), which is a type of database that links facts together without circular references:

  1. Which document — the exact SEC filing (Apple 10-K for fiscal year ending September 28, 2024)
  2. Which section — Item 8 (financial statements) vs Item 7 (MD&A prose)
  3. How it was found — machine tag, cross-checked tag, text quote, or AI extraction
  4. Confidence level — did the machine tag and the AI prose extraction agree?

A concrete example: Apple total revenue

Here is how the same number lives in two places inside the same filing:

Where the number comes from Two sources, same answer
Source 1 — Machine-readable tag (iXBRL) in Item 8: <ix:nonFraction name="us-gaap:RevenueFromContractWithCustomerExcludingAssessedTax" contextRef="c-1" unitRef="usd" decimals="-6">391035000000</ix:nonFraction> → Tagged as "Total Revenue" · Annual context · Unit = US dollars · Value = $391,035,000,000 Source 2 — Management prose (MD&A) in Item 7: "Total net sales were $391.0 billion in 2024..." What the AI extracted from the prose: {"metric": "Total net sales", "value": "391035", "unit": "MILLION", "scale": "billion"} Comparison: Tag says $391,035 million. AI says $391.0 billion. Same number. Error: 0.0%.

This is not cherry-picking. It is the first and only headline revenue figure in the MD&A Results of Operations section. The AI found it without being told where to look.

Results

MetricPhase 2 (raw)Phase 3 (aliases + computed)
AAPL FY23 coverage74.3%88.6%
AAPL FY24 coverage80.0%88.6%
MSFT FY24 coverage74.3%88.6%
GOOGL FY23 coverage80.0%91.4%
MMM FY23 coverage68.6%88.6%
Kill threshold (60%)All 5 filings cleared
Coverage improvement 5 companies, 2 fiscal years
iXBRL coverage improvement chart
Raw iXBRL extraction (Phase 2) vs after concept aliases and computed totals (Phase 3). The ~10 pp lift comes from resolving label variants and deriving totals from components.

Phase 4 LLM extraction — ground truth validated: The model extracted "Total net sales: $391.0 billion" from Apple FY24 MD&A prose. The iXBRL annual total (context c-1) for the same period is $391.035 billion. Error: 0.0%. This is not cherry-picking — it is the first and only headline revenue figure in the MD&A Results of Operations section.

Ground truth validation Apple FY2024 · 6 line items
Ground truth validation chart
iXBRL annual totals (deterministic, context-filtered) vs LLM-extracted values from MD&A prose. All six revenue line items match to within rounding precision.
Validation

Pass/fail gates: what we feared, what we tested, what we found

Each gate was written as a kill criterion — a test that, if failed, means the approach is fundamentally broken and should stop. These are not after-the-fact rationalizations; they were defined before any code was written.

Gate 1: Can we extract enough concepts from machine-readable tags alone?

The fear: If raw tag extraction misses more than 40% of the concepts we need, the deterministic foundation is too weak to build on. We would need to rely entirely on AI for everything.

The test: Measure coverage of 35 canonical financial concepts across 5 company filings without any label normalization. Pass threshold: >60%.

What we found: Raw coverage ranged from 68.6% to 80.0%. After adding alias mappings and computed totals, coverage rose to 88.6–91.4%. Result: PASS. The deterministic foundation is solid enough to build on.

Gate 2: Can the data model handle a restatement cleanly?

The fear: If Apple restates FY2023 revenue in a future filing, most systems overwrite the old number. We would lose the original value and the reason it changed.

The test: The schema must support two versions of the same concept (as originally reported vs as recast), with flags for why it changed — amendment, accounting standard adoption (e.g., ASC 606, the accounting standard for revenue recognition), segment reorganization, or simple error correction.

What we found: The RestatementLink object supports all four cases. Result: PASS. A restatement creates a new linked node; the old value remains queryable.

Gate 3: Does AI extraction from prose hallucinate headline numbers?

The fear: A small local AI might invent numbers, confuse fiscal years, or extract a quarterly figure instead of the annual total. If error exceeds 50%, prose extraction is unreliable.

The test: Extract revenue and segment figures from Apple FY24 MD&A. Compare against the iXBRL annual totals (context-filtered, not first-match). Pass threshold: <5% error.

What we found: Total revenue matched at 0.0% error ($391.0B vs $391.035B). All five geographic segments matched. Result: PASS.

Note: "Context-filtered" means filtering the iXBRL tags to only the annual (not quarterly or segment-specific) values. An unfiltered first-match would return $294.9B — a segment total, not the company-wide figure. This is exactly why careful prose extraction matters: the AI correctly identified the headline annual number.

Gate 4: Is a competitor already doing this cheaper and better?

The fear: If Calcbench, idaciti, or another incumbent already solves this problem at scale, a solo build is futile.

The test: Honestly assess competitive position. This is a portfolio artifact, not a product pitch.

What we found: Calcbench ($50M+ funding, 10+ years) and idaciti have massive head starts. This project is a proof of concept for a hiring committee, not a commercial play. Result: NOT APPLICABLE. The artifact still demonstrates systems thinking, data architecture, and validation discipline.

Limits

  • Only 5 companies tested. Not a market-wide benchmark.
  • The remaining ~10% coverage gap is conditional absence (M&A activity, zero foreign-exchange effects, segment reorganizations) — not extraction failure.
  • AI validation is single-company (Apple FY2024). Multi-company validation pending.
  • Segment discussions extracted drivers and risks but not quantitative segment-level profit margins.
  • No user interface. All output is JSON, typed Python objects, and test assertions.
  • Not investment advice. Not a production investor-relations product.

Related notes

Executive Track Record Engine. Knowledge graph + vector search over SEC proxy filings.
Earnings Call Interrogator. Fine-tuned Qwen3-8B vs frontier APIs on earnings-call extraction.
Adversarial Thesis Testing. Bull/bear split vs equal-budget memo. 450 runs.