Project Mantis · Research note
SEC Filing Provenance Engine
A lineage system for U.S. public-company financial data. Machine-readable tags (iXBRL) provide deterministic extraction; a local AI reads management prose (MD&A) and extracts headline numbers validated against those tags.
SEC · LINEAGE
August 2026
Author: Alex Behm
Tag coverage (5 filings)88.6–91.4% · 35 concepts
Revenue extraction error0.0% · AI vs tags
Pipeline errors0 · 2 chunks tested
AI model cost$0 · local inference
SEC filing → section extraction → tag resolution → alias mapping → computed totals → lineage graph Ollama qwen3:8b · JSON mode

Thesis: Deterministic extraction from machine-readable financial tags can be pushed past 80% coverage by mapping variant tag names to a single canonical concept. AI extraction from management prose can match those headline numbers with zero error — if you eliminate model "thinking tokens," ground-truth every number against the tags, and define pass/fail gates before running the experiment.

1. What this problem looks like in practice

Every public company in the United States files an annual report called a 10-K with the Securities and Exchange Commission (SEC). Inside that report are two kinds of numbers:

  • Tagged numbers — wrapped in machine-readable codes called iXBRL (Inline eXtensible Business Reporting Language). Think of iXBRL as barcodes embedded in every figure: the computer knows "this number is Total Revenue" without a human reading the surrounding text.
  • Untagged numbers — buried in paragraphs of management commentary. The MD&A (Management's Discussion and Analysis) section is where executives explain results in prose. These numbers have no machine-readable labels.

Tagged extraction is deterministic and accurate but incomplete — it captures headline statement figures but misses segment breakdowns, non-standard metrics, and forward-looking commentary. Most systems pick one lane. I wanted both, with every number traced back to its exact source.

2. A concrete example: Apple total revenue

Here is how the same number lives in two places inside Apple's fiscal year 2024 10-K filing:

Where the number comes from Two sources, same answer
Source 1 — Machine-readable iXBRL tag in Item 8 (financial statements): <ix:nonFraction name="us-gaap:RevenueFromContractWithCustomerExcludingAssessedTax" contextRef="c-1" unitRef="usd" decimals="-6">391035000000</ix:nonFraction> → Tagged as "Total Revenue" · Annual context · Unit = US dollars · Value = $391,035,000,000 Source 2 — Management prose in Item 7 (MD&A): "Total net sales were $391.0 billion in 2024..." What the AI extracted from the prose: {"metric": "Total net sales", "value": "391035", "unit": "MILLION", "scale": "billion"} Comparison: Tag says $391,035 million. AI says $391.0 billion. Same number. Error: 0.0%.

This is not cherry-picking. It is the first and only headline revenue figure in the MD&A Results of Operations section. The AI found it without being told where to look.

3. Pipeline

The system has two parallel paths that converge on the same typed fact structure:

Path 1 — Deterministic (rules-based, no AI):

SEC 10-K filing → HTML parser → iXBRL resolver → Concept aliases → Computed totals → Typed fact object

Path 2 — AI-assisted (local model, $0 cost):

SEC 10-K filing → Section segmenter → Semantic chunker → AI extraction (Ollama) → Value parser → Typed fact object

Both paths feed into a lineage graph — a Directed Acyclic Graph (DAG), which is a type of database that links facts together without circular references — so you can ask "what was this number before the restatement?" and get a precise answer tied to the original document.

MetricPass thresholdMeasured
iXBRL coverage (raw tags only)≥60%68.6–80.0% (5 filings)
iXBRL coverage (with aliases + computed totals)≥80%88.6–91.4% (5 filings)
Revenue extraction error (AI vs iXBRL annual total)<5%0.0% ($391.0B vs $391.035B)
Pipeline errors (parse failures, JSON breaks)00 (2 chunks tested)

4. Coverage improvement: why tags miss concepts

Different companies use different iXBRL tag names for the same financial concept. Apple calls revenue RevenueFromContractWithCustomerExcludingAssessedTax. Microsoft calls it RevenueFromContractWithCustomerIncludingAssessedTax. To a computer, these look like completely different things. A human analyst knows they both mean "Total Revenue."

I built a lookup table of 35 canonical financial concepts with their variant tag names across filers, plus formulas for derived totals:

  • Gross Profit = Revenue minus Cost of Goods Sold (COGS) — when the GrossProfit tag is absent, compute it from Revenue and COGS tags
  • Total Assets = Current Assets plus Non-Current Assets
  • Total Liabilities = Current Liabilities plus Non-Current Liabilities
iXBRL coverage improvement chart

The ~10 percentage point lift comes from recognizing that "ExcludingAssessedTax" and "IncludingAssessedTax" are the same concept to an analyst, and from deriving totals when the total tag is missing but the components are present. The residual ~10% gap is conditional absence — concepts like "Payments to acquire businesses" only appear when a company makes an acquisition that year.

5. AI extraction validation

I ran a local 8-billion-parameter model (Qwen3, via Ollama on my own machine) over Apple FY2024 MD&A prose. The model extracted six headline numbers. All six matched the iXBRL annual totals.

Ground truth validation chart
Total revenue
$391.0B
AI vs tag: 0.0% error
Americas region
$167.0B
matches tag c-149
Europe region
$101.3B
matches tag c-152
Greater China
$67.0B
matches tag c-155

Design rule The AI must return structured JSON via an API call with format: "json". The command-line tool ollama run emits "thinking tokens" inline, which breaks JSON parsers. The HTTP API eliminates this entirely.

Corpus statistics

Filings parsed
5
AAPL FY23/FY24, MSFT FY24, GOOGL FY23, MMM FY23
Canonical concepts
35
revenue, gross profit, operating income, total assets, etc.
Coverage after aliases
88.6–91.4%
up from 68.6–80.0% raw tag extraction
Code files
16
~3,200 lines Python; 31 test assertions

6. Pass/fail gates (kill criteria)

Each gate was written as a kill criterion — a test that, if failed, means the approach is broken and should stop. These were defined before any code was written.

Gate 1: Can we extract enough concepts from machine-readable tags alone?

The fear: If raw tag extraction misses more than 40% of concepts, the deterministic foundation is too weak. We would need AI for everything.

The test: Measure coverage of 35 canonical concepts across 5 filings without label normalization. Threshold: >60%.

What we found: Raw coverage 68.6–80.0%. After aliases + computed totals: 88.6–91.4%. Result: PASS. The deterministic foundation is solid enough to build on.

Gate 2: Can the data model handle a restatement?

The fear: If Apple restates FY2023 revenue, most systems overwrite the old number. We lose the original and the reason it changed.

The test: The schema must support two versions — as originally reported vs as recast — with flags for why: amendment, accounting standard adoption (e.g., ASC 606, the accounting standard for revenue recognition), segment reorganization, or error correction.

What we found: The RestatementLink object supports all four cases. A restatement creates a new linked node; the old value remains queryable. Result: PASS.

Gate 3: Does AI extraction from prose hallucinate headline numbers?

The fear: A small local AI might invent numbers, confuse fiscal years, or extract a quarterly figure instead of the annual total. If error exceeds 50%, prose extraction is unreliable.

The test: Extract revenue and segment figures from Apple FY24 MD&A. Compare against iXBRL annual totals (context-filtered, not first-match). Threshold: <5% error.

What we found: Total revenue matched at 0.0% error. All five geographic segments matched. Result: PASS.

"Context-filtered" means filtering the iXBRL tags to only the annual (not quarterly or segment-specific) values. An unfiltered first-match would return $294.9B — a segment total, not the company-wide figure. This is exactly why careful prose extraction matters: the AI correctly identified the headline annual number.

Gate 4: Is a competitor already doing this cheaper and better?

The fear: If Calcbench, idaciti, or another incumbent already solves this at scale, a solo build is futile.

The test: Honestly assess competitive position. This is a portfolio artifact, not a product pitch.

What we found: Calcbench ($50M+ funding, 10+ years) and idaciti have massive head starts. This project demonstrates systems thinking, data architecture, and validation discipline for a hiring committee. Result: NOT APPLICABLE.

7. Limits

  • Only 5 companies tested. Not a market-wide benchmark.
  • The remaining ~10% coverage gap is conditional absence (mergers and acquisitions [M&A] activity, zero foreign-exchange effects, segment reorganizations) — not extraction failure.
  • AI validation is single-company (Apple FY2024). Multi-company validation pending.
  • Segment discussions extracted drivers and risks but not quantitative segment-level profit margins.
  • No user interface. All output is JSON, typed Python objects, and test assertions.
  • Not investment advice. Not a production investor-relations product.

Full source: ~/Projects/sec-provenance/ · 16 files, ~3,200 lines · 31 test assertions