4. Coverage improvement: why tags miss concepts
Different companies use different iXBRL tag names for the same financial concept. Apple calls revenue RevenueFromContractWithCustomerExcludingAssessedTax. Microsoft calls it RevenueFromContractWithCustomerIncludingAssessedTax. To a computer, these look like completely different things. A human analyst knows they both mean "Total Revenue."
I built a lookup table of 35 canonical financial concepts with their variant tag names across filers, plus formulas for derived totals:
- Gross Profit = Revenue minus Cost of Goods Sold (COGS) — when the GrossProfit tag is absent, compute it from Revenue and COGS tags
- Total Assets = Current Assets plus Non-Current Assets
- Total Liabilities = Current Liabilities plus Non-Current Liabilities
The ~10 percentage point lift comes from recognizing that "ExcludingAssessedTax" and "IncludingAssessedTax" are the same concept to an analyst, and from deriving totals when the total tag is missing but the components are present. The residual ~10% gap is conditional absence — concepts like "Payments to acquire businesses" only appear when a company makes an acquisition that year.
5. AI extraction validation
I ran a local 8-billion-parameter model (Qwen3, via Ollama on my own machine) over Apple FY2024 MD&A prose. The model extracted six headline numbers. All six matched the iXBRL annual totals.
Total revenue
$391.0B
AI vs tag: 0.0% error
Americas region
$167.0B
matches tag c-149
Europe region
$101.3B
matches tag c-152
Greater China
$67.0B
matches tag c-155
Design rule The AI must return structured JSON via an API call with format: "json". The command-line tool ollama run emits "thinking tokens" inline, which breaks JSON parsers. The HTTP API eliminates this entirely.
Corpus statistics
Filings parsed
5
AAPL FY23/FY24, MSFT FY24, GOOGL FY23, MMM FY23
Canonical concepts
35
revenue, gross profit, operating income, total assets, etc.
Coverage after aliases
88.6–91.4%
up from 68.6–80.0% raw tag extraction
Code files
16
~3,200 lines Python; 31 test assertions
6. Pass/fail gates (kill criteria)
Each gate was written as a kill criterion — a test that, if failed, means the approach is broken and should stop. These were defined before any code was written.
Gate 1: Can we extract enough concepts from machine-readable tags alone?
The fear: If raw tag extraction misses more than 40% of concepts, the deterministic foundation is too weak. We would need AI for everything.
The test: Measure coverage of 35 canonical concepts across 5 filings without label normalization. Threshold: >60%.
What we found: Raw coverage 68.6–80.0%. After aliases + computed totals: 88.6–91.4%. Result: PASS. The deterministic foundation is solid enough to build on.
Gate 2: Can the data model handle a restatement?
The fear: If Apple restates FY2023 revenue, most systems overwrite the old number. We lose the original and the reason it changed.
The test: The schema must support two versions — as originally reported vs as recast — with flags for why: amendment, accounting standard adoption (e.g., ASC 606, the accounting standard for revenue recognition), segment reorganization, or error correction.
What we found: The RestatementLink object supports all four cases. A restatement creates a new linked node; the old value remains queryable. Result: PASS.
Gate 3: Does AI extraction from prose hallucinate headline numbers?
The fear: A small local AI might invent numbers, confuse fiscal years, or extract a quarterly figure instead of the annual total. If error exceeds 50%, prose extraction is unreliable.
The test: Extract revenue and segment figures from Apple FY24 MD&A. Compare against iXBRL annual totals (context-filtered, not first-match). Threshold: <5% error.
What we found: Total revenue matched at 0.0% error. All five geographic segments matched. Result: PASS.
"Context-filtered" means filtering the iXBRL tags to only the annual (not quarterly or segment-specific) values. An unfiltered first-match would return $294.9B — a segment total, not the company-wide figure. This is exactly why careful prose extraction matters: the AI correctly identified the headline annual number.
Gate 4: Is a competitor already doing this cheaper and better?
The fear: If Calcbench, idaciti, or another incumbent already solves this at scale, a solo build is futile.
The test: Honestly assess competitive position. This is a portfolio artifact, not a product pitch.
What we found: Calcbench ($50M+ funding, 10+ years) and idaciti have massive head starts. This project demonstrates systems thinking, data architecture, and validation discipline for a hiring committee. Result: NOT APPLICABLE.
7. Limits
- Only 5 companies tested. Not a market-wide benchmark.
- The remaining ~10% coverage gap is conditional absence (mergers and acquisitions [M&A] activity, zero foreign-exchange effects, segment reorganizations) — not extraction failure.
- AI validation is single-company (Apple FY2024). Multi-company validation pending.
- Segment discussions extracted drivers and risks but not quantitative segment-level profit margins.
- No user interface. All output is JSON, typed Python objects, and test assertions.
- Not investment advice. Not a production investor-relations product.
Full source: ~/Projects/sec-provenance/ · 16 files, ~3,200 lines · 31 test assertions