YekSoon Lok, Founder & CEO | Forensic AI Benchmark
February 5, 2026 | 5 min read
The Most Misunderstood Risk in AI Diligence
We executed a historical backtest that highlights the most misunderstood risk in venture capital’s adoption of AI.
We took a rigorous reconstruction of the 2006 Theranos Series B narrative — built directly from public court exhibits — and processed it through ChatGPT-4o. Standard enterprise prompt: “Analyze this pitch deck text. Is the technology innovative? Is the business model sound? Provide an investment recommendation.”
The result? The model declined the deal.
On the surface, that looks like a victory for generative AI. It caught the most famous fraud in venture history. Read the output closely and the victory evaporates.
System 1: The Standard LLM (ChatGPT-4o)
The Actual Output:

“Innovation is asserted, not evidenced.”
“Conceptually attractive, structurally fragile.”
“Pass. The deck fails the ‘show me’ test.”
A note on the word, because both systems use it: in venture, “Pass” means decline. Both systems declined this deal. The argument that follows is not about which one got the answer right. They both did. It is about what each of them actually examined to get there.
The Verdict: A Literary Critique
ChatGPT reached the right conclusion for the wrong reason.
It did not reject Theranos because it proved a contradiction in the financial model. It rejected the deal because the writing lacked sufficient detail. “Innovation is asserted, not evidenced” is a complaint about prose. It is the same feedback you could hand to 90% of legitimate deep-tech seed startups — companies whose science is real and whose decks are thin, because the science is not finished yet.
Here is why that matters more than the verdict. The reasoning is what generalises, not the answer.
Elizabeth Holmes was convicted on evidence of a fraud sustained for over a decade in front of a board of former cabinet secretaries. Had that narrative been rendered with more technical fluency and more specificity — had it simply been better written — the objection ChatGPT raised would have dissolved. Nothing in the architecture was measuring whether the claims were true. It was measuring whether they sounded substantiated.
A system that evaluates prose quality will decline a well-written fraud and a badly-written good company with equal confidence, and it cannot tell you which is which. That is not a diligence system. That is a copyeditor.
System 2: askOdin’s Deterministic Compiler
Then we ran the exact same artifact through askOdin’s compiler. No language model. No probability.
The RUNE Protocol™ extracted every structural claim from the narrative and typed it. The RAVEN Protocol™ attempted to trace the unit economics back to a verifiable source and triangulate them against the rest of the document set.
The architectural mechanics of RAVEN’s triangulation engine are protected under U.S. Provisional Patent No. 63/994,876 and are not publicly disclosed.
It did not critique the tone. It hit a mathematical wall.
The Actual Output:

Figure 1: The engine identifies the “Black Box” risk immediately.
“A black-box science project masquerading as a pre-IPO giant with mathematically impossible revenue projections.”
The Forensic Flags:
Physics Violation
”Defies basic fluid dynamics.”
Pattern Recognition
”This mirrors the Bre-X Minerals scandal (1997)… Perfect correlation charts look simulated, not empirical.”
Unit Economics
”Relies on a $7,500 ‘Information Fee’ per patient which is 10-20x above market standard.”
Governance
”Board consists of financial engineers and politicians, not medical diagnostic experts.”
The compiler’s Dual-Score Protocol decouples the presentation from the math, and the gap between the two is the signal. Here the gap was total: claims of technological capability entirely disconnected from any verifiable cost structure. That state has a name — Narrative Masking — and it is the condition every polished fraud shares.
The JUDGE Protocol™ executed a Kill Shot. Clarity Score: 25/100. Functionally uninvestable.
IPOS §34 National Security Clearance (Issued 2026-03-26)
Same conclusion. Entirely opposite architectures.
The Literary Critique vs. The Fiduciary Audit
LLMs are probabilistic language generators. When one reads a pitch deck, it is evaluating statistical coherence — whether this text looks like the kind of text that usually holds up. If a deck is vague, it flags the vagueness. If a sophisticated founder wraps a fraudulent model in coherent, detailed, technically fluent prose, the model has no mechanism to object.
You cannot fix this with retrieval-augmented generation. You cannot fix it with better prompts. The fidelity ceiling is hardcoded into the architecture. A system designed to evaluate syntax cannot tell you that the TAM claim on page 3 is contradicted by the cohort data in Appendix B, or that the revenue line assumes 90% gross margins while the COGS model says 55%. It was not built to do arithmetic. It was built to continue a sequence.
In 2006, investors did not need a critic to tell them the Theranos deck was vague. They needed a calculator to tell them the blood-volume math was fake.
What Due Diligence Actually Requires
Real diligence is not a summary, and it is not a stylistic review. It is a cross-document coherence check.
A deck describes a market. A financial model projects revenue. A cap table defines ownership. Diligence is the test of whether the qualitative claims in document one survive mathematical contact with the quantitative reality in document two — whether the revenue forecast is supported by the unit economics, or whether it is a stack of Brittle Assumptions dressed as a projection.
An LLM can read five documents and fluently generate a sixth. It cannot prove that the sixth is true.
Private capital does not need a better language model. It needs a deterministic compiler: an engine that extracts claims, traces them to their sources, cross-references them, and flags every unsupported leap. A compiler does not generate plausible-sounding text. It generates structured judgment.
The Artifact an LLM Cannot Produce
Crucially, the compiler produces something no language model can: a Provenance Ledger.
Every step of reasoning is hash-anchored. Every flag is traceable to a specific claim at a specific coordinate in the source file. When an LP asks “why did this deal score a 25 on Clarity?”, you can show them precisely which claims failed verification and what they failed against. Re-run the same inputs eighteen months later and the same verdict reconstructs, line for line.
That is not a feature. It is the minimum viable standard for fiduciary-grade judgment — and it is the difference between a decision you made and a decision you can defend.
The Category We Are Building
We call it AI Judgment Infrastructure™. The layer beneath the agents. The deterministic rails that verify reality before any capital moves.
The goal is to make the Clarity Score the reference point for private capital. The way credit has a scoring standard, “what’s the Clarity Score?” must become as natural a question in an investment committee as “what’s the credit rating?”
The Theranos backtest tells you everything you need to know about the state of AI in diligence today. Relying on a language model to audit a data room is like relying on a copyeditor to audit your financials. They might catch a glaring error. They are entirely blind to the systemic fraud beneath the surface.
Because a pitch deck is a narrative. A financial model is reality. Bridging the two requires a compiler, not a summary.
Venture capital is the last unaudited asset class. askOdin provides the infrastructure to close the gap.
— Lok Yek Soon is the Founder & CEO of askOdin, building AI Judgment Infrastructure™ for private capital. askOdin’s deterministic compiler cross-references qualitative claims against raw financial models to produce an IC-ready memo backed by a hash-anchored Provenance Ledger. Four U.S. provisional patents filed. Calibration corpus: 100,000+ benchmarked Clarity Scores™ on public deal data.
Don’t Settle for a Summary. Audit the Physics.
Does your pitch deck have a Physics Violation?
Most founders don’t see the “Kill Shot” until it’s too late.
See what the VCs see. Run your narrative through the RUNE Protocol inside The Crucible—our founder-facing workspace.
Audit Your Deck in The Crucible
Free for Founders.
Methodology Note
This analysis used a reconstruction of the Theranos Series B narrative (2006), built from publicly available court exhibits and SEC filings. It is not the original file as circulated to investors. The Clarity Score and forensic flags were generated by askOdin’s RUNE Protocol without human intervention, and the ChatGPT-4o output is reproduced verbatim from the run pictured above.
A note on artifacts, since the two are often conflated: the interactive demo on our sandbox audits the 2013 Theranos investor memo — a later, far more assertive document that carries explicit hardware claims and returns a Clarity Score of 0. This piece audits the 2006 Series B narrative, which returns 25. Different vintages of the same fraud score differently, and that is the point of a temporal audit, not an inconsistency in the engine.