The Model Was Confidently Wrong: Deterministic Tools as Judge
AI deepdive
0:00 / 0:00
The Model Was Confidently Wrong: Deterministic Tools as Judge
5 просмотров · 2 нед. назад
AI deepdive
10 подписчиков
5 просмотров · 2 нед. назад
Ask a language model what a Windows system function looks like and it will tell you. Confidently. In exactly the right terminology. And on seventy-one real Windows system binaries, it was wrong about sixty-nine of them.
The interesting part is not that it was wrong. The interesting part is that nothing inside the model could have told it so. A model predicts text. Its answers are shaped by what was common in what it read. A compiled binary is not text — it is bytes produced by one specific compiler, at one specific optimization level. The model has a prior for what files like this usually look like. It has no prior for this file.
So the architecture question stops being "how do we make the model more accurate" and becomes "where do we put the judgement". This video is about reverify, a tool built around one rule: the model may propose, but it may not assert. Every claim is expressed in a constrained grammar and checked against the artifact by deterministic code, which returns one of three verdicts — VERIFIED, REFUTED, or INCONCLUSIVE — each carrying the evidence behind it.
The numbers are the reason this is worth fifteen minutes:
On 71 Windows system binaries, a fixed textbook prior was wrong in 69 of them (97 percent).
Of those 71 wrong claims, the number accepted as VERIFIED was zero. The repo's own reading is that zero of seventy-one puts the false-accept rate below about five percent with 95 percent confidence — not that it is zero. The honest bound is stated in the repo, not extracted under pressure.
In CI on three platforms — 40 Linux ELF, 77 macOS Mach-O, 68 Windows PE files — the prior was wrong on all of them and zero wrong claims were accepted.
Pooled with a third-party ARM64 Linux run: 275 binaries, four formats and architectures, zero wrong claims accepted.
And the control is the part that convinced me. Two small C libraries, compiled on each platform at low and high optimization. At low optimization with the frame pointer kept, the textbook prologue verified nine out of nine. At high optimization, wrong nine out of nine. On the Microsoft compiler, wrong nine out of nine. On ARM64, wrong nine out of nine. The benchmark verifies the prior exactly where compilers really emit it, and refutes it everywhere else.
We also cover the parts that are usually left out: why "all claims verified" is trivially reachable and what a per-claim information weight does about it, why every harness loses state when it summarizes and resets, and how a ledger keyed by file hash survives a crash, a context clear or a brand new session. Then the same rigour applied to ordinary source code, where the oracle moves from bytes to behaviour — including a satisfiability-solver tier that establishes equivalence rather than sampling it.
And the limits, stated plainly. This only works where a deterministic oracle exists. Structural and behavioural facts about an artifact: yes. Claims about the world, judgements about design quality, anything where reasonable people disagree: no oracle, therefore no verification. The guarantee is soundness, not completeness. Pretending otherwise is worse than admitting it, because a verification gate that cannot check anything is a green light with nothing behind it.
Chapters:
00:00 — Cold Open: Wrong 97 Percent of the Time
01:22 — Priors Are Not Ground Truth
02:45 — Make the Tool the Judge
04:17 — The Toolkit Under the Verdict
05:47 — Sixty-Nine of Seventy-One, and Zero of Seventy-One
07:51 — Grounded Means Informative
09:34 — State That Survives the Reset
11:43 — The Same Rigour Without a Binary
13:01 — What This Actually Is
14:31 — Close: Stop Auditing Confidence
Sources:
reverify — https://github.com/2akouwu/reverify (MIT)
README.md — claim grammar, verdict tiers, information weighting, ledger, orchestration
BENCHMARK.md — the 69/71 reference run, three-platform CI, the -O0/-O2 control corpus
EXAMPLE.md — kernel32.dll worked example, textbook prologue refuted at entry RVA 0x2c500
PyPI — https://pypi.org/project/reverify/
#AI #LLM #MachineLearning #AIEngineering #AIAgents #Verification #Hallucination #DeterministicTools #ReverseEngineering #BinaryAnalysis #Python #MLOps #AISafety #AIDeepDive