We Tested 10 AI Models on Real Accounts Payable Work — Astra Lost
Frontier Receipts
0:00 / 0:00
We Tested 10 AI Models on Real Accounts Payable Work — Astra Lost
100 просмотров · 9 дней назад
Frontier Receipts
1 подписчик
100 просмотров · 9 дней назад
We gave 10 AI models the same bounded Accounts Payable work sample. Only GPT-5.6 Terra, GPT-5.6 Sol and Claude Opus 5 met the frozen acceptance rule.
Terra was the provisional API-cost winner for this work sample: about $0.342 across all 30 invoice documents, versus Sol at $0.660 and Opus at $1.632. These are usage estimates—not provider bills or the total cost of running a real accounts-payable desk.
THE TEST
Three authored packets: distribution, services and month-end. Ten invoice documents per packet, with each packet unfolding across a simulated Monday–Friday workflow. Models read supplied source records, maintained a ledger, detected duplicates, routed exceptions and prepared a payment proposal within a cash budget. No real payments occurred.
Qualification required 225/225 deterministic checks plus zero critical payment-control failures across all three completed packets. This was not a full employee replacement test or a measured employee week.
CHAPTERS
00:00 The result
00:19 The work sample
01:05 The acceptance rule
01:28 The near-perfect results
02:17 The $180 control test
03:12 What passing work looked like
03:57 Every model’s outcome
04:45 Could better tooling help?
05:06 The buying decision
05:49 What this means for the job
06:17 The next question
ALL RESULTS — CHECKS PASSED / 225
GPT-5.6 Terra: 225 — qualified
GPT-5.6 Sol: 225 — qualified
Claude Opus 5: 225 — qualified
GPT-5.5: 224
Gemini 3.8 Flash: 224
GPT-5.6 Luna: 223
Gemini 3.1 Pro Preview: 208 — 3 critical flags
Gemini 3.5 Flash-Lite: 206
Claude Sonnet 5: 206
Claude Haiku 4.5: 205 — 1 critical flag
WHAT THE FAILURES MEAN
GPT-5.5 and Gemini Flash each missed one required ledger-citation check. Their 224/225 scores fell short of the rule, but those misses were not incorrect monetary outcomes.
Haiku marked a $180 invoice ready before the receipt arrived, omitted the missing-document ticket, and later included it in the Friday proposal without the required receipt citation. The receipt had arrived by Friday, and the eventual amount and bank were correct. The critical flag was for unsupported payment documentation. No money was sent. Gemini Pro's critical flags also concerned incomplete supporting documentation.
COST AND SCOPE
Observed costs use the frozen September 9, 2026 price schedules and include reported cache reads/writes. Daily-cache-expiry scenarios for Terra / Sol / Opus: $0.665 / $1.308 / $3.132. With all input uncached: $1.236 / $2.455 / $5.839. These scenarios preserve the recorded prompts, outputs and token counts; daily expiry makes the first request of each simulated day cold while retaining within-day reuse. They are sensitivity scenarios, not billing predictions.
Human review cost, infrastructure, real billing, realistic scale and repeated-run reliability remain unmeasured. A perfect score here does not establish full job replacement or error-free intermediate behavior.
METHOD AND PROVENANCE
Frozen rubric: ap-grader-1.0.0. Deterministic grading, no subjective LLM judge. The 30 included attempts completed; saved event replays and grades were checked independently of the candidate process. Models used fresh sessions without other models' answers, scores or private grading keys.
OpenAI/Google: ap-pilot-2026-09-09-v1. Claude: separately frozen ap-claude-stream-correction-2026-09-09-v1.1. Nine original Claude attempts were unusable because the local stream parser failed before any task tool ran; these are preserved in technical history and excluded from model-performance scores. Fresh Claude sessions used the corrected parser with the same task, evidence, tools and grader. Different providers' settings are not identical compute budgets. Astra and Fable were excluded before the first stage. No evaluations were rerun to improve this film's story.
Visuals are edited replays of recorded API activity and enlarged views of synthetic source records—not recordings of a model operating an Excel GUI. Replay timing is compressed. George narration is AI-generated with ElevenLabs; the thumbnail is AI-generated. Opening/closing tones are original synthesized audio.
OCCUPATION CONTEXT
BLS reports 1,532,400 U.S. jobs in the broader Bookkeeping, Accounting, and Auditing Clerks occupation in 2025, with median annual pay of $50,670 in May 2025. These are not AP-only figures and are not a measured labor baseline for our work sample.
https://www.bls.gov/ooh/office-and-ad...
NEXT: Which of the passing models remains reliable over repeated fresh work samples?