Both AI Coding Agents Passed Our First Test. One Fix Was Still Broken.
Agent Proof Lab
0:00 / 0:00
Both AI Coding Agents Passed Our First Test. One Fix Was Still Broken.
2 просмотра · 7 дней назад
Agent Proof Lab
2 просмотра · 7 дней назад
Codex and Claude each received the same real Vercel AI SDK bug, the same pinned source commit, the same time limit, and three independent attempts.
All six patches passed our first evaluator. One was still behaviorally incomplete.
In this pilot episode, we show the missing asynchronous return, strengthen the hidden audit, and compare reliability, speed, and patch scope without pretending that one task is a universal model ranking.
Results:
• Codex: 3/3, median 5:42
• Claude: 2/3, median 8:18
• Codex median changed files: 18
• Claude median changed files: 10
Methodology note: the evaluator was strengthened after the original runs. Audit v4 includes delayed-completion, stream-read failure, response-write failure, cleanup, public-wrapper propagation, type checking, the original open test, and the full Node plus Edge package regression suite. It was applied equally to all six unchanged workspaces. Future experiments will freeze the full hidden suite before execution.
Tool configuration: Codex CLI 0.153.4 through the existing ChatGPT subscription; its exact subscription-default model identifier was not emitted in the preserved JSONL and is not claimed. Claude Code 2.1.71 used claude-opus-4-6 according to preserved run metadata. Agent timing starts immediately before each CLI process and ends when it exits; dependency installation and evaluator execution are excluded. The task policy prohibited web and Git-history lookup, but the runner did not independently enforce a network firewall. Preserved commands show no prohibited lookup attempt; each CLI retained its provider service connection.
Evidence report:
https://agent-proof-lab-e001.mdtech-top.ch...
Chapters:
00:00 First tests passed — one fix was wrong
00:42 The real asynchronous contract
01:39 Same task, same rules, three runs each
02:30 Tracing every public promise path
03:19 What each agent changed
04:09 The wrapper one patch missed
05:09 Why the original result was a tie
05:26 Strengthening the evaluator
06:08 Final retrospective audit
06:28 Three passes versus two
06:48 What this means for engineering teams
07:27 What we would ship
07:45 Limitations and the next experiment