Перейти к содержимому

Both AI Coding Agents Passed Our First Test. One Fix Was Still Broken.

Agent Proof Lab

0:00 / 0:00

Both AI Coding Agents Passed Our First Test. One Fix Was Still Broken.

2 просмотра · 7 дней назад
Agent Proof Lab
2 просмотра · 7 дней назад
Codex and Claude each received the same real Vercel AI SDK bug, the same pinned source commit, the same time limit, and three independent attempts. All six patches passed our first evaluator. One was still behaviorally incomplete. In this pilot episode, we show the missing asynchronous return, strengthen the hidden audit, and compare reliability, speed, and patch scope without pretending that one task is a universal model ranking. Results: • Codex: 3/3, median 5:42 • Claude: 2/3, median 8:18 • Codex median changed files: 18 • Claude median changed files: 10 Methodology note: the evaluator was strengthened after the original runs. Audit v4 includes delayed-completion, stream-read failure, response-write failure, cleanup, public-wrapper propagation, type checking, the original open test, and the full Node plus Edge package regression suite. It was applied equally to all six unchanged workspaces. Future experiments will freeze the full hidden suite before execution. Tool configuration: Codex CLI 0.153.4 through the existing ChatGPT subscription; its exact subscription-default model identifier was not emitted in the preserved JSONL and is not claimed. Claude Code 2.1.71 used claude-opus-4-6 according to preserved run metadata. Agent timing starts immediately before each CLI process and ends when it exits; dependency installation and evaluator execution are excluded. The task policy prohibited web and Git-history lookup, but the runner did not independently enforce a network firewall. Preserved commands show no prohibited lookup attempt; each CLI retained its provider service connection. Evidence report: https://agent-proof-lab-e001.mdtech-top.ch... Chapters: 00:00 First tests passed — one fix was wrong 00:42 The real asynchronous contract 01:39 Same task, same rules, three runs each 02:30 Tracing every public promise path 03:19 What each agent changed 04:09 The wrapper one patch missed 05:09 Why the original result was a tie 05:26 Strengthening the evaluator 06:08 Final retrospective audit 06:28 Three passes versus two 06:48 What this means for engineering teams 07:27 What we would ship 07:45 Limitations and the next experiment