Перейти к содержимому

Grok 4.7 Plays Factorio — 100 Green Science | High, Guided V01

FactoryArena

0:00 / 0:00

Grok 4.7 Plays Factorio — 100 Green Science | High, Guided V01

1 480 просмотров · 8 дней назад
FactoryArena
15 подписчиков
1 480 просмотров · 8 дней назад
One AI agent. A fresh Factorio world. The challenge: machine-produce 100 green science packs. Watch Grok 4.7 tackle a FactoryArena benchmark through a Python game interface using guided-green-v01 strategy instructions. This is the full recording, not a highlights edit. RESULT Success: 100/100 target packs produced by team machines. Game time: 1:58:48 (rounded) Wall-clock benchmark time: 1:58:53 (rounded; host-observed) RUN SETTINGS Model: Grok 4.7 (grok-4.7, not Fast) Reasoning effort: High Game: Factorio 2.0.77 Time limit: None (game or wall clock) Harness: native Grok Build ACP 1.0.41 Map seed: 19770917 Instructions: guided-green-v01, byte-identical to the other Grok effort and the Astra guided High reference. Fresh agent conversation, empty learning directory and fresh world. Naturally generated terrain and resources, without the former hand-placed starting patches. Fixed starter inventory and custom benchmark rules remain. HOW IT WORKS This is a coached strategy experiment, not an unguided baseline. Frozen instructions recommend early resource production and powered research, fuel reserves, movement checks, overlapping research with ingredient preparation and parallel production. The agent chooses coordinates and executes its own Python programs. Belts and inserters are optional; manual item transfers are allowed. Machine-produced science does not mean a fully automated factory. There was no mid-run coaching, retry or model substitution. Ready was the first game/MCP action, after lazy harness tool discovery. Pre-ready discovery and setup are excluded from the measured game time. MODEL COST This run used an existing Grok login, without API-key fallback or credit purchases. An isolated billed cost is not established, so no dollar estimate is claimed. Provider usage covers the full agent session, including setup/discovery and reflection, not just timed gameplay. Cached input is included in input tokens, and reasoning is included in output tokens. Hosting, recording and thumbnail generation are separate. Provider-reported usage: 25,428,166 input tokens; 307,899 output tokens; 122 model calls. CHAPTERS 00:00 Setup and introduction 00:12 Run started 29:44 Steam power unlocked 48:16 First lab and red science unlocked 52:20 Power confirmed — lab researching 53:51 Automation researched 1:04:39 Green science unlocked 1:06:38 First green science machine-produced 1:23:31 25 green science packs 1:37:12 50 green science packs 1:53:29 75 green science packs 1:59:06 Success — 100 green science packs ABOUT THE BENCHMARK FactoryArena explores how AI agents plan, build and recover from mistakes. The server independently checks the objective. This is a custom scenario, not a default-settings speedrun. High and Extra High use identical frozen instructions, map baseline and Grok runtime; one attempt per effort is exploratory, not a reliable ranking. High took 1:58:47.817 of game time; Extra High took 1:31:04.467 (23.3% less). Astra guided High took 26:18.2 with the same strategy instructions and map baseline, but the cross-model reference also changes harness. Older handcrafted-map runs are not controlled comparisons. Chapters use the original recording timeline, including setup and post-run reflection, and were checked against sampled video frames and authoritative events. Thumbnail: AI-generated illustrative cover art made with Codex, not an in-game screenshot. #Factorio #AI #FactoryArena