Grok 4.7 Plays Factorio — 100 Green Science | High, Guided V01
FactoryArena
0:00 / 0:00
Grok 4.7 Plays Factorio — 100 Green Science | High, Guided V01
1 480 просмотров · 8 дней назад
FactoryArena
15 подписчиков
1 480 просмотров · 8 дней назад
One AI agent. A fresh Factorio world. The challenge: machine-produce 100 green science packs.
Watch Grok 4.7 tackle a FactoryArena benchmark through a Python game interface using guided-green-v01 strategy instructions. This is the full recording, not a highlights edit.
RESULT
Success: 100/100 target packs produced by team machines.
Game time: 1:58:48 (rounded)
Wall-clock benchmark time: 1:58:53 (rounded; host-observed)
RUN SETTINGS
Model: Grok 4.7 (grok-4.7, not Fast)
Reasoning effort: High
Game: Factorio 2.0.77
Time limit: None (game or wall clock)
Harness: native Grok Build ACP 1.0.41
Map seed: 19770917
Instructions: guided-green-v01, byte-identical to the other Grok effort and the Astra guided High reference.
Fresh agent conversation, empty learning directory and fresh world.
Naturally generated terrain and resources, without the former hand-placed starting patches. Fixed starter inventory and custom benchmark rules remain.
HOW IT WORKS
This is a coached strategy experiment, not an unguided baseline. Frozen instructions recommend early resource production and powered research, fuel reserves, movement checks, overlapping research with ingredient preparation and parallel production. The agent chooses coordinates and executes its own Python programs. Belts and inserters are optional; manual item transfers are allowed. Machine-produced science does not mean a fully automated factory. There was no mid-run coaching, retry or model substitution.
Ready was the first game/MCP action, after lazy harness tool discovery. Pre-ready discovery and setup are excluded from the measured game time.
MODEL COST
This run used an existing Grok login, without API-key fallback or credit purchases. An isolated billed cost is not established, so no dollar estimate is claimed. Provider usage covers the full agent session, including setup/discovery and reflection, not just timed gameplay. Cached input is included in input tokens, and reasoning is included in output tokens. Hosting, recording and thumbnail generation are separate.
Provider-reported usage: 25,428,166 input tokens; 307,899 output tokens; 122 model calls.
CHAPTERS
00:00 Setup and introduction
00:12 Run started
29:44 Steam power unlocked
48:16 First lab and red science unlocked
52:20 Power confirmed — lab researching
53:51 Automation researched
1:04:39 Green science unlocked
1:06:38 First green science machine-produced
1:23:31 25 green science packs
1:37:12 50 green science packs
1:53:29 75 green science packs
1:59:06 Success — 100 green science packs
ABOUT THE BENCHMARK
FactoryArena explores how AI agents plan, build and recover from mistakes. The server independently checks the objective. This is a custom scenario, not a default-settings speedrun. High and Extra High use identical frozen instructions, map baseline and Grok runtime; one attempt per effort is exploratory, not a reliable ranking. High took 1:58:47.817 of game time; Extra High took 1:31:04.467 (23.3% less). Astra guided High took 26:18.2 with the same strategy instructions and map baseline, but the cross-model reference also changes harness. Older handcrafted-map runs are not controlled comparisons.
Chapters use the original recording timeline, including setup and post-run reflection, and were checked against sampled video frames and authoritative events.
Thumbnail: AI-generated illustrative cover art made with Codex, not an in-game screenshot.
#Factorio #AI #FactoryArena