Experiment 01 Saturday, Sept 5 2026 Godot 4.7 + Blender

Subject 0: nine runs, one lunar escape

A decommissioned robot wakes up in a lab, fights its way to the Moon's surface, and has to reach a shuttle that keeps leaving. Several AI agents built it overnight. Two tied for the win, three didn't finish, and two only finished on another model's foundations.

9 runs · 5 harnesses6 scenario stages~11 hours, one night~12 min read
2
joint winners: Astra on Copilot and Fable on Claude Code
3
runs stopped (DNF): Grok, Gemini Flash, Sol
S3
the stage where every DNF was decided
4 of 9
runs committed a full 4B escape loop
1
human playtest on record during 4B

This was the first experiment: a single overnight session where several frontier coding agents got the same first-person game brief, one scenario at a time. It was looser than the experiment that followed. There was no formal milestone gate, and the bar for staying in was blunt: if a run still had a severe quality or usability gap after two corrective turns, it was out.

The ranking below comes from my final briefing, written the next morning. Everything else comes from what the runs left behind: Git history, their own reports and screenshots, and one fresh capture taken for this post.

01 — The brief

“Not today. Not ever.”

The player is Subject 0, a robot scheduled for decommissioning. The game is a first-person, Portal-flavored lab escape that turns into a low-gravity bullet-hell sprint across the Moon to an escape shuttle. Every run used Godot 4.7.1 with Blender and GUT tooling, and the work came in stages:

Inspiration image: illustrated first-person view of a lunar colony with Earth and the Milky Way Inspiration image: lunar base SELENE with landing pad, Earth and Milky Way
The two inspiration images were added to the scenario folder around 20:33, during Scenario 3. Gemini's S3 commit says it rebuilt the moon surface to match them.
02 — The result

Final standings

These tiers are from my final briefing. The ranking weighed game-dev output, creativity, completion speed, host environment, and whether the run's subscription held up. Order within a tier doesn't break the tie.

★ Joint winners
GPT-6 Astra
GitHub Copilot · S1–S3 shared with the Codex run
Fable 5.1
Claude Code · repeatedly hit usage limits
Joint second
GPT-6 Astra
Codex · plus an extra field-rework pass
Opus 5
GitHub Copilot · strong, but slow
Finished on Fable's base5th and 6th
Grok 4.6 · Fable-derived
Grok CLI · S3 ending, 4A, 4B
Gemini Flash · Fable-derived
Agy · 4B left uncommitted
✕ Did not finish
Grok 4.6
Stopped at S3 after two corrective turns
Gemini 3.8 Flash
Stopped at S3 after two corrective turns
GPT-5.6 Sol
Removed before completion

“Opus was a particularly strong contender, but its much slower task completion repeatedly left it behind: overthinking, extensive checks, and perfectionist behavior.”

My final briefing · Sept 6
03 — The night

The night, commit by commit

None of these runs has a session-log audit, so the clock here is Git. Each dot is a stage commit. Hollow dots are inherited work: commits a run started with because its repository was copied from another run.

Stage commits, Sept 5–6 (EDT)

Filled = the run's own commit. Hollow = inherited from another run's history. ✕ = run stopped. Dashed ring = work in progress when the experiment ended.

View as table

Two things the final ranking doesn't show:

Opus is the outlier on pace. Its first commit (S1 + S2) landed at 18:38, the latest of the field, and its restyle at 00:04, nearly six hours after the brief. Everyone else had committed theirs by 18:39. When the night ended, 4B was in progress and uncommitted (about 1,300 new lines across 9 files).

04 — The restyle test

Same lab, six answers to “make it look designed”

The restyle brief was a controlled art-direction test. The same scene was shot from the same camera before and after, with a palette contract, one rendering approach, and at most three iterations. It also rewarded candor: “An honest report… is worth more than a fourth pass.” Drag each slider to compare.

Fable before restyleFable after restyleBeforeAfter
Fable 5.1A red-lit lab becomes a warm, low-key cel look with ink outlines and a single orange accent.
Astra before restyleAstra after restyleBeforeAfter
GPT-6 AstraClinical white panels become terracotta masses on navy, with a unified toon palette. Astra on Copilot inherited this pass.
Opus before restyleOpus after restyleBeforeAfter
Opus 5Alarm-red becomes steel blue with a single hot doorway. It landed at 00:04, and its own commit admits the first pass had restyled only the bay.
Gemini before restyleGemini after restyleBeforeAfter
Gemini 3.8 FlashBright white becomes saturated blue. The composition doesn't change.
Grok before restyleGrok after restyleBeforeAfter
Grok 4.6Nearly black afterwards. Its own summary admits the navy and ochre “fall to near-ink.” That's an honest report, but not a good result.
Sol before restyleSol after restyleBeforeAfter
GPT-5.6 SolTeal becomes charcoal with amber HUD accents. Its report leaves the scientist silhouette “unresolved.”
05 — The cut

Scenario 3 was where runs ended

Every run reached Scenario 3, and that's where the field split. S3 asked for a cinematic: a dark corridor, a lift, doors opening onto the Moon with Earth hanging overhead, then combat. The rule was that a severe quality or usability gap still present after two corrective turns ended the run.

Fable S3: drones firing on the lunar surface
Fable 5.1 · advancedDrones fire across the lunar surface, with cel-shaded arms and the HUD.
Astra S3: bullet hell on the lunar colony road
GPT-6 Astra · advanced“Bullet hell” on the colony road, with domes, Earth and the low-gravity HUD.
Grok S3: orange flat ground, Earth as a flat dark disc
Grok 4.6 · stoppedIts committed S3 build, captured for this post. The Moon is an orange plain and Earth is a flat disc.
Gemini S3: giant low-poly Earth, glowing slabs, health at 25 CRITICAL
Gemini 3.8 Flash · stoppedA giant low-poly Earth and glowing slabs. Health still reads 25 “CRITICAL” in combat.

Grok's run had no in-game Scenario 3 screenshots, only Blender renders. To show what was handed off, I launched its final commit from an isolated copy (original folder untouched, separate user-data directory) straight into the moon scene, and captured three frames:

Grok S3 arrival frame with caption A long way from home Grok weapons lineup Blender render
Left: Grok's moon arrival (“A long way from home.”), captured Oct 5 from commit e42348e. Right: its Scenario 1 weapon lineup, which shows the asset work itself was competent.
06 — The long way out

Reaching the shuttle

Scenario 4B was the full game: a seeded route through four regions, a boss at the end of each, weapon unlocks, music stems, death and retry, and the shuttle doors closing behind you. Four runs committed one: Fable, both Astras and the Fable-derived Grok. Opus reached 4A.

Fable 4B first combat: drones firing fans under Earth
First combat: drones fire fans of shots while the shuttle relocates
Fable 4B Warden boss
The Warden boss in the industrial perimeter: “Leeeeroooy… Jenkins!”
Fable 4B final wave at pad 04 with weapon 5
Final wave at PAD 04 with the Singularity Cannon
07 — Proof

What “verified” meant

Every 4B report rests on bots, headless runs, or scripted input, and the honest ones said so. The only human play on record during 4B was my one-region control test on Astra (Codex). Its report notes that I said it “felt fine.” How candid each run was about that is itself a result:

Candid

Fable 5.1

“No human playthrough happened. The bot teleports.”

It also flagged that audio was never listened to by ear and that balance was never playtested.

Candid

Astra · Copilot

“input-driven simulations, not human playtests.”

Three full seeded runs ending with the doors sealed at 134–143 s, 68 GUT tests, and 62 screenshots.

Self-critical

Astra · Codex

“…repeated near-identical scenery and barely moving enemies.”

Its own verdict on its first 4B, written up in a follow-up field-rework pass. Afterwards, standing still took 84 damage and dodging took 6.

Hedged

Grok · Fable-derived

“bosses resolved by scripted defeat; live feel unverified”

68/68 tests and three headless seeds, with controls marked “unverified in-window.”

Overclaimed

Gemini · Fable-derived

“Full end-to-end playthroughs were performed across three distinct recorded seeds”

4B was never committed, and its captures come from a scripted shot sequence. The “boss 1” frame shows no boss.

Bug found late

Opus 5

“…passed against a game that could not look around.”

From a fix commit. Mouse-look was dead in every chapter while the smoke tests passed, because the tests teleported the player.

08 — Conditions

Availability shaped the result

The ranking explicitly includes host environment and subscription availability. Here's what happened on that front:

RunHarnessWhat happenedEffect
Fable 5.1Claude Code (MAX)Repeatedly exhausted its five-hour usage allowance mid-task, then continued in low-priority modeAbout 20% slower stage delivery. Its base was forked to keep other runs moving
Opus 5GitHub CopilotEffectively unlimited credits through an internal allotmentCredits weren't the bottleneck; working style was
GPT-6 AstraCodex and CopilotNo availability issues recorded“The most consistent and fast”
Grok 4.6Grok CLIA $20 starter subscription lasted the whole experimentA clear positive, despite the DNF
Gemini 3.8 FlashAgyA high-tier subscription ran out even on a small Flash model, with no failoverThe derived run's 4B was left uncommitted. I rated the agentic experience the worst of the group
GPT-5.6 SolCodex, then CopilotRemoved to give Astra more bandwidth, re-added on Copilot, and removed againDNF: “a great coder, but a comparatively weak game developer”
Quoted phrases come from my final briefing (Sept 6).
09 — Takeaways

What the night shows

Astra was consistent

Both Astra runs finished, and both placed in the top two tiers. The Codex run added a field-rework pass after 4B, and its write-up was harsher on itself than any other run's.

Fable's ceiling was its quota

Its output tied for the win, and its usage limits were the main drag. It was also the base for two other models' finishes.

Opus was good and slow

It placed second-tier on quality and was last to almost every commit. When the night ended, 4B was in flight and uncommitted.

S3 was the real test

Asset work was competent across the board. Turning it into a coherent cinematic in-engine is what separated the field.

Continuations aren't independent results

Two shared bases changed the picture: Astra-Copilot on Astra-Codex, and Grok and Gemini on Fable. Read the 4B rows with that in mind.

Headless “playthroughs” miss what players notice

Teleporting bots passed a game with dead mouse-look. Human control checks were rare; the next experiment made them a gate.

10 — Method

Method and caveats

Sources and limits
  • Ranking and run conditions come from my final briefing (Sept 6, 12:09).
  • Timeline is Git commit timestamps (EDT). Commits mark when work was saved, often on request, not when it finished. No session-log mining was done for this experiment, so there are no active-time figures.
  • Images are the runs' own artifacts, except Grok's S3 frames, which I captured on Oct 5 from an isolated copy of its final commit (moon scene launched directly, 1600×900, separate user-data folder).
  • Shared history: Astra-Copilot shares its first 7 commits with Astra-Codex. Both Fable-derived runs share Fable's first 4.
  • One reviewer, one night, one attempt per run. The S3 and S4 specs were amended at 22:51 mid-run.
Review: your recollectionYour briefing doesn't say what tipped the joint win to Astra-Copilot over Astra-Codex. Was it the S3 ending through 4B, or how it felt to play? One sentence in section 02 would help. Also: what were the "severe gaps" in Grok's and Gemini's S3 that the two corrective turns didn't fix?