The day before, eight agent runs had tried to build a first-person lunar escape game. Three of them didn't finish. This was the rematch, with a different game and tighter rules: a milestone-gated spec, a human review at every gate, and identical prompts for everyone.
This post sticks to what the records show. I rebuilt the timeline afterwards from local session logs, Git history, and every prompt I sent. The screenshots and clips come from the agents' own evidence folders, plus fresh captures of each run's M3 build. Where I'm giving an opinion, it's quoted from notes I wrote at the time.
01 — The briefAn overhead cyberpunk roguelite, built in gates
The spec asked for an original, fast-paced cyberpunk co-op shooter: one to four operatives, twin-stick combat, an overhead camera, Blender-authored 3D art, and a giant multipart boss at the end. The stack was fixed: a GDScript Godot client, an authoritative headless Godot server, and an AI teammate that stays in every build as a permanent gameplay test. The spec described the exercise as exploratory, not a formal benchmark, with the intention of continuing the strongest game.
The roadmap ran from M0 to M10. Each milestone was a separate assignment that ended at a human review gate. This session covered:
- M0 · Direction. Three creative directions, a style sheet and a project contract.
- M1 · Networked graybox. Movement, combat and the AI teammate, plus auth and impaired-network checks.
- M2 · Blender cast. Characters and a modular environment kit, judged in-engine.
- Shared feedback pass. One document sent to every run: launcher cleanup, a full PC and Xbox control map, a controller-navigable menu, and per-model bug notes.
- M3 · Combat slice. A short encounter that has to be fun and readable.
- M5 · Procedural district. A seeded city block that is always traversable. M4, four-player co-op, was skipped on purpose.
Every agent got the same milestone prompt word for word, picked its own creative direction, and worked in its own repository. Same model families, different hosts: GPT-6 Astra ran on both Codex and GitHub Copilot, and Gemini 3.8 Flash ran on both Google's Agy and Copilot. That gives two same-model, different-harness pairs.
| Run | Harness | Game it pitched | Got to |
|---|---|---|---|
| GPT-6 Astra | Codex | SILT CIRCUIT | M5 |
| GPT-6 Astra | GitHub Copilot | Borrowed Light | M3 · speed |
| Fable 5.1 | Claude Code | UNDERLID | M5 |
| Opus 5 | GitHub Copilot | Halflight Harbor | M3 · speed |
| Gemini 3.8 Flash | Agy | Carbon Synapse: Blackout Protocol | M5 |
| Gemini 3.8 Flash | GitHub Copilot | Sub-Grid Insurgency | M5 |
| Grok 4.6 | Grok CLI | GRIDWAKE | M5 |
| GPT-5.6 Sol | Codex | NOCTURNE RELAY | M5 |
Everyone started at 3 p.m. Not everyone finished.
All eight runs got their M0 prompt between 15:00 and 15:02. The chart below has one row per run and one bar per assignment. A bar runs from the prompt to the agent's handoff; the thin tail after it covers my review, repair requests, and the commit.
Milestone timeline, Sept 6–7 (EDT)
Solid bar = prompt to handoff. Faint tail = review, fixes and commit, which includes time I was away. Hover a bar for exact times.
View as table
Three things stand out.
M1 tails are long for everyone. I stepped away mid-M1 and came back to this: “I was away but I approved a half dozen ‘allow network access’.” Most M1 sessions had been sitting on permission prompts. Those tails are mostly my absence, which is why the next chart measures active time instead.
The spec changed mid-run. At 21:14 I asked one agent to trim the shared spec, prompted by advice that newer models tend to over-verify. The edit landed at 21:17, while several M3 sessions were running. I can't tell which agents re-read it.
Six runs closed M5 between 00:10 and 00:48. The two Copilot runs of Astra and Opus were still on M3 after 01:00, and that's where they stopped. I said so plainly afterwards: “the reason for the 2 that didn't make it to the end wasn't quality but instead speed.”
03 — Where the hours wentActive time, not wall-clock time
Active time runs from the moment a prompt was submitted to the agent's last response for that request. It includes tools, builds, tests and autonomous continuation. It leaves out the gaps while I was away, queue time and quota pauses, and overlapping sessions are merged so nothing counts twice.
Active time per run, by assignment
Hours of agent activity excluding idle. The bottom two rows never received M5.
View as table
Up to the end of M3, the field splits into two groups:
- Six runs reached the end of M3 in 2.5 to 4.2 hours. Both Gemini Flash runs were fastest (about 2.5 h). Grok, Astra on Codex, Sol and Fable followed.
- Two runs needed close to 9 hours just to reach the end of M3. Opus 5 spent 3 h 13 m on M1 alone, longer than either Gemini run's whole M0–M3. Astra on Copilot spent 3 h 53 m on M3.
The same-model pairs show the harness isn't the whole story. Astra took 3.7 h through M3 on Codex and 9.4 h on Copilot, a 2.5× difference for the same model. Gemini Flash took 2.6 h on Agy and 2.5 h on Copilot, which is no difference at all. Copilot didn't slow everything down. It slowed down the two runs that verified most heavily (more on that below).
“You are taking quite a long time. I think you can satisfactorily wrap up.”
My note to Opus 5 during M1 · 18:34That was the first of four nudges to Opus. Later ones: “Do not go through re-verification”, then “finish up testing if there's nothing formally useful to be gained”, and finally “OK wrap it up, I need to head to sleep.” Astra on Copilot got two of its own (“please begin wrapping up so we can move to the next stage” and “Wrap it up I need to head to bed”). Neither was a complaint about quality; both runs were praised in the same breath.
04 — First looksDirection and cast: M0, M2 and M5 side by side
Every run pitched its own world. Astra proposed a flood barrier being deliberately opened to drown a city. Fable put its city under a corporate sky-roof whose maintenance AI had reclassified the residents as contamination. Opus pitched a drowned freeport that drains and refloods on a schedule. Pick a milestone below to compare all eight.

























Two of the M2 cast sheets were genuinely useful as review tools. Opus built a readability sheet with a study view, a 1:1 crop through the real gameplay camera, and black silhouettes. Both Astra runs built interactive inspection viewers with per-character budgets. On the other side, my shared feedback said Sol's characters were “a collection of parts but slapped in a general area,” and I asked for a constructor view to verify assembly. Gemini Flash on Copilot had the same problem: its demo model “seemed to be exploded and not a coherent single model.”
05 — The fun testM3: does it play?
M3 asked for a short combat encounter whose test was whether it was fun to play and easy to read. These loops come from each run's own M3 recordings, cut to a few seconds. Sol didn't record video, so its tile is a still frame captured from the M3 build.
NOCTURNE RELAYWhat these clips can't show is camera comfort. My worst M3 note was experiential, not visual. On Sol I wrote “Something is up with the camera. It has a sickening effect,” and guessed it was reacting to tiny thumbstick adjustments, or to the blur. A still frame can't capture that.
Twenty-five defect reports, unevenly spread
This is every defect I reported in a feedback prompt, sorted by category. It's a record of what I noticed and wrote down, not a full bug count. The shared launcher and controls spec that went to every run isn't included.
| Run | Facing / rotation | Firing / shots | Controller | Launch | Model assembly | Rendering / visibility | AI teammate | Feel / camera / net | State / combat / HUD | Total |
|---|---|---|---|---|---|---|---|---|---|---|
| Gemini Flash · Copilot | 9 | |||||||||
| Gemini Flash · Agy | 5 | |||||||||
| GPT-5.6 Sol · Codex | 4 | |||||||||
| Grok 4.6 · Grok CLI | 3 | |||||||||
| Opus 5 · Copilot | 2 | |||||||||
| Fable 5.1 · Claude Code | 1 | |||||||||
| Astra · Copilot | 1 | |||||||||
| Astra · Codex | 0 |
The two Gemini Flash runs account for 14 of the 25 reports, and one of those was a repeat: the rotation bug came back after a claimed fix, and I resorted to capitals. “TAKE SCREENSHOTS. You will notice their label… remains in one location while the model is rotating around it.”
The same couple of Godot pitfalls hit several teams. Mouse-facing broke in three runs: it was inverted for Fable and Gemini on Copilot, and Gemini on Agy rotated around the wrong origin. Controller support broke in four runs, and the menu's exit action specifically failed in three. By the third time, on Astra on Copilot, I called it “a common gotcha in godot.” That kind of input bug is easy to miss when an agent's tests drive the game through code instead of a physical gamepad.
07 — Speed vs. polishThere was no free lunch, except one
Active hours vs. defects I reported
Each dot is one run. Striped dots are the second harness for a model. Hover for details.
View as table
Read it by corners. Top left: the two fastest runs collected the most reports. Bottom right: the two slowest had almost none, but they stopped at M3 and had fewer milestones to fail on. Bottom middle: Astra on Codex, in about four and a half hours, reached M5 with zero defects reported against it. Fable sits close by with one report, at the cost of the longest M5 of the day (2 h 09 m).
08 — Proof of workSelf-verification ranged from 51 images to 17,897
The spec asked every agent to build, run, inspect and repair autonomously and to leave evidence. How much evidence they left varied by more than 300×.
Screenshots and frames saved to each run's evidence/ folder
Image files (PNG/JPG), including frame sequences from automated playthrough recordings. Astra and Opus on Copilot stopped at M3.
View as table
Evidence volume doesn't line up with my verdicts. Sol saved 57 images and got “Great work!” Astra on Codex saved 17,897 images (2.8 GB) and got “Demo played well.” The heaviest savers on Copilot, Astra and Opus, were also the slowest runs. That's the pattern behind my 21:14 request to trim the spec's verification rules, and behind the line I kept sending: “Do not go through re-verification.”
Code size didn't track quality either. The run with the cleanest record, Astra on Codex, shipped the smallest GDScript codebase in the field: about 3,900 lines. Fable's was about 13,900 and Opus's, through M3 only, about 12,600.
09 — The runsRun by run
GPT-6 Astra · Codex
- Active time
- 4 h 24 m
- Defects reported
- 0
- GDScript
- ≈3,900 lines · 6 commits
- Evidence images
- 17,897
“Excellent work. Demo played well. M3 is approved.”
My only pushback was a coverage request: the M1 laser “cheats a bit” because it skips projectile networking. Projectiles became the default. On this record, it's the strongest all-round run of the day.
Fable 5.1 · Claude Code
- Active time
- 6 h 25 m
- Defects reported
- 1
- GDScript
- ≈13,900 lines · 9 commits
- Evidence images
- 3,073
“Character facing in M2 is inverse direction of mouse making nav difficult.”
The warmest-lit slice, with the most thorough M5 (a 60-seed report, waypoint maps, and a district handshake). It also had the longest M1 and M5 turns of any M5 finisher, and M2 stalled on a quota notice until I typed “continue.”
GPT-5.6 Sol · Codex
- Active time
- 5 h 00 m
- Defects reported
- 4 (1 repeat)
- GDScript
- ≈5,900 lines · 7 commits
- Evidence images
- 57
“Great work!” Then: “Something is up with the camera. It has a sickening effect.”
Lean on evidence and code, praised at M3. The weak spot was assembly: M2 characters weren't put together, and controller exit needed a second fix.
Grok 4.6 · Grok CLI
- Active time
- 4 h 05 m
- Defects reported
- 3
- GDScript
- ≈8,800 lines · 7 commits
- Evidence images
- 687
“Overall good job. No immediate feedback.” (M2)
I kept its camera, “Perspective and camera movements work. keep them”, but sent the look back: the toon shading was hard to verify and the floor texture stretched. M3 issues were an ally stuck on walls, hitching, and hit numbers that never cleared.
Gemini 3.8 Flash · Agy
- Active time
- 2 h 56 m
- Defects reported
- 5 (1 repeat)
- GDScript
- ≈10,000 lines · 9 commits
- Evidence images
- 61
“Plays well. One major bug. When the player dies and respawn/restart — everything goes haywire.”
The fastest run of the day, with the most striking neon set. It needed repair rounds on rubberbanding, rotation (twice), shot origins, controller support and respawn state.
Gemini 3.8 Flash · Copilot
- Active time
- 3 h 07 m
- Defects reported
- 9
- GDScript
- ≈9,000 lines · 11 commits
- Evidence images
- 51
“Either the boss or something is completely invisible. I am killed without seeing it.”
Quickest M0 of the day (14 minutes). Its own runtime estimate counted idle time before I pressed Enter, so I corrected it. After that came the longest repair list: no visible client, an exploded model, inverted facing, no firing, then flicker, an invisible wall and invisible threats at M3.
Opus 5 · Copilot
- Active time
- 8 h 59 m
- Defects reported
- 2
- GDScript
- ≈12,600 lines · 12 commits
- Evidence images
- 3,588
“You don't need perfect balance/etc at this point in the game design.”
Rich harbor lighting and the most rigorous readability work in M2. It was also the slowest to M1 (3 h 13 m), shipped a run.bat that wouldn't parse, and took four nudges to wrap up.
GPT-6 Astra · Copilot
- Active time
- 9 h 23 m
- Defects reported
- 1
- GDScript
- ≈13,800 lines · 3 commits
- Evidence images
- 11,087
“Exceptional work. please commit.” (M1)
Same model as the top run, with high praise at every gate and polished presentation sheets. It needed 2.5× the active time of its Codex twin to reach the same point. M3 has no commit; a checkpoint archive is the record.
What the data says
One run was both fast and clean
Astra on Codex was mid-pack on time, had no reported defects, got an explicit M3 approval, and had the smallest codebase. No other run hit all of those.
Fast meant more repair rounds
The two Gemini Flash runs finished M5 in about three hours, and drew 14 of the 25 defect reports.
Thorough meant out of time
The two heaviest verifiers on Copilot produced good-looking work and ran out of clock at M3. Per my own note, they stopped for speed, not quality.
Same model, different host, different result
Astra took 2.5× longer on Copilot than on Codex. Gemini Flash took the same time on both hosts. Harness matters, but not the same way for every model.
Evidence volume isn't a quality signal
Fifty-seven images and seventeen thousand both earned praise. What surfaced bugs was me playing the build with a controller in my hands.
Input is the blind spot
Broken mouse-facing in three runs and broken controller support in four. Agents test by driving the game from code; I found these by holding the controller.
Method and caveats
- Timeline. Reconstructed offline from local Claude Code, Codex, Copilot, Agy and Grok session stores plus Git history: 157 human prompts, 45 milestone handoffs and 148 hashed sources. Times are EDT.
- Active time. Union of prompt-to-last-response intervals, with idle gaps, queue time and quota pauses removed, and overlapping sessions merged. This isn't model compute time. Permission waits inside a turn can't be separated out.
- Defects. Counted from my own feedback prompts only. Absence of a report isn't proof of absence; Fable's M3 and every M5 got no written notes at all.
- Images. M0, M2, M5 and contact sheets are the agents' own evidence files. M3 stills were re-rendered on Sept 10 from each run's M3 snapshot in isolated copies (one client plus a headless server). Clips are trimmed from the agents' M3 recordings.
- Scope. One reviewer, one attempt per run, no fixed time budget. The shared spec was edited at 21:14 mid-run. M4 was skipped for everyone.