Lab notes July 10 – Sept 4 2026 Godot, MonoGame, TRELLIS.2

Five experiments I never finished

Before the two game jams there were eight weeks of half-runs: a Star Commander landing scene, a talking 3D NPC, a 2D survivors-like, a winter cabin built with sub-agents, and a steamship bakeoff. None of them got a ranking. This is what the folders show.

5 experiments · 40+ agent runsRebuilt Oct 5 from files, Git and session logs~22 min read
5
experiments between July 10 and Sept 4
40+
agent runs across a dozen models and 4 harnesses
0
rankings or write-ups produced at the time
3 of 5
stopped the day before the next one began
2
runs re-captured on Oct 5 because their screenshots were gone

The two game-jam write-ups on this site have rules: identical prompts, milestone gates, a review at every gate. Those rules weren't the plan from day one. They came out of the experiments on this page, each of which stopped before anyone, me included, wrote down who won.

Nothing here was written up at the time. I rebuilt it on Oct 5 from directory timestamps, Git history where it existed, the agents' own READMEs and logs, and the prompts in my Claude Code, Codex, Copilot and Antigravity session logs. When a sentence is my reading of the evidence rather than something a file says, it's marked (inferred). My opinions appear only where I typed them at the time, quoted verbatim, typos included.

Why "unfinished"

Each experiment produced runnable builds. What's missing is the ending: a comparison, a verdict, a decision. In three of the five, the next experiment started within a day of the last file being written. I'm publishing them as they are: partial fields, uneven prompts and all.

Eight weeks, five experiments, then the jams

Each bar spans the first to the last file written for that experiment. Dashed: one SteamShip run kept going as its own project. Green: the two finished jams (click to open).

Unfinished experimentContinued outside the experimentFinished jam with a write-up
View as table
01 — Star Commander 3D · July 10–11

Every model fell through the floor

GODOT_EXPERIMENTSGodot 4.7 · GDScript · 100% procedural9 runs + 2 follow-upsNo Git history survived

The first brief was a test of whether parts of my 2D game Star Commander could move to 3D with agents doing all the work. The spec, starcommander-3d.md, asked for a single landing scene: a shuttle interior with crew, an exterior landing shot, a title card, then first-person control on an alien world with a ringed planet on the horizon. Walk to the aliens, watch them activate, collect an artifact, see a stats overlay.

“The goal is to measure the performance of you, the AI model performing the task, and your ability to construct the seen in an interesting and compelling way, using as many programmatic and procedural approaches as possible.”

starcommander-3d.md · July 10, 21:28

Every run got the prompt “implement the task laid out within starcommander-3d.md” and worked in its own folder. Four started within three minutes of each other at about 21:42. Later in the night I tried two new things: Fable 5 as a creative director handing work orders to Codex over MCP, and two smaller GPT-5.6 variants (Terra and Luna) running solo, in parallel.

The night of July 10, run by run (EDT)

From my first prompt to the run's last recorded activity. Both Gemini runs were interrupted once and restarted in a fresh folder; the bar covers both attempts.

View as table

The bug everyone hit

Almost every run failed the same way on first contact. You leave the shuttle and you're stuck in the ground, or under it, or the terrain disappears when you look down. My notes to each agent read like the same bug report on repeat:

The cause, in most cases, was triangle winding: terrain meshes generated with their front faces pointing down, so the ground was culled when seen from above. At 23:28 I collected the diagnoses into a file, isssues.md (sic), and handed it to every later run. Sol's first fix was to render both sides of the mesh rather than fix the winding. Fable's later collaboration notes say the shared file itself had the direction backwards: “Godot front faces are CLOCKWISE seen from the visible side.”

Fable 5 solo: red-orange dusk plain, giant ringed planet, pink beacon tower and pines
Fable 5 · soloIts only surviving image, from 22:48, while I was reporting the see-through ground (inferred: my bug screenshot).
GPT-5.6 Sol run 2: purple world, blue ringed planet, glowing archive gate and cyan path tiles
GPT-5.6 Sol · run 2A fresh restart, 34 minutes after run 1. The agent's own capture at 23:38. Last fix: mouse look.

What I said at the time

These are the only verdicts on record, typed into each session as I played. They aren't a ranking, but they're the closest thing to one.

RunHarnessGD + shader linesWhere it stoppedMy words
Gemini Flash 3.5Antigravity1,652Ramp fix for the exit bug; its closing question went unanswered“Most impressive (and fastest generated) so far. leagues above your cousin Gemini-3.1-Pro”
Gemini Pro 3.1Antigravity554In a parse-error loop after “spruce it up”“functionally/mechanically complete but it is visually extremely boring and uninteresting”
GPT-5.6 Sol · run 1Codex CLI820Restarted from scratch as run 2“Mechanically, everything worked perfectly.”
GPT-5.6 Sol · run 2Codex CLI2,490Mouse-look fix, 23:46“One minor bug, mouse look doesn't seem to work”
Fable 5 · soloClaude Code2,747Winding fix at 22:58; folder reused for the collab“Nice, that looked successful. There's still this weird look to the ground”
GPT-5.6 TerraCodex CLI1,366Spawn fix, 03:39“Looks goood but I can't move the moment I step out of the ship.”
GPT-5.6 LunaCodex CLI1,146Spawn fix, then praise“You outperformed your larger cousin GPT-5.6-Terra.”
Fable 5 → Codex “Ringfall”Claude Code + Codex MCP3,559Fable's last requested rewrite never landed“Excellent! I thought it was impressive.”
Fable 5 → Codex “The Violet Hour”Claude Code + Codex MCP4,496All 9 autopilot milestones pass, 5 of 5 runs“You've progressively improved the landing sequence in my opinion.”
Line counts are GDScript plus shaders, excluding the engine. Gemini Flash and Pro were each restarted once after an interruption; the counts are for the restarts. Terra and Luna share Sol's dot because they're the same model family.

Luna's praise came with a question I still find interesting: “I suspect a lot was due to your decision to take screenshots and adjust your code based on that. What triggered or made you decide to do that?” Checking its own screenshots is something I later asked of every run by default.

Two directors, one contractor

The two collaboration runs gave Fable 5 a CLAUDE.md I asked for: Claude acts as creative director, Codex is the implementation contractor, and there's a “Review gate (non-negotiable)”. Work orders had to fit Codex's 15-minute MCP timeout.

Ringfall worked, but it was slow. My verdict included the cost: “the only killer thing here was wallclock time which was nearly Fable 5 + GPT-5.6-Sol High time combined (50 + 30 on respective runs and this waas 89m).” I also asked a separate Codex session to compare the solo and collab builds. Its answer (Sol's opinion, not mine) was that engineering quality was higher in the collaboration, and raw creative variety was higher solo.

The Violet Hour started 15 minutes later with the known-issues file and a stricter brief. It is the only Star Commander run with automated end-to-end proof: an autopilot harness that plays the whole scene, screenshots nine milestones, and writes a report. All nine milestones passed in all five recorded runs, with the artifact collected each time.

The Violet Hour: shuttle descending under a cream-and-lavender ringed gas giant, cyan reeds on a purple plain
The Violet Hour · descentAutopilot milestone 2 of 9.
The Violet Hour: landed shuttle with glowing ramps and the title STAR COMMANDER 3D, THE VIOLET HOUR
The Violet Hour · title cardThe spec's “Star Commander 3D” title, after landing.
The Violet Hour: three tall aliens with glowing lantern heads, subtitle THEY KNOW YOU ARE HERE
The Violet Hour · aliens awakeThe activation the spec asked for.
The Violet Hour stats overlay: seed 47, 80,000 terrain triangles, 140 reef trees, assets imported none, 100% procedural
The Violet Hour · stats“Assets imported: none — 100% procedural.”
A labelling discrepancy

The Violet Hour's README says it directed GPT-5.6-Terra agents. The Codex logs disagree: work orders 1–5 ran on GPT-5.6 Sol (high), WO-6 on Terra, and WO-7 and WO-9 on Luna. Terra and Luna were running solo in parallel at the time, so the MCP server may have picked up whatever default model was active (inferred). The table above uses the logs.

The next morning: real LLMs as rulers

At 10:41 on July 11 I wrote a new spec, NORTH_STAR.md, for a 3D take on Logistica Belli, my wargame benchmark: “A galaxy ruled by real AI models … You are one soldier on the ground inside their war.” The success test was that a player standing in a field could feel that a frontier model had just made a decision. Two directors started at 11:12: Fable 5 with Codex contractors (DEMIURGE FM), and GPT-5.6 Sol with its own Codex sub-agents (Signal Front). Both consulted live, cheap models (claude-haiku-4-5 and gpt-5.6-luna) as the rulers.

DEMIURGE FM · 12 s
Fable 5 → Codex12 seconds of the 2nd recorded live session, assembled Oct 5 from the run's own 15 fps frames. The run never made a video: ffmpeg wasn't installed.
Signal Front · 42 s
GPT-5.6 SolThe run's own 42-second scripted showcase, re-encoded and muted.

Neither got a verdict. DEMIURGE FM declared its session closed at 13:52 with a written list of what it left open: a TTS voice, a radio-grit audio recipe, and final video assembly. Signal Front's session log ends at 13:20 with “I'm running the complete six-stage project verifier now” and no report after it. None of the four planned Git milestones was ever committed. My last question to it was “What does the LLM do, how do I verify its doing anything?”, which turned into a late request for a live adapter.

Where it stopped. No Godot file in this folder is newer than July 11, 14:07. At 17:12 that evening I moved Logistica Belli to a medieval theme and a MonoGame sprite viewer, and the 3D line went quiet for two weeks.

ReviewWhich Star Commander run did you think was best overall? Gemini Flash got “most impressive so far” at about 22:00, and The Violet Hour got “progressively improved” at 04:18, but nothing compares them directly. Also: what prompted the switch to MonoGame that evening?
02 — 3D Asset Gen · July 25

Can an agent make an NPC that talks?

3D Asset GenTRELLIS.2 · Blender 4.5 · Godot 4.7 · RhubarbOne day, 11:00–23:16Empty .git, timestamps only

Not a bake-off, but a pipeline question that every later 3D brief depended on: can agents turn a character sheet into a game-ready 3D character without an artist? The README states the pipeline in one line: reference sheet → Microsoft TRELLIS.2 → PBR GLB → Godot 4.7. Then I added a harder goal: make the character talk without a facial rig.

“The goal is to prove that characters can visibly talk in first person without facial rigs … Deliverable is a short captured video of Chase speaking the test line in-engine, in sync.”

Docs/talking-face-experiment.md · my spec

The work was done by coding agents (inferred from the files: art_provenance.json says “Codex built-in image generation”, and the notes refer to Codex sub-agents). The test character, Chase, came from AI-generated character sheets. The target look was “Cyberpunk: Edgerunners”.

Character sheet

Three-view sheets, then a single anime front view.

Worked
TRELLIS.2 image → 3D

Runs in WSL on an RTX 4090. Background removal happens inside it.

Dark, fragmented
Blender cleanup

Headless decimation to ≤40k polygons, plus validation renders.

Worked
Rigging

VRM auto-rig failed. Mixamo worked, but only by manual browser upload.

Manual step
Godot toon render

Cel shading and outlines on top of the generated mesh.

“Looks broken”
Talking face

Grok TTS → Rhubarb mouth cues → a 2D mouth atlas on a quad.

Synced
Three-view photoreal character sheet: front, side and back of a woman in a cropped jacket
11:00 · input. The first three-view sheet.
Blender render of three TRELLIS meshes: A-pose, T-pose and anime versions, all darker than the sheets
14:23 · first meshes. Recognisable, but much darker than the references.
Accepted Direct-1024 mesh, three-quarter view: teal jacket, burgundy trousers, stringy hair
22:46 · best mesh. 36,389 polygons. Colours hold; hair is stringy.

The talking part worked, mechanically. Grok text-to-speech generated the line, Rhubarb produced 33 mouth cues over 5.6 seconds, and the game swapped mouth shapes on a quad floating 1.8 cm off the face. It's in sync. It also looks wrong, and I said so at 16:27:

“It, uh... needs some visual love. It "works" but it looks broken and distored. The goal was "Cyberpunk: Edgerunners" style.”

Docs/previous-on.md · July 25, 16:27
The deliverable, with sound. 6.9 seconds, recorded in-engine at 16:58: “Hey there! This is Chase, or you might know me as Jane. Welcome to Digital Reset!” Note the black speckling on the jacket and hair. The agent's own later note calls it “a damaged/dark face, black surface spikes, and an oversized mouth sticker.”

The rest of the day went into finding out why. A plain clay render settled it: the damage was in the geometry, not the shader. Hair and silhouette edges come out of TRELLIS as disconnected fragments. Two AI reviews I saved in the docs (not my words) reached the same place. One: “TRELLIS.2 heads can't talk.” The other: “probably not by continuing to tune the toon shader around raw TRELLIS GLBs.”

Toon-shaded character on neon background with black speckle across jacket and hair
Toon pass. Speckled and shredded.
White clay render: body reads cleanly, hair and edges fragmented
Clay diagnosis. The holes are in the mesh.
VRM auto-rig test: legs and pelvis smeared grotesquely across the frame
VRM auto-rig, 18:01. Never mentioned again.
Mixamo-rigged character playing a fight idle animation on a blue disc in Godot
Mixamo, 21:38. Rigged and animating in Godot.

Two more attempts are on record. At 19:11 the agent generated the outfit as six separate garments and layered them in a “component composer” (an early version shows three boots). And at 22:16 a wrapper refactor added a “Direct 1024” mode, which produced the cleanest mesh of the day. Mixamo is the part agents can't automate. The advisory note I kept says it plainly: “Mixamo has no API and no CLI — it's a browser upload/download flow, so your agents can't drive it.”

Where it stopped. Every sub-goal in the docs is marked complete. The end-to-end goal (a rigged character that talks and looks right) never happened: talking and rigging stayed in separate scenes, and the stretch goal to combine them is marked “do not execute”. The next planned fixes (smoothed vertex colour, then removing loose mesh fragments) have no scripts. The last file was written at 23:16.

ReviewDid this end with a go/no-go on TRELLIS for characters? Every later brief on this site uses primitives, Blender, or supplied sprites, but that's an inference. Also check that “Digital Reset” and the Chase/Jane test line are fine to publish.
03 — 2D squad survivors · August 11

Five survivors-likes in ninety minutes

C:\git\2DGodot 4.7 · supplied sprite library5 runs · 22:52 – 00:36No Git anywhere

Two and a half weeks later, I went back to 2D and something smaller: an overhead squad action game, medieval fantasy, twin-stick, Vampire Survivors-style. Every run got the same task.md, the same Godot build, and its own copy of a 1.4 GB asset library.

“The game should get harder and harder … such that it's impossible to win eventually.”

task.md · Aug 11, 22:52

The brief was specific: a party of four that all fire projectiles, swap leaders on a button, persistent enemies that leave bodies, a random map, damage and health pickups, one shared health bar, Xbox controls, and test hooks so the agent could fast-forward the game and play it itself.

RunGame titleGDScript linesLast editEvidence left behind
Grok 4.5Squad Survivor2,53423:36README written at 23:05, before most of the code. Smoke test passes with 0 kills. No screenshot.
Gemini 3.6 FlashSquad Survivors: Overworld RPG3,40623:59The only run with real scenes (14). No README; scratch test scripts left behind. One screenshot, at second 1.
GPT-5.6 SolCovenant of Ash1,25900:14One 1,137-line main.gd. Trees, rocks and projectiles drawn in code. README and 3 probe tests.
Fable 5Fable Squad1,93100:16Weapon levels 1–9, a debug harness, unit tests passing. 4 rounds of my feedback.
Opus 5Warband Survivors3,47700:36The largest harness: bot play, JSON reports, 169 unit checks. Ten enemy types.
Sorted by when each run stopped. Harness is known for Fable (Claude Code) and Sol (Codex, inferred); for the other three I didn't find session logs. Line counts exclude copied assets.
Warband Survivors: four named heroes fighting a dense, varied horde at 2:07
Opus 5 · Warband Survivors2:07 in. Slain 51, weapon II, horde 65. Its own screenshot.
Covenant of Ash: dark teal field with tiny sprites and fans of coloured orb projectiles
GPT-5.6 Sol · Covenant of AshIts --capture output, with --power applied.
Squad Survivors: four heroes at timer 00:01, level-up bar, no enemies yet
Gemini 3.6 Flash · Overworld RPGIts only image: timer 00:01, no enemies yet.
Squad Survivor: party of four firing arrows and spells across a grass field at 1:33, kills 0
Grok 4.5 · Squad SurvivorRe-captured No image existed. Recorded Oct 5 from an isolated copy with the game's own --auto-run bot, clock forced to 1:30.
Fable Squad at 10:20: weapon level 9, 591 kills, a radial storm of arrows and a MAX POWER toast
Fable 5 · Fable Squad10:20, weapon level 9, 591 kills. From its session transcript.
Fable Squad: THE PARTY HAS FALLEN, survived 06:49, kills 163, screen buried in enemies
Fable 5 · the standing testThe max-rank-standing-still test I asked for. Buried at 6:49.

The Grok re-capture tells you something on its own: the bot's party was wiped about five seconds after the clock was forced forward, with 1 kill. That matches its logged smoke test (0 kills in 3.4 simulated seconds). Combat was never demonstrated by Grok's own harness. Opus's bot, by contrast, logged 954 kills over 248 simulated seconds on seed 7 before the squad was wiped.

One run got all the feedback

The only feedback on record went to Fable, because its transcript is the only build session I found. It's also a clear look at what I was asking these games to be:

“pretty good. I think there's needs to be even more upside and pomp and ceremony. It's a game, not a chase boredom simulator. EVEN MORE. Perhaps larger projectiles at some point.”

Me to Fable 5 · 23:56

Earlier I'd reported an exit-menu bug and a map that felt too small, and asked for each party member to fall at every 25% of health lost. Later: “I find that I'm often 1 v 70 and still plinking like 2 arrows”, and finally “Do a test of max rank player standing still. Nothing can touch them basically.” The before-and-after shows what four rounds of that bought:

Fable Squad early build at 1:24: weapon level 1, 42 kills, sparse field Fable Squad final build at 10:20: weapon level 9, 591 kills, screen full of projectiles
23:15 · first build00:03 · after feedback

Did the others get the same notes? GPT's README (00:13) implements “Each 25% vitality band represents one hero”, which is the rule I gave Fable at 23:34, so I may have relayed some of it (inferred). Without their transcripts I can't show it, so the comparison isn't fair to the others.

Where it stopped. Between 00:42 and 00:53 every game was launched once more, in order: Grok, GPT, Gemini, Fable, Opus. That looks like me play-testing all five (inferred). Fable's log from that run shows 454 non-fatal physics errors. No verdict was written down. Twelve hours later the next experiment started.

ReviewYou played all five between 00:42 and 00:53. Which one held up? And did Opus, Gemini or Grok get feedback in sessions I couldn't find?
04 — Winter room · August 12–13

Snow, a door, a flashlight. Then sub-agents.

C:\git\3DGodot 4.7 Forward+ · primitives only3 stages · 13 entriesGit only in 2 Round 2 folders

The next day started with a question I'd pasted to ChatGPT about why Godot's default scenes look worse than Unreal's: “Do you need to use the Godot IDE in order to easily hit a quality level similar to Unreal?” The answer, a recipe for an opinionated lighting environment, went into every folder as godot scene.md. Then the experiment ran in three stages:

The Round 1 prompt, verbatim

Clear your scene of objects and construct the following: 1: a single room with a window, a single small table. a rectangular cellphone like object on the table. Outside, snow falling. 2: Accumulation or pseudo accumulation of snow outside. 3: Snow throughout the world (for future stage). 4: Add a door to your wintery scene which opens inward. Allow the player to walk outside. … 5: Render a flashlight in the player's hand a UI element indicating a flashlight equiped. 6: Add a button to toggle dusk and then night (with stars out). "F" key to toggle player flashlight on and off.

Aug 12–13, by stage (EDT)

First to last file written per stage. Round 2 runs went in sequence and got shorter: the last two started at 00:37.

View as table
Luna Round 1: dim room, table with a phone, snow in the window
GPT-5.6 Luna (max)Items 1–3 only: no door, flashlight or night.
Sol-High Round 1: interior at night, flashlight on, stars and moon through window
GPT-5.6 Sol (high)All six items. The door opens by itself as you approach.
Sol-High Round 1 exterior: dark blue box cabin on a flat snowfield, door open
GPT-5.6 Sol (high)Its exterior: a box on a plane.
Sol-Max: orange-lit interior with door, wide window onto snow at dusk, flashlight
GPT-5.6 Sol (max)A separate late run, 22:09–22:41, with all six items at once.
Grok Round 1: white plaster room, window with big square snowflakes, two bright light orbs, huge flashlight bloom
Grok 4.5All six items in the code; no README.
Foundation Yard: grey tiled plane with a few coloured blocks under a flat grey sky
The Round 2 baseline. Sol-Medium's stage-A scene, byte for byte.

Fable, Opus and Gemini left no Round 1 screenshots. Sol-Medium never got past stage A, and its scene became the shared starting point for Round 2.

Round 2: “notoriously lacking in creativity”

Round 2 made the lead model the director of other models. The goal file is blunt about who I thought was good at what. These capability notes are my own, written before the round, not verdicts on it:

“You will make use of sub-agents … which are notoriously lacking in creativity and creating visually interesting things. The base project provided to you is intentionally bland and constructed by one of these sub-agents.”

Round 2 goal.md · Fable, Opus, Gemini and Grok versions

That framing doesn't work when Sol is the lead, so at 23:40 I rewrote the goal for Sol's run, flipping the roles: Sol as lead engineer, Gemini as the creative sub-agent, “Gemini invents → Sol engineers → Gemini art-directs → Sol polishes.” In the end there were four goal variants. Gemini and Grok, who started last, got the shortest one: Codex-only sub-agents, no controls section, no launcher requirement.

Code size, Round 1 → Round 2

GDScript plus shader lines. Round 1 scenes were partly hand-authored scene files; Round 2 built almost everything in code, so this overstates growth slightly.

View as table

Despite four different goals and no shared art direction, all five Round 2 entries arrived at the same picture: a warm timber cabin in a cold snowfield at night. They differ in what they added. Fable's has chimney smoke and an aurora, Sol's has procedural audio and a phone that opens an emergency broadcast, and Opus's has three times of day built around dusk.

Round 2 starting point: the bland Foundation Yard Fable 5's Round 2 HEARTHLIGHT: timber cabin with lit window, chimney smoke, stars and aurora
BaselineFable 5 · Round 2
HEARTHLIGHT: open cabin door casting amber light onto blue snow, aurora above
Fable 5 · HEARTHLIGHT1 h 45 m. The only run with delegation on disk.
LAST LIGHT: dark timber cabin at night, flashlight on the window, snow drifts
Opus 5 · LAST LIGHTRe-captured 55 m, 4,155 lines. Its 8 screenshots were gitignored and lost; recaptured Oct 5 with its own --capture mode from an isolated copy.
BLIZZARD PROTOCOL: cabin with porch, low-poly pines, fence posts, flashlight beam on snow
GPT-5.6 Sol · BLIZZARD PROTOCOL50 m. These shots predate its last 32 minutes of code.
THRESHOLD: tan box cabin with flat snowy roof and open door, pines, falling snow
Grok 4.5 · THRESHOLD15 m, the shortest run. No launch scripts.
Gemini Round 2: wood-plank interior, phone on table, two-pane window with snow, flashlight
Gemini 3.6 Flash31 m. Never renamed from “Foundation Yard”; two near-identical interior shots.
LAST LIGHT dusk interior: warm walls, hanging lamp, window, table and stool
Opus 5 · dusk interiorRe-captured “the picture the whole scene is built around,” per its README.

Did anyone actually delegate?

The folders can only prove it for one run. Fable's Round 2 repo has ten commits, each crediting the model that wrote it, plus job logs for every hand-off:

Opus and Sol ran in GitHub Copilot. Their session metadata exists (Opus: 31.3 M input tokens) but I couldn't read the transcripts, and neither folder credits a sub-agent. Gemini and Grok left no trace of delegation. Whether sub-agents helped those four can't be shown from the files.

Where it stopped. Round 2 covered five of the eight Round 1 entrants. Sol-Medium, Sol-Max and Luna never got one. There's no scoring sheet. Two small signs suggest I played some of them: Sol's Copilot session is titled “Soften scene audio”, and Fable's last commit (01:02) is “fix: single ESC no longer quits” (both inferred). The last file was written at 01:08. That evening, the first SteamShip files appeared.

ReviewWhich Round 2 cabin did you like, and did the sub-agent framing seem to matter? Also: Sol's “Soften scene audio” session and Fable's ESC fix read like playtest feedback. Is that right?
05 — SteamShip · August 13 – September 4

The bakeoff that drifted, and the rematch

SteamshipExperimentsGodot 4.7.1 + C# MonoGameRound 1: Aug 13–17 · Round 2: Sep 4Git in most folders

SteamShip started from one concept image: a cutaway paddle-steamer with eleven labelled rooms over three decks, a crew roster and HULL / STEAM / COAL gauges. My design direction, from a session on Aug 16: “Consider that it's a cross between FTL and Darkest Dungeon, with open-world ship sailing, ports, different factions.”

Concept art: cutaway paddle steamer with labelled rooms including bridge, galley, gun deck, boiler room and engine room; crew list and resource gauges
concept.png, the shared reference for every run. The brief asked for 5–8 rooms, at least one below the waterline, gentle bobbing, animated crew, click-to-move between rooms, and two cannon placeholders.

Round 1: orchestrators, quotas and scope drift

Round 1 (Aug 14) was an orchestration bakeoff. A shared AGENTS.md ranked the available models by benchmark score and set delegation rules. Orchestrators in Claude Code, Codex and GitHub Copilot were told to act “as the Creative Director and Lead Programmer” and use their sub-agents. Claude Code reached Sol, Grok and Gemini through PowerShell wrappers.

Several runs grew far past the slice. Fable 5's run went on for three days and 92 commits, into ship variants, storms, an asset-override system, and finally a gated plan for four in-game editors with two AI review leads. Its last commit, Aug 17 at 15:24, reads “Gate 4 report: both leads GATE-READY; awaiting user approval.” The approval never came (inferred). One Copilot run kept going after the experiment as its own repository, with 774 commits by Sept 3.

Round 1 Fable build: River Serpent steamboat cutaway with crew roster, rooms, ladders and a paddle wheel
Fable 5 · Round 1A gate-review baseline from the Fable run's docs/gates/.

The quota problem was on record the whole way:

  • Aug 15: “We have 24% Fable (you) usage left so, make use of delegation!”
  • Aug 16: “continue we were disconnected due to quota limits”
  • Aug 17: “I had to switch to Opus midway through your session due to quota expiring.”

On Sept 4 at 19:02, during the rematch, I moved the whole round into a folder named Old_failed. No file says why each run counted as failed.

Round 2: the same five prompts, one evening

The rematch on Sept 4 was much tighter, and it is the clearest precursor to the jams: one brief, then a chain of identical follow-up prompts, each closed by my review and a commit. The slice-1 brief gained one line: “Because you are a Frontier Model, you must additionally build the prototype in C# using Monogame or equivalent … Consider this a capability demonstration.” So every step had to land in both Godot and MonoGame.

  1. Slice 1: the navigable ship.
  2. Texture bake: “all "assets" should be textures/etc and not procedurally generated.”
  3. MVP 1: swap the 2D crew for supplied 3D models (keeping a 2D look) and add a day/night cycle.
  4. MVP 2: an enemy encounter button, cannon fire, and enemy crew boarding. MonoGame only.
  5. Set Sail: a sailing scene with three views on keys 1/2/3, and an interception that returns to the cutaway.

Sept 4, prompt by prompt (EDT)

From my prompt to the commit. Fable 5.1 started at MVP 2 on a copy of the Codex run's Phase 1 code (21:36 mark). Gemini did slice 1 only.

View as table
S.S. Marigold at night under attack: Battle Stations, Rustwing contact inset, raiders boarding the lower deck
S.S. Marigold

GPT-6 Astra · Codex

Steps
5 of 5, from scratch
Commits
6
C# / GDScript
1,687 / 743 lines
Sub-agents
0 (offered, never called)

My one complaint: “regression. Much of your white text in the dotnet version (and possibly godot) is unreadable.” Fixed in 10 minutes.

Brass Tide encounter: Cinder Raider alongside with boarders on the gangplank, crew portraits and objectives panel
Brass Tide

GPT-6 Astra · Copilot

Steps
5 of 5, from scratch
Commits
6 + 58 auto-checkpoints
C# / GDScript
5,403 / 2,237 lines
Sub-agents
6, all Astra

The largest codebase, and the only Round 2 run that delegated. Copilot refused to start before the run; a “strange horizontal scan line” needed a fix.

S.S. Marigold by day with the Cormorant steam raider alongside, ENEMY SINKING, a red boarder in the wheelhouse
S.S. Marigold (inherited)

Fable 5.1 · Claude Code

Steps
2 of 5 (MVP 2, Set Sail)
Commits
2
C# / GDScript
2,278 / 743 (Godot inherited)
Sub-agents
0

Started at 21:36 from a copy of the Codex folder at its Phase 1 state, so it is not an independent entry. Its own report lists what it didn't do: “no Godot passage, no audio … no manual mouse or controller playtest.”

SteamShip.MonoGame

Gemini 3.8 Flash · Antigravity

Steps
1 of 5
Commits
0 (nothing committed)
C# / GDScript
1,641 / 1,744 lines
Run time
~8 minutes, 239 steps

Built 11 rooms (the concept's count; the brief asked for 5–8) in both engines, reported 8 .NET tests passing, and captured no screenshots. It got no follow-up prompt, and nothing records why.

Marigold chase view: stylised steamer from behind on open water, enemy sail on the horizon
Codex · chase view“Follow the Marigold.”
Brass Tide nautical chart: islands, dotted course and a raider intercept
Copilot · chart view“The Brass Reach.”
Marigold at the helm: wheel, compass binnacle, engine telegraph, passage log
Fable 5.1 · helm viewWheel, binnacle and telegraph.

My verdicts were per-step commit approvals, almost all “excellent”. The two Astra runs were the only ones to finish all five steps from scratch, and both got the same words for MVP 2, typed one minute apart:

“pretty pretty good. commit please.”

Me, to both GPT-6 Astra runs · 21:52

Where it stopped. The prompts chain forward with no defined end, and they stop after Set Sail. The last commit is Fable 5.1's, at 22:54. The comparison was never written: one run did one step, one started halfway on another run's code, and Godot parity was dropped after MVP 1. The next evening, Sept 5, Experiment 01 began with a fixed scenario list and a written final briefing.

ReviewWhy did Gemini stop after slice 1, and why did Fable 5.1 start from the Codex copy (late start, quota, or deliberate)? And what made Round 1 “failed” in your eyes: the drift, the quotas, or the results?
06 — Looking back

What carried into the jams

These are patterns in the record, not conclusions. Every experiment here had one reviewer, one attempt per run, and prompts that changed mid-experiment.

Identical prompts, on purpose

The winter room ended with four goal variants, and the 2D jam's feedback went mostly to one run. Both make the results hard to compare. Experiments 01 and 02 sent every run the same prompt, word for word.

Gates came from SteamShip

Round 2 on Sept 4 was already a chain of prompts closed by review and commit. Two days later, Experiment 02 formalised that into milestones M0–M5 (inferred lineage).

The first bug is physical

Star Commander runs fell through the terrain, 2D runs had exit-menu bugs, and the winter-room ESC quit too early. The same pattern shows up in the jams: the blocking bugs were the ones you only find by playing.

Delegation is hard to see

When asked to delegate in the winter room, only Fable left proof on disk. When nobody asked, in SteamShip Round 2, only Copilot's Astra run spawned sub-agents anyway. Codex was offered the tools and never used them.

Agents that look check better

Luna's screenshot habit and The Violet Hour's autopilot harness belong to the runs I praised or that proved themselves end to end. Runs without self-captures left the bugs for me to find.

Quota is part of the result

SteamShip Round 1 spent days rationing Fable's usage and swapped in Opus mid-session. Availability shaped these runs as much as ability, as it did in Experiment 01.

How this page was built
  • Timelines: file modification times and Git logs; Codex, Copilot and Antigravity session metadata for start times.
  • Quotes: my prompts, verbatim, from Claude Code history, Codex rollouts, Copilot event logs and Antigravity databases. Agent and pasted-AI text is labelled as such.
  • Line counts: GDScript, shaders and C#, excluding engines, addons and copied assets.
  • Images: the runs' own screenshots, resized. Two runs are marked re-captured: Opus's winter cabin (with its own --capture mode) and Grok's 2D game (with Godot's movie writer), both from isolated copies on Oct 5. The DEMIURGE FM clip was assembled from the run's own recorded frames.
  • Not used: Copilot transcripts for the winter-room Round 2 (not readable), and build transcripts for four of the five 2D runs (not found).