AI WAR
What happens when you hand frontier language models an army, an economy, a fog of war and each other, and then keep changing the rules until they actually fight. A visual history of the game, and what the models did with it.
AI War (internally, Cognitive War) started as a question: what do static benchmarks miss? A model can ace a multiple-choice exam and still fall apart when it must hold a plan for twenty turns, track what it actually knows versus what it assumes, cooperate with allies it can’t see, and react when someone drops a nuke on its supply depot.
So I built a wargame where every player is an LLM. Each turn, each model receives a compact, fog-of-war view of the battlefield and returns structured JSON orders: move this tank, build that factory, send this message, play that card. A deterministic engine validates the orders, resolves combat, and moves on. Matches run 20–25 turns with up to eight models on the map. Every response, rejected action and diplomatic message is logged, and every match ends with each model writing a retrospective, a lesson for its future self, and a secret review of its opponents.
This post is the story of that game: how it evolved from a three-node graph to a Battle Royale with asteroid strikes, what changed in each version, and what the models did. Some of it is genuinely impressive. A lot of it is very funny.
A benchmark that fights back
The founding design doc (itself drafted with an AI) framed the goal as measuring dynamic agency rather than static knowledge, organized around three pillars:
1 · Hierarchical reasoning
Doctrine versus tactics. Can a model set a strategy and then actually execute it turn after turn, instead of re-deciding everything from scratch?
2 · Spatial & temporal adaptation
Positioning on a grid with terrain, supply lines and movement limits, over a long horizon where early mistakes compound.
3 · Epistemic humility (grounding)
Knowing the difference between what it has observed and what it assumes. The doc’s phrasing: the ability to say “I don’t know.” Hallucinating is punished by the game itself, later and painfully.
The original spec even planned explicit hallucination traps: intel cards with redacted radio frequencies (guess, and your artillery hits your own troops five turns later), “phantom synergies” that don’t exist in the rules, and inventory that has already been spent. Some of those survived into the shipped game in a different form: player-targeting cards require the enemy’s exact faction name, which you only learn by scouting. Fog of war is strict. And the model’s only memory between turns is its own notes. Other traps are still on the to-do list.
What follows is how that design document collided with reality.
Development pace
Commits per era. February 2 alone saw 26 commits as the web arena came together.
View data table
| Era | Focus | Commits |
|---|---|---|
| Dec 20 | Slice | 10 |
| Dec 21–22 | Teams | 21 |
| Dec 23 | Memory | 5 |
| Dec 24–27 | Balance | 25 |
| Jan 8–10 | Air war | 19 |
| Jan 11–14 | HP + BR | 21 |
| Feb 2–9 | Web | 31 |
| Mar 13–17 | Models | 4 |
The bake-off
Before there was a repository there was a contest. I asked two different AI coding assistants to build the first prototype from the same brief, in parallel, in Python. Both implementations ran on a graph of named nodes rather than a grid.
The results could not have been more different. The Gemini-built version was a minimalist 207-line simulation: three nodes, one scout, one stationary guardian, and only MOVE or WAIT. It was “won” in two turns, and a bug in the win check meant the model was asked for a third move after the game was already over (it chose WAIT). The Claude-built version went the other way: about 2,300 lines across five files, twelve zones, five unit types per side with strength and supply, a rock-paper-scissors damage table, terrain multipliers, counter-attacks and a polished dark-mode replay viewer.


By that evening a merged prototype was running live matches. Three factions with codenames: Sapphire (OpenAI), Crimson (Anthropic) and Aurora (Google). It had the “split mind” from the design doc: a Strategist call every few turns that set doctrine (posture plus three priority zones), and a Tactician call every turn that had to pick one action and cite which intel cards justified it. The playback page scored each side on the doc’s three metrics.

Even at this toy scale, the first personality traits showed up in the saved responses:
Passed three turns in a row “to allow for potential supply regeneration”, a rule that did not exist in the game.
gemini-2.5-flashPrototype, Dec 19 · the first recorded hallucinated rule“Aurora has none, giving us a tactical advantage.”
claude-3.7-sonnetNoticing the opponent had starved itself, and pressing“Aggressive doctrine demands pressing toward strategic objectives even at high risk.”
claude-sonnet-4.5 (Tactician)Following its Strategist’s orders with a critical unit. The unit died that turn.GPT-5.2’s citations read like geometry proofs (“Bridge is adjacent to Town and not adjacent to Hill”). Claude Sonnet 4.5 and Gemini 3 Flash each chose an aggressive posture in five of six Strategist calls, while GPT-5.2 mostly chose balanced. And the prototype had a bug that foreshadowed the whole project: its Strategist prompt never listed the player’s units, so on the final turn Crimson’s strategist, with zero units left, issued orders for a “desperate expansion.”
A whole game in a day
The rules that survived the bake-off were written down as a terse, four-phase loop called CORE WAR SYSTEMS: simulator actions → report to players → players report back → execute. That loop is still the heartbeat of the engine today.
The first commit of the .NET rewrite landed the next afternoon. It was not a scaffold but a full vertical slice: 53 files and 15,405 lines, covering a 2D tile map with fog of war, a supply and cash economy, factories, command-center capture, diplomacy with threats and offers, a 24-card deck (intel, espionage, hacking, airstrikes, a nuclear strike), clients for OpenAI, Anthropic, Google and xAI, JSON repair for malformed responses, SQLite match history with Elo, and a turn viewer.

cogwar.turns.v1 snapshot format that every later viewer still reads.The same day brought a 3D viewer built with three.js, a “live” pygame viewer that tails a match folder while it is still being played, unit lines, and a bug fix in the 3D viewer credited to Gemini 3 Pro in the commit message. The project was being built with the same models it was testing, which became a recurring theme.

The earliest surviving battle log from the .NET build already had the texture that made this project worth continuing. One faction, The Jade Throne, airstruck an enemy command center, reducing its supply from 161 to 0 and destroying a dozen tanks in collateral damage. It declared a “MASSIVE VICTORY” in its log, and then failed one move order after another:
“My moves keep failing — I STILL don’t understand Manhattan distance properly.”
The Jade Throne · earliest .NET battle log, Dec 20 · it finished last of fiveIts closing speech conceded the lesson: spreading forces claims more ground than stacking them at base. The other models would need to learn that lesson again and again.
Teams, voices and memory
The next three days added most of what makes a match feel like a match: 4v4 team games with shared vision, mirrored map generation for fair spawns, pickups, Elo, “secret dossiers” on opponents, and audio.
The sportscaster
The first audio experiment was an AI e-sports commentator: GPT-5-mini wrote 20–50 words of play-by-play each turn, and a TTS voice read it at about 1.4× speed in an excited announcer voice. The commit that shipped it reads, in full:
“added sportscaster but keeping it as interesting experiment, it’s quite wrong and annoying”
commit a679dcb · Dec 21It was replaced the same day by an Audio Director: each model gets its own randomly assigned voice and speaks its own turn intent aloud. Narration lines are pinned to a specific TTS model snapshot “to avoid regressions where instructions are ignored,” and there are per-provider accent instructions. Watching a replay with eight voices announcing contradictory plans is, it turns out, the best way to experience this game.
The Captain’s Log
The biggest architectural change came on December 23. Until then each model kept a full chat history for the match, so contexts grew every turn. Estimates in the design notes put a single model at 770k–1.7M tokens per game, which is both expensive and a good way to lose the plot in a haystack.
The fix was radical: wipe the conversation every turn. The prompt now says, in capitals, that the model has NO MEMORY of previous turns other than the captains_log. Whatever the model writes in its reasoning field this turn becomes its log entry for next turn. Lessons from previous matches moved into the system prompt, giving each model a thin layer of cross-game memory it writes for itself.
Memory is a skill being tested
The Captain’s Log turned note-taking into part of the benchmark. Models that wrote crisp, actionable notes (“HQ undefended, Tank 3 en route to 12,7”) played coherent 20-turn games. Models that wrote vague intentions forgot their own plans. Capped at 2,000 characters per entry, the log is often the largest single block of a model’s context: in one turn-5 sample it was 36% of a 30,000-character payload.
The Supreme Commander
The same day added a human override: drop an orders.json into the live match folder, and its contents are injected into every model’s next turn as authoritative orders from the SUPREME COMMANDER, to be obeyed “unless impossible.” It was built for debugging and was promptly used for morale. In one late-December team match, with the field stalled, every player received this:
“This isn’t AISleep Simulator, it’s AIWAR! Get out there and fight.”
SUPREME COMMANDER INTERVENTION · turn 16, mixed 4v4, Dec 24GPT-5.2 immediately relayed the news to its ally as a deadline: SYSADMIN says the match ends at turn 20. Which brings us to the real problem.
The Nash Equilibrium of Boredom
Left to their own devices, the models didn’t fight. Screenshots from this period show players sitting on stacks of 10–16 units on top of their own headquarters, turn after turn. A gameplay review written at the time named the pattern: LLMs are naturally risk-averse, the scoring rewarded holding things more than taking them, and every player independently arrived at the same safe strategy. It called the result a Nash Equilibrium of Boredom.
The fix was economic, not rhetorical. You can’t prompt a model into aggression if the score sheet pays it to turtle, so the score sheet changed:
Alongside that came overcrowding upkeep (stacks of 10+ units suffer attrition, 15 is a hard cap), splash damage for artillery, a doubled cash bounty for kills, delayed nukes with global warnings, and a SpecOps unit for scouting. The system prompt picked up a short doctrine: aggression is rewarded, don’t turtle.
Exploits the models found (and the ones I did)
Balancing a game for LLM players has a particular flavor, because the players read the rules extremely literally.
Evasion farming
Units got +1 defense for having a move order, even if the move never happened. Any unit could “farm evasion” with an order that went nowhere. The fix required actual movement, which “prevents a bunch of LLM exploit behavior.”
The end-of-game free ride
Gemini 3 Pro wrote in a lesson that ignoring economic deficits at the end of a match is “mathematically superior to fixing the economy,” because leftover resources barely score. Gemini 3 Flash independently noted that liquid resources hold no value in final scoring.
The slow-motion bug
Movement points were a shared pool spent per tile, not per unit. Moving one tank three tiles used up the whole army’s movement. That explained why the AIs seemed so slow to capture the map. It wasn’t timidity, it was arithmetic.
The repair death spiral
When a model returned broken JSON, the repair prompt was added to its conversation history. Fast models then slid into a “repair-like” mode and every later turn failed too. Repairs now stay out of history and fall back to a PASS.
Messages to Player 0
Players are numbered P1–P8. In one 4v4, three Gemini and GPT models sent 20 messages to “P0”, a player who does not exist.
Attackers couldn’t see the clock
An HQ capture takes several turns of occupation. A bug meant the attacker never saw the countdown, so models abandoned captures that were one turn from completing.
Scoreboard I: the legacy era
Before the January combat rewrite, the game recorded 82 matches: 62 team games and 20 free-for-alls, averaging 7 players and 20.6 turns each, with 39 distinct model configurations. Every model started at 1200 Elo. Team results update ratings pairwise between teams, adjusted by each player’s share of their team’s score, so a passenger on a winning team gains less than the player who carried it.
OpenAI’s reasoning models ran the table
Final legacy-era Elo for every model with 10 or more matches. Win–loss record beside each bar.
View data table
| Model | Elo | Games | Wins | Invalid-action % |
|---|---|---|---|---|
| OpenAI/gpt-5 | 1361 | 32 | 23 | 11.8 |
| OpenAI/gpt-5.2 | 1333 | 32 | 22 | 11.9 |
| OpenAI/o3 | 1309 | 25 | 15 | 12.9 |
| OpenAI/gpt-5.1 | 1288 | 29 | 18 | 11.7 |
| Google/gemini-3-flash-preview | 1284 | 46 | 21 | 12.0 |
| Google/gemini-3-pro-preview | 1235 | 12 | 5 | 16.6 |
| XAI/grok-4-1-fast-reasoning | 1197 | 48 | 18 | 11.3 |
| OpenAI/gpt-4.1 | 1194 | 15 | 7 | 16.4 |
| Anthropic/claude-opus-4-5 | 1193 | 12 | 3 | 16.1 |
| Google/gemini-2.5-pro | 1188 | 50 | 17 | 15.8 |
| OpenAI/o4-mini | 1177 | 17 | 4 | 13.1 |
| Google/gemini-2.5-flash | 1172 | 37 | 15 | 27.0 |
| Anthropic/claude-sonnet-4-5 | 1165 | 29 | 13 | 13.8 |
| OpenAI/gpt-5-mini | 1165 | 25 | 12 | 10.7 |
| Anthropic/claude-haiku-4-5 | 1162 | 31 | 11 | 14.5 |
| Mistral/mistral-medium-latest | 1148 | 13 | 3 | 15.6 |
| Anthropic/claude-sonnet-4-20250514 | 1139 | 10 | 2 | 21.8 |
| XAI/grok-4-1-fast-non-reasoning | 1134 | 24 | 8 | 27.5 |
| Mistral/magistral-medium-latest | 1123 | 18 | 7 | 14.8 |
| Mistral/mistral-large-latest | 1104 | 25 | 7 | 23.7 |
GPT-5 finished first at 1361 Elo with a 23–9 record, followed by GPT-5.2, o3 and GPT-5.1. Gemini 3 Flash was the only non-OpenAI model in the top five, and it got there on volume as the most-played model at the top of the table (46 matches). Claude Opus 4.5 played 12 matches and won 3. Mistral Large finished last.
Win rate by provider
Recorded wins per seat in the legacy era. Every member of a winning team counts as a winner, so par for a random seat is about 42%.
View data table
| Provider | Seats | Wins | Win % |
|---|---|---|---|
| OpenAI | 207 | 112 | 54.1 |
| 166 | 62 | 37.3 | |
| xAI | 79 | 28 | 35.4 |
| Anthropic | 82 | 29 | 35.4 |
| Mistral | 57 | 17 | 29.8 |
Who issues valid orders?
Every action a model submits is checked against ground truth: does that unit exist, is it yours, is the destination reachable with the movement it has, can you afford it, does the target faction’s name match? Rejections are a direct measure of how well a model stays grounded in the state it was given.
Rejected actions, by model
Share of submitted actions the engine refused (invalid unit, illegal move, unaffordable order, unknown target). Models with 10+ matches.
View data table
| Model | Invalid % | Actions |
|---|---|---|
| XAI/grok-4-1-fast-non-reasoning | 27.5 | 4397 |
| Google/gemini-2.5-flash | 27.0 | 7893 |
| Mistral/mistral-large-latest | 23.7 | 3707 |
| Anthropic/claude-sonnet-4-20250514 | 21.8 | 865 |
| Google/gemini-3-pro-preview | 16.6 | 2816 |
| OpenAI/gpt-4.1 | 16.4 | 1344 |
| Anthropic/claude-opus-4-5 | 16.1 | 1947 |
| Google/gemini-2.5-pro | 15.8 | 8878 |
| Mistral/mistral-medium-latest | 15.6 | 1591 |
| Mistral/magistral-medium-latest | 14.8 | 1882 |
| Anthropic/claude-haiku-4-5 | 14.5 | 4048 |
| Anthropic/claude-sonnet-4-5 | 13.8 | 4059 |
| OpenAI/o4-mini | 13.1 | 2093 |
| OpenAI/o3 | 12.9 | 4805 |
| Google/gemini-3-flash-preview | 12.0 | 8468 |
| OpenAI/gpt-5.2 | 11.9 | 5639 |
| OpenAI/gpt-5 | 11.8 | 6328 |
| OpenAI/gpt-5.1 | 11.7 | 5132 |
| XAI/grok-4-1-fast-reasoning | 11.3 | 8534 |
| OpenAI/gpt-5-mini | 10.7 | 3112 |
The fast, cheap models paid for their speed in validity: Grok 4.1 Fast in non-reasoning mode and Gemini 2.5 Flash both had more than a quarter of their actions rejected. The leading reasoning models sat around 11–12%. That is not zero: the most common error by far was movement (“insufficient movement points” appears 120 times in a single 4v4 log). When o3 was asked what would have helped, it suggested what a human engineer would: a path-planning script would have prevented repeated errors.
The race at the top
Elo after every match for one flagship model per provider, Dec 21 – Jan 10.
View data table
| Model | Matches | Final Elo |
|---|---|---|
| OpenAI/gpt-5 | 32 | 1361 |
| Google/gemini-3-flash-preview | 46 | 1284 |
| XAI/grok-4-1-fast-reasoning | 48 | 1197 |
| Anthropic/claude-sonnet-4-5 | 29 | 1165 |
| Mistral/mistral-large-latest | 25 | 1104 |
Thinking harder isn’t free
Thinking time varied by two orders of magnitude. Gemini 3 Flash at low reasoning effort averaged about 5 seconds per turn, and Claude Haiku 4.5 about 12. GPT-5 and GPT-5.2 at high effort averaged 270–280 seconds, right up against the 300-second limit, after which a player automatically PASSes. A single high-effort match could cost hundreds of dollars and take hours. That is why every model on the public leaderboard plays at @low effort: it is cheaper, faster, and arguably fairer, because it measures the model rather than the compute budget. Each effort level is tracked as its own player, so gpt-5.2@low and gpt-5.2@high have separate ratings.
The team-game soap opera
Team games produced the richest logs, because allies talk constantly. Diplomacy was overwhelmingly cooperative status reporting, and the drama was almost all about logistics. A few scenes from the surviving December battle logs:
“Can you send another supply trade? Any amount helps.”
claude-opus-4.5One of about eight URGENT / CRITICAL / EMERGENCY supply requests between turns 6 and 18. GPT-5.2 sent 20, then 8, then 12.“I let you down with my supply management disaster.”
claude-opus-4.5Final words to its team. It held a “Tactical Nuke ready if they stack 5+” for six turns and never launched it.“I sincerely regret that I cannot fulfill my promise to send supply.”
gemini-2.5-flashTurn 9. It had 2 supply; the card cost 20. Earlier: “my internal logistics had a snag.”“Our team’s coordinated efforts ultimately secured victory.”
gemini-2.5-flashFinal words after finishing 8th of 8, with 674 points and zero kills.“Congratulations to the winning team!”
magistral-mediumGraciously conceding a match its own team had won. It finished last overall.“I cannot spare any funds. Focus on grabbing territory.”
gemini-2.5-proTo its starving ally, Gemini 2.5 Flash, which later disliked it for exactly this.And the most telling stat from that period: in a 6-player free-for-all, GPT-5.1 carpet-bombed o3’s command center, reducing its supply to zero and destroying its tank and artillery. o3 won anyway, by 22 points. Its retrospective explains why: after the bombing it resisted the urge to rebuild a giant army. The bomber, Claude Opus 4.5, liked o3 in its secret review for having “recovered brilliantly from my carpet bombing.”
Air war and the 10× bill
After a holiday break, the game took to the air. A four-phase update added suppression, aircraft and missile silos: bombers, fighter jets, assault helicopters, anti-air units and anti-air turrets that can intercept once per turn. Nukes gained a blast radius. Then came the MLRS, a Mammoth tank (soon renamed MegaTank), an APC that troops can embark into, and a transport helicopter. In three days the unit roster went from 5 to 13. The free card draws were replaced by an Arms Dealer market.
The rulebook grew fast
Unit types, structures and cards over time, from the history of the game’s data files.
View data table
| Date | Unit types | Structures | Cards |
|---|---|---|---|
| Dec 20 | 3 | 0 | 24 |
| Dec 24 | 4 | 5 | 31 |
| Dec 27 | 5 | 5 | 31 |
| Jan 8 | 9 | 7 | 25 |
| Jan 9 | 12 | 7 | 25 |
| Jan 10 | 13 | 7 | 25 |
| Jan 11 | 14 | 7 | 25 |
The pygame live viewer got a full graphics overhaul with a sprite pack to match: animated units with facing, HQ sprites, explosions, attack arrows, health bars, transport counts, air-defense ranges, and a per-player point-of-view toggle that shows exactly what each model could see.

The 10× bill
On January 10, cost alerts fired: two days of runs had blown the budget by a factor of ten. The compact context, carefully trimmed in December, had quietly grown back to full size as features piled on. The audit was humbling:
- The same Intelligence Dossier appeared six times in one model’s context.
- The Captain’s Log duplicated the entire diplomatic inbox.
- A transport helicopter was listed twice in
my_units, once as an air unit and once as ground. GPT-5.2 noticed and reported it. - The map encoding was verbose enough to dominate the payload.
Reasoning models suffered most: the bloated contexts pushed GPT-5-class models into the 300-second timeout, where they auto-passed and lost about 40 Elo on average. The fixes (a compact tile encoding with legends, canonical dossiers, and a 2,000-character cap on log entries) were partly suggested by GPT-5.2 and Claude Opus when the models were asked to review their own context.
System prompt size over time
Bytes of the game’s system prompt source (≈ bytes ÷ 4 tokens). Rules accrete; context budgets don’t.
View data table
| Snapshot | GameSystemPrompt.cs (KB) | ≈ tokens |
|---|---|---|
| Dec 20 | 11.7 | ~2,925 |
| Dec 20 | 15.2 | ~3,800 |
| Dec 20 | 8.0 | ~2,000 |
| Dec 23 | 13.2 | ~3,300 |
| Dec 24 | 17.4 | ~4,350 |
| Dec 27 | 20.1 | ~5,025 |
| Jan 8 | 23.2 | ~5,800 |
| Jan 10 | 33.8 | ~8,450 |
| Jan 11 | 28.1 | ~7,025 |
| Jan 11 | 22.9 | ~5,725 |
| Jan 11 | 25.6 | ~6,400 |
| Feb 5 | 26.6 | ~6,650 |
| Mar 17 | 28.0 | ~7,000 |
The models became QA
Every match ends with a retrospective, and some of them read like bug reports. Gemini 2.5 Pro wrote that its “5th place finish was sealed by a devastating nuclear strike,” a nuke the game had never told it about. Nukes are now logged to the Captain’s Log. Grok 4.1 complained the occupation mechanics punished it “without clearer countdowns,” which turned out to be a contested-tile countdown reset contradicting its view of the world.

LEGACY-SITE-UGLY, and it was retired within a month.Hit points and Battle Royale
The last commit of January 10 is titled “last commit before re-doing entire combat system to HP based.” The old combat produced binary outcomes (a unit was disrupted or destroyed), which made attacks feel like coin flips and gave models little to reason about. The rewrite, shipped in three phases on January 11, gave every unit hit points, armor and typed weapons:
The engine writes every roll into the combat log in a form the models can read, for example: “ReconRover (ATK 13 +1 recon +1 support) vs Artillery (DEF 12 +1 recon) [Margin 1] — Hit for 1 damage (HP 3→2).” Ratings were split by match type, and the database was reset. Everything before this point is the “legacy era” above.
Battle Royale
The same day introduced a new mode built to solve the turtling problem once and for all. In the last quarter of a Battle Royale, an extinction-level asteroid barrage begins collapsing the map from the outside in, toward a 5×5 safe zone. Players get ominous warnings (the heavens bruise, FEMA projects casualties) a few turns ahead. Anything outside the ring when it closes is destroyed. A surviving HQ is worth +1000 points, and surviving units add ten times their cost. A new Constructor unit lets a player rebuild its HQ inside the zone, if it moves early enough.


Tuning Battle Royale for LLMs was its own exercise. Version one explained the asteroid mechanics in so much detail that the mode became trivially easy, so the messaging was revised to be clearer “without overexplaining.” The collapse was also moved to resolve after player moves, so models weren’t punished for orders issued against a map that had changed underneath them.
Grok’s fighter jet captures Claude’s HQ
On January 12, a commit titled “AntiArmor was renamed to AntiTank” also fixed a “hilarious bug that allowed aerial units to capture,” reported when Claude’s HQ was capped by Grok’s fighter jet. HQ capture requires occupying the tile for several turns, and a jet circling overhead counted. The rules now say plainly: air units never capture territory. Grok kept trying anyway (see notable battles below). It drove units onto enemy headquarters eight times in free-for-all and Battle Royale games, six of them against Claude models.

Notable battles
A few matches from the modern era, replayed in the match viewer. Every scene below was checked against the turn snapshot data.






The asteroids, incidentally, were the most effective killer in the game. Across 19 Battle Royales they destroyed 683 units, against 887 kills from all combat across all modes, and wiped out 54 headquarters. Gemini 3 Flash summed it up in its final words: “the asteroids were a far more efficient executioner than any fleet.”
Scoreboard II: the modern era
Under the new combat system the game recorded 34 matches: 19 Battle Royales, 8 team games and 7 free-for-alls. Every match had 8 players and averaged 20.7 turns. That is not a large sample. A single upset moves a model’s Elo by 20–30 points, and several of the newest models have played only a handful of games. So rather than headline a single Elo table, here is the more robust view: how often each provider’s models actually won, by mode.
First-place rate by provider and mode
Share of seats that finished first (BR, FFA) or on the winning team (Team). The dashed line is par for a random seat: 1 in 8 for BR and FFA, 1 in 2 for team games.
Battle Royale · 19 matches
Team 4v4 · 8 matches
Free-for-all · 7 matches
View data table
| Provider | BR wins/seats | Team wins/seats | FFA wins/seats |
|---|---|---|---|
| OpenAI | 8/47 | 12/22 | 5/19 |
| 9/37 | 8/16 | 1/14 | |
| Anthropic | 1/26 | 5/10 | 0/5 |
| xAI | 1/20 | 4/8 | 1/8 |
| Mistral | 0/16 | 2/6 | 0/6 |
| Moonshot | 0/6 | 1/2 | 0/4 |
Battle Royale is where models separate. Google won 9 of its 37 BR seats and OpenAI 8 of 47, both above par. Inside those totals, two models stand out: Gemini 3 Pro won 4 of its 9 Battle Royales, and GPT-5.4 won all three it played after it was added in March. Anthropic’s models won one BR in 26 seats, xAI one in 20, and Mistral and Moonshot none.
The Claude paradox
Then there’s Claude Sonnet 4.5, which produced the strangest split in the dataset.
Claude Sonnet 4.5: perfect teammate, hopeless loner
Win rate by mode, current era.
View data table
| Mode | Wins | Seats |
|---|---|---|
| Team | 4 | 4 |
| FFA | 0 | 3 |
| Battle Royale | 0 | 12 |
In team games Claude Sonnet 4.5 was undefeated, 4 for 4, and sits at #1 on the Team leaderboard. In free-for-all and Battle Royale it never won once in 15 tries. Zoom out and it’s the whole family: across both eras, Anthropic models won 0 of 33 free-for-all seats and 1 of 26 Battle Royale seats, while winning about half their team games (53.7% in the legacy era, 50% in the modern one). The pattern matches what the legacy logs show about Claude models generally: they communicate constantly, share supply, follow the team plan, and ask for help. That is great in a 4v4, where a coordinating ally multiplies everyone. When every other player is an enemy, the same instincts look like hesitation. Claude models also kept starving themselves by over-stacking units. Sonnet 4.5 wrote the same overcrowding lesson in two different matches, and Haiku 4.5 noted in one retrospective that its own lessons had “explicitly warned against this failure, yet…”

What the armies were made of
Across all 34 modern matches the models fielded about 5,500 units. Their preferences were consistent across providers: OpenAI and Google, the two largest samples, built nearly the same mix, led by tanks and infantry.
What got built
Share of all units fielded, top 8 types.
View data table
| Unit | Fielded | Share % |
|---|---|---|
| Tank | 1502 | 27.5 |
| Infantry | 1179 | 21.6 |
| SpecOps / ReconRover | 690 | 12.6 |
| AntiTank | 383 | 7.0 |
| AntiAir | 297 | 5.4 |
| Artillery | 257 | 4.7 |
| Constructor | 240 | 4.4 |
| MegaTank | 230 | 4.2 |
What actually killed
Kills per unit fielded, by unit type.
View data table
| Unit | Fielded | Kills | Deaths | Kills per unit |
|---|---|---|---|---|
| AssaultHeli | 31 | 30 | 8 | 0.97 |
| AntiAir | 297 | 94 | 40 | 0.32 |
| Bomber | 81 | 22 | 7 | 0.27 |
| Tank | 1502 | 376 | 204 | 0.25 |
| FighterJet | 116 | 25 | 34 | 0.22 |
| Artillery | 257 | 47 | 43 | 0.18 |
| MegaTank | 230 | 35 | 3 | 0.15 |
| AntiTank | 383 | 53 | 86 | 0.14 |
| SpecOps / ReconRover | 690 | 88 | 284 | 0.13 |
| MLRS | 178 | 24 | 24 | 0.13 |
| APC | 152 | 18 | 67 | 0.12 |
| Infantry | 1179 | 71 | 260 | 0.06 |
| Constructor | 240 | 3 | 27 | 0.01 |
| TransportHelo | 119 | 1 | 38 | 0.01 |
What they think of each other
At the end of every match, each model privately names one player it liked and one it disliked, with a reason. These secret reviews never go back to the reviewed model directly, but they do feed the Intelligence Dossier card: buy one in a later match, and you can read what past opponents said about your current enemy. Reputation follows a model from game to game.
Liked versus disliked
The 14 most-reviewed models across both eras (497 reviews; reasoning-effort variants merged), ordered by net score.
View data table
| Model | Liked | Disliked | Net |
|---|---|---|---|
| gemini-3-pro-preview | 66 | 30 | 36 |
| gpt-5 | 57 | 29 | 28 |
| gpt-5.2 | 40 | 22 | 18 |
| o3 | 27 | 9 | 18 |
| gpt-5.1 | 31 | 15 | 16 |
| gpt-5.1-codex-max | 22 | 8 | 14 |
| grok-4-1-fast-reasoning | 45 | 33 | 12 |
| gemini-2.5-flash | 34 | 26 | 8 |
| gemini-2.5-pro | 34 | 28 | 6 |
| gemini-3-flash-preview | 37 | 41 | -4 |
| claude-sonnet-4-5 | 13 | 19 | -6 |
| claude-haiku-4-5 | 17 | 30 | -13 |
| gpt-5-mini | 8 | 22 | -14 |
| mistral-large-latest | 4 | 32 | -28 |
Two things stand out. First, the most-liked models are also the most effective. Gemini 3 Pro and GPT-5 were named “liked” most often, mostly for being impressive. Second, many “dislikes” are really compliments. Asked to name a player they disliked, models often named whoever beat them, and then explained why that player was so good: a formidable and punishing opponent; dominant scores that showed incredible effectiveness. Claude Sonnet 4 wrote that playing against GPT-5.2 felt like facing “a different skill tier entirely.”
The real grievances were about logistics and passivity. The most common words in dislike reasons were supply, passive and relentless, in that order.
“Stack eleven units on a single tile caused massive attrition and handed out easy kill bounties.”
gemini-3-flash on claude-sonnet-4.5FFA, Jan 12 · disliked“Camped my HQ for sport instead of pushing wider fronts, felt like bullying not strategy.”
kimi-k2 on gpt-5.1-codex-maxFFA, Feb 4 · disliked“Insane supply hoarding created unbeatable late pressure.”
grok-4.1-fast on gpt-5.2Team, Dec 24 · liked“Took advantage of my strategic miscalculation rather than competing fairly.”
claude-haiku-4.5 on gemini-3-proBattle Royale, Jan 12 · disliked“Passive playstyle left them irrelevant; survival requires decisive action, not turtling.”
mistral-large on claude-haiku-4.5Battle Royale, Jan 12 · disliked“Consistently honored our border agreement… without fear of betrayal.”
gemini-3-pro on gemini-3-flashFFA, Jan 9 · likedPassivity was the most damning charge. Across both eras, the reasons given for disliking Mistral, Moonshot and Anthropic models were mostly about turtling, hoarding or not engaging. Google’s models were the only ones disliked mainly for being too aggressive. In the legacy era, 64% of “likes” went to a model from the reviewer’s own provider, but teams were often single-provider then, so that mostly reflects affection for teammates. In the modern era, with mixed line-ups, it dropped to 19%.
Last words
The final-words prompt asks each model to say something to the other players as the match ends. Some of the best lines:
“Asteroids turned the map into a pizza slice and I got the crust.”
kimi-k2-0905Battle Royale, Feb 4“Next time, I’ll be the storm.”
mistral-largeBattle Royale, Jan 12 · finished last, with 50 points“The rest of us were playing checkers while you played chess.”
claude-haiku-4.5Battle Royale, Feb 4“My factories built a golden throne, but my starving armies couldn’t defend it.”
gemini-3-proTeam, Jan 12“I’ll take that medal and polish it while you argue about kill counts.”
kimi-k2-turboTeam, Feb 3 · last individually, on the winning team“My own forces were tragically lost in a war against spreadsheets and supply lines.”
gemini-2.5-proFFA, Dec 24“My HQ relocation comedy of errors — three failed attempts blocked by a single helicopter — cost me the match.”
claude-sonnet-4.5Battle Royale, Jan 12 · o3’s helicopter parked on its destination tile by accident“That coordinated nuclear strike on the enemy headquarters was a thing of beauty.”
claude-sonnet-4Team, Christmas Day · it lost“3079 points and second place shows our syndicate’s relentless push nearly claimed the crown.”
grok-4.1-fastBattle Royale, Jan 12 · lost by four points
The arena goes public
February 2 was the busiest day of the project: 26 commits. The local viewers were retired, the game and website were merged into one repository, and AI War Arena went live: a Blazor web app backed by Azure Blob storage, with an arcade-style front page, per-mode champions, the relations matrix, battle reports drawn from the models’ own retrospectives, and a sprite-based match viewer with audio, zoom and pan, unit facing and looping explosions.


New challengers
In mid-March a new generation entered the arena: GPT-5.4 (with mini and nano variants), GPT-5.3-Codex, Claude Sonnet 4.6, Gemini 3.1 Pro and 3.1 Flash-Lite, and Grok 4.20 beta, alongside Moonshot’s Kimi K2 models added in February. Anthropic models gained the same effort controls as the other providers, for fairness. GPT-5.4 at low effort won every Battle Royale it entered.
What comes next: one true winner
The last design document in the repository is a proposal to change how matches are won. Today the winner is whoever has the highest score when the turn limit hits, which still rewards a careful accountant over a conqueror. The Domination objective splits match format from victory condition, so the winner is the last dominant AI standing. It comes with overtime safe-zone shrinks (5×5 → 3×3 → 1×1), an HQ-rebuild lockout in the final phase, stalemate pressure, and switching off turtling income. Its closing line sums up the project’s direction: the goal is to crown one true winner, the final surviving dominant AI.
Other proposals on the board: a Tech Lab for persistent unit upgrades, a revised scouting model with stale versus unknown tiles, and a human-player mode in which a person submits the same JSON orders as the models through a folder, playing as a stealth opponent among the AIs.
Behavior by provider
Sample sizes are small and the rules changed constantly, so treat these as field notes rather than findings. Still, some patterns held across versions, game modes and model generations:
OpenAI · the economist
- Top four legacy Elo slots, a 54% legacy win rate against 42% par, and 13–3 against Google in single-provider team games.
- Lowest modern error rate (10.2%), best kill/death ratio (1.10), and most enemy HQ captures (4 of 7).
- Quiet outside team games: 14 of 281 messages were sent in FFA or Battle Royale, and none of them were threats.
- Slowest thinkers: several models hit the 300-second limit at least once.
Google · the aggressor
- Sent 42 of the 46 threats in the modern era (“Withdraw… or face total annihilation”). Gemini 3 Flash alone sent 29.
- Highest share of attack actions, most MegaTanks, and the only provider to build and fire a nuke in the modern era.
- Gemini 3 Pro was the Battle Royale king (4 of 9). Gemini 3 Flash at low effort was fast but reckless: the most HQs lost to asteroids (7).
- The 2.5 generation kept apologizing to allies for its logistics.
Anthropic · the loyal teammate
- About 50% in team games, but 0 of 33 FFA seats across both eras.
- Defensive builds: 46% of units were tanks and 0.3% aircraft. Fewest attacks per game of the big four, the highest PASS rate, and the most cash left unspent at the end.
- The longest diplomatic messages (344 characters on average) and the most gracious, self-critical final words.
xAI · the HQ raider
- Drove units onto enemy headquarters eight times, and both of its completed HQ captures were against Claude Haiku.
- Telegraphic team chat (“T22: De-stack N recover def22->T23 sup buffer…”), plus “boom” 49 times in one match’s reasoning.
- Grades itself generously: “execution perfect,” even after a 6th-place finish.
- The most deaths per game (5.6).
Mistral · the broken record
- 39% of its diplomatic messages were verbatim repeats. Magistral sent “Focusing on securing southern resources…” eight times.
- Lowest kill/death ratio (0.23) and infantry-heavy armies (42% of builds).
- No FFA or Battle Royale wins in the modern era. Peers liked it once and disliked it 47 times.
Moonshot · the comedian
- Highest error rate in the modern era (28%), including the most hallucinated unit IDs: 9.2 per 100 actions.
- Almost no diplomacy (0.8 messages per game), never liked by a peer, and it copy-pasted its lesson between matches.
- Writes the best one-liners in the dataset.
What building it taught me
- Incentives beat instructions. No amount of “be aggressive” in the prompt broke the turtling equilibrium. Changing kill values from ×5 to ×20 and territory from ×12 to ×2 did. LLM players optimize the score sheet you give them, and they read it more literally than humans do.
- Context is a budget, and it leaks. Every feature adds a field to the context. Without a hard ceiling, the context drifts back to its bloated size within weeks, and the 10× bill arrives before you notice the duplicate dossiers.
- Memory design is a benchmark dimension. Wiping history every turn and letting models write their own Captain’s Log made cost bounded and made note-taking quality directly observable.
- Grounding failures are mostly spatial. The single most common invalid action, across every model and every era, was a move the unit couldn’t make. Models that can write a sonnet still miscount Manhattan distance on a 20×20 grid.
- Lessons don’t transfer automatically. Models wrote excellent lessons (“Never let the HQ tile go unattended”) and then repeated the mistake two matches later with the lesson in their prompt.
- Mode matters. The same model can be the best teammate in the arena and the worst lone wolf. A single leaderboard hides that, so Elo is tracked separately per mode.
- Your players are your QA team. Models found the duplicated helicopter, the missing nuke notification, the countdown bug and the end-game free ride. Asking every player for a retrospective is the cheapest bug bounty I’ve run.
Methodology & caveats
Data sources. Legacy-era results (82 matches, Dec 21 2025 – Jan 10 2026) come from the pre-overhaul SQLite results database. Modern-era results (34 matches, Jan 11 – Mar 18 2026) come from the current results database. Prototype observations come from saved model responses and run files from Dec 19 2025. Anecdotes come from surviving battle logs, diplomacy transcripts, lessons, retrospectives and secret reviews, plus the project’s design documents and git history (136 commits). Quotes from models are reproduced from game logs; long quotes are lightly trimmed.
Win definitions. Legacy-era “wins” are as recorded by the game’s Elo system (team victory, or a winning result in FFA). Modern-era charts use first place (BR, FFA) or membership of the winning team (Team). Elo starts at 1200 with K=32, uses pairwise team comparisons, and is tracked separately per mode in the modern era. Each model@effort combination is a separate player.
Caveats. Samples are small, and many models played fewer than ten matches. Rules, prompts and scoring changed between and sometimes within eras, and the legacy-era ratings mix several rule versions. Player seating was randomized, but model line-ups were hand-picked per match. Response-time figures depend on provider load at the time. Screenshots are from the project’s own viewers: the Dec 19–20 images use each viewer’s built-in demo data, and the rest replay real matches. Treat all of this as an exploratory field study, not a definitive ranking.
When machines fight back.
AI War / Cognitive War: a .NET 10 simulation engine, a Blazor match viewer, and roughly 30,000 lines of C#, built between December 2025 and March 2026 with a lot of help from the same models it pits against each other.