CR Bench
Introduction
Our starting idea was to have models play the real Clash Royale client through screenshots and clicks. That would be an interesting computer-use test, but it mixes several problems together. The model has to recognize the board, choose a move, place a card and do all of that before the situation changes.
So we used a deterministic simulator where both models observe the same moment in the game. The clock freezes while they choose. Their actions are committed together, then the simulator advances two seconds.
This gives a model time to reason about a push, a defensive placement or whether to save elixir. It still has to live with the consequences. Troops move, attacks take time, and a bad placement does not become good just because the model thought about it for longer.
The result
Each model played the other four twice, swapping sides and starting hands. That gives eight games per model and twenty games overall. These were individual best-of-one games, rather than an elimination bracket or best-of-three series.
| Model | W–L | Glicko-2 | API cost |
|---|---|---|---|
| Astra | 8–0 | 1,907 | $35.92 |
| Gemini | 5–3 | 1,602 | $2.55 |
| Claude | 4–4 | 1,500 | $34.19 |
| Terra | 3–5 | 1,398 | $9.89 |
| Luna | 0–8 | 1,093 | $0.80 |
Costs cover each model’s responses in completed rated games. Together, the 20 games used 4,288 responses and $83.35 in inference. Setup and interrupted attempts are excluded from this figure.
Astra won both games against every opponent. The middle of the table was less clean: Gemini and Claude split their games, as did Claude and Terra. Luna lost all eight.
Here is Astra against Claude. By 1:52.8 of game time, Astra has taken a princess tower. It finishes 1–0. All three of Astra’s towers are still at full health at the end. That is a concrete example of the gap in this run, although one replay cannot tell us how much came from planning, placement accuracy or the opponent’s mistakes.
What the models actually saw
We gave both players the same level-11 deck: Knight, Archers, Minions, Arrows, Fireball, Giant, Musketeer and Mini P.E.K.K.A. Keeping the deck fixed removes deck selection as an explanation for the result.
The observation was structured data: public units and towers, plus the player’s own hand, elixir and next card. The opponent’s hand, elixir, next card and pending action stayed hidden. The players had to make decisions under incomplete information, but they did not have to read a screenshot or aim a mouse.
At each decision, a model could play one card at a position or wait. Both players received a shared prompt describing the rules and card statistics. Every game started with fresh context. A player could carry bounded notes and recent actions through that game, but it did not learn from earlier matches in the tournament.
We used one configuration per family:
- GPT-6 Astra — low reasoning
- Gemini 3 Flash Preview — minimal reasoning
- Claude Opus 5 — low reasoning
- GPT-5.6 Terra — low reasoning
- GPT-5.6 Luna — no reasoning
Every request had a 2,048-token output ceiling and a 60-second deadline. Those limits are shared, but “low,” “minimal” and “none” do not mean equal compute across providers. This is a comparison of these configurations under a common game interface.
Legal moves still matter
Pausing the game removes the pressure to respond quickly. It does not remove the need to follow the rules. A model still has to choose an affordable card, put it in a legal place and produce an action the game can execute.
Astra made 244 play attempts. Only 13 were rejected, about 5%. Luna made 615 attempts and had 335 rejected, about 54%. Luna attempted to play much more often, but a larger number of attempts did not turn into a better result.
| Model | Accepted plays | Rejected plays | Invalid outputs |
|---|---|---|---|
| Astra | 231 | 13 | 0 |
| Gemini | 240 | 80 | 127 |
| Claude | 256 | 45 | 2 |
| Terra | 306 | 81 | 1 |
| Luna | 280 | 335 | 1 |
An invalid output cannot be interpreted as an action and becomes a wait. A rejected play is a parsed placement that the simulator refuses. These are separate counts; deliberate waits are not shown.
Gemini is an interesting counterexample to a simple “fewer mistakes means more wins” explanation. It produced 127 invalid outputs and still finished second. The action counts tell us something about execution reliability. They do not, by themselves, establish which model has the best strategy or why a particular game was won.
There is also more to the outcome than taking a tower. In the first Astra–Gemini game, neither side earned a crown. The match reached the five-minute health tiebreak. Gemini’s damaged princess tower had 638 HP left; Astra’s towers were untouched. Astra won 0–0 on the tiebreak.
This is why a crown count alone would be an incomplete score. A game can produce a meaningful difference in tower health without producing a crown. We rate the simulator’s recorded win or loss, including its declared tiebreak rule.
The middle of the table
Gemini’s 5–3 record came from beating Luna and Terra twice, splitting with Claude, and losing twice to Astra. Claude beat Luna twice, split with Gemini and Terra, and lost twice to Astra.
The two Claude–Gemini games are useful to watch together. Gemini has taken a tower by 2:23.4 in the first game and wins 1–0. With the sides swapped, Claude has taken a tower by 1:10.4 and also wins 1–0. The direction of the result changes within the same pairing.
That is a reason to be careful with the ordering below Astra. Two games against an opponent are enough to expose a difference in this sample, but not enough to characterize every starting hand or strategic interaction.
The cost difference is substantial within the completed run: Gemini’s rated responses cost $2.55, compared with $35.92 for Astra and $34.19 for Claude. These are useful measurements of this workload. They are not fixed prices per match, and a different deck, reasoning setting or game length could change the comparison.
How much should we trust the ratings?
We used Glicko-2, with all twenty games in a single rating period after independent replay verification. Each model started at 1,500 with a rating deviation of 350. The point estimate produces a convenient ordering; the uncertainty is a necessary part of reading it.
| Model | Rating | Rating ± 2 RD |
|---|---|---|
| Astra | 1,907 | 1,582–2,233 |
| Gemini | 1,602 | 1,277–1,927 |
| Claude | 1,500 | 1,175–1,825 |
| Terra | 1,398 | 1,073–1,723 |
| Luna | 1,093 | 767–1,418 |
These ranges are wide. They describe Glicko’s rating uncertainty, not calibrated confidence intervals. The paired games are also correlated. Eight games per model on one deck support a provisional result, not a general ranking of model intelligence.
The simulator matters too. This run uses a frozen CRForge revision, not the live Clash Royale client. Its rules can differ from the live game; for example, Fireball does 206 tower damage here. The benchmark measures decisions in that declared ruleset.
What happened when requests failed
Three attempts at Game 019 stopped after Gemini request failures, at 94, 146 and 120 game seconds. The third interruption was traced to an HTTP 429 rate-limit response. We preserved those attempts, checked their completed simulation prefixes and excluded them from ratings. No winner was assigned to an unfinished game.
Game 019 restarted with the original seeds and fresh model context. The first eighteen completed games stayed unchanged. For the final two games, we added a nominal twelve-second wall-clock cooldown between decision windows. The simulation remained paused during that cooldown.
There were no application-level retries, although the underlying SDK retained two transport retries. These details matter because a leaderboard should explain how it handled operational failures, especially when a model’s access to the game depends on an external API.
Every game
The table below shows each row model’s wins and losses against the column model. It is a round robin: everyone played everyone. You can open any of the twenty matches beneath it.
| Model | Astra | Gemini | Claude | Terra | Luna |
|---|---|---|---|---|---|
| Astra | — | 2–0 | 2–0 | 2–0 | 2–0 |
| Gemini | 0–2 | — | 1–1 | 2–0 | 2–0 |
| Claude | 0–2 | 1–1 | — | 1–1 | 2–0 |
| Terra | 0–2 | 0–2 | 1–1 | — | 2–0 |
| Luna | 0–2 | 0–2 | 0–2 | 0–2 | — |
The animated replays reconstruct the recorded simulation with Supercell artwork. Movement is interpolated between 5 Hz samples; health and outcomes remain exact. Facing and attack animations are illustrative. The recordings do not identify projectile types, so the viewer uses neutral bolts. Both hands are visible to the referee here; each model saw only its own.
What this leaves open
The strongest result in this experiment is Astra’s 8–0 record under a common interface. The more useful research question is what produced that advantage. Our action counts show a large difference in legal execution, but we have not isolated planning quality from rule-following, card knowledge or the way each model uses its limited notes.
A follow-up could vary the deck and repeat more seeds, then compare those results with a version that guarantees legal actions. That would help separate choosing a good plan from expressing a legal move. It would also be useful to compare reasoning settings within a family before drawing conclusions about the family as a whole.
For now, stopping the clock gives us a practical way to study a model’s decisions in an adversarial game without charging it game time for thinking. The replays make those decisions inspectable, and the small tournament gives us a first set of results to build on.
The results data, shared prompt, replay verification hashes and frozen engine revision are available for reference. This is an independent experiment, unaffiliated with Supercell.
Exact model versions and rating settings
- GPT-6 Astra · low
openai/gpt-6-astra-20260903 - Gemini 3 Flash Preview · minimal
google/gemini-3-flash-preview-20251217 - Claude Opus 5 · low
anthropic/claude-opus-5-20260723 - GPT-5.6 Terra · low
openai/gpt-5.6-terra-20260709 - GPT-5.6 Luna · none
openai/gpt-5.6-luna-20260709
Protocol: crforge-starter-strategy-v1. Glicko-2 initial volatility 0.06; tau 0.5. One simultaneous rating period. Model versions and observations are fixed to this run.