The bench
The games from the catalog as a comparison: what Alan rated each build, which model built it, and what the build took.
All builds
Bar colour shows the model
Costs are list-price equivalents at published API rates; Method names the rates used for each build. Each bar compares a figure with the largest in its column. Token counts are rounded to three figures; the exact count is in each game's details. A dash means no figure is listed.
Alan's reviews
What Alan wrote after playing each build, highest rated first.
Method
-
One prompt per build
Each model gets one prompt and builds the whole game from it.
-
Every build is playable
Every build is hosted. You can play any of them from the catalog.
-
Alan plays and rates
Alan plays each build and rates it out of 10.
-
Costs are list-price equivalents
Each cost is a list-price equivalent at published API rates. The notes below name the rates used for each build.
The briefs
How each figure was measured
Effort is the effort setting the model ran at, named the way its own tool names it. Agents is the number of agents the build's main session started to do parts of the work. Build time and cost are measured for each build as noted here.