Games

The bench

The games from the catalog as a comparison: what Alan rated each build, which model built it, and what the build took.

All builds

Costs are list-price equivalents at published API rates; Method names the rates used for each build. Each bar compares a figure with the largest in its column. Token counts are rounded to three figures; the exact count is in each game's details. A dash means no figure is listed.

Alan's reviews

What Alan wrote after playing each build, highest rated first.

    Method

    1. One prompt per build

      Each model gets one prompt and builds the whole game from it.

    2. Every build is playable

      Every build is hosted. You can play any of them from the catalog.

    3. Alan plays and rates

      Alan plays each build and rates it out of 10.

    4. Costs are list-price equivalents

      Each cost is a list-price equivalent at published API rates. The notes below name the rates used for each build.

    The briefs

    How each figure was measured

    Effort is the effort setting the model ran at, named the way its own tool names it. Agents is the number of agents the build's main session started to do parts of the work. Build time and cost are measured for each build as noted here.