GRAIN BENCHMARKS · DESIGN BENCH

One brief. Every model.

Scores can’t tell you whether a page looks right. Design Bench gives every model the same brief and shows you what each one built, side by side, with what the run cost.

First results comingNothing is published until a full campaign finishes

STATUS · NO RESULTS PUBLISHED YET

First results coming.

The gallery fills in once every model has built the first brief and each page has been rendered. Until then there is nothing to put side by side, and we won’t mock it up.

Read how it’s measured
  1. 01

    One written brief

    A single, fixed design brief, published here in full, so you can judge each result against what was actually asked.

  2. 02

    Every model builds it

    Each model gets the same brief and the same tools and produces a working HTML page, not a mock-up.

  3. 03

    Rendered, not described

    Each page is rendered to a screenshot, and where available you can open the HTML the model wrote.

  4. 04

    With the bill attached

    The cost of each run sits under its output, so good design and a good price can be weighed together.

HOW IT’S MEASURED

Measured by running the work.

Grain’s benchmark harness hands each model real tasks and grades what it produced, so a score reflects a model doing the work, not answering a quiz about it.

Grain’s own harness
Every result comes from Grain’s benchmark harness, run on the models Grain actually serves. We do not import scores from model providers or other leaderboards.
Fixed task sets
Each board is a fixed set of tasks with an automated grader. Every model gets the same tasks, the same tools and the same limits.
Every number has an n
Each score shows how many graded runs it rests on. Below three runs it is marked “early result”, and a 95% interval is shown where we have one.
No rounding up
Scores are truncated, never rounded up, and models with equal scores share a rank. Prices come from Grain’s model catalog; an unknown price shows “—”, never zero.
Only measured boards
A board appears only once it has data. The overall ranking appears only when at least two boards are measured, as the mean of normalised board scores.
Re-run on a schedule
Nerf Bench repeats a fixed task set on a schedule so a quiet regression shows up as a line going down, not as an anecdote.

GRAIN MODELS

Brief a model yourself.

Every model on these boards is one you can run in Grain. Give it a real Mission and judge the result yourself.

EN

Language