GRAIN BENCHMARKS · DESIGN BENCH
One brief. Every model.
Scores can’t tell you whether a page looks right. Design Bench gives every model the same brief and shows you what each one built, side by side, with what the run cost.
STATUS · NO RESULTS PUBLISHED YET
First results coming.
The gallery fills in once every model has built the first brief and each page has been rendered. Until then there is nothing to put side by side, and we won’t mock it up.
Read how it’s measured- 01
One written brief
A single, fixed design brief, published here in full, so you can judge each result against what was actually asked.
- 02
Every model builds it
Each model gets the same brief and the same tools and produces a working HTML page, not a mock-up.
- 03
Rendered, not described
Each page is rendered to a screenshot, and where available you can open the HTML the model wrote.
- 04
With the bill attached
The cost of each run sits under its output, so good design and a good price can be weighed together.
HOW IT’S MEASURED
Measured by running the work.
Grain’s benchmark harness hands each model real tasks and grades what it produced, so a score reflects a model doing the work, not answering a quiz about it.
- Grain’s own harness
- Every result comes from Grain’s benchmark harness, run on the models Grain actually serves. We do not import scores from model providers or other leaderboards.
- Fixed task sets
- Each board is a fixed set of tasks with an automated grader. Every model gets the same tasks, the same tools and the same limits.
- Every number has an n
- Each score shows how many graded runs it rests on. Below three runs it is marked “early result”, and a 95% interval is shown where we have one.
- No rounding up
- Scores are truncated, never rounded up, and models with equal scores share a rank. Prices come from Grain’s model catalog; an unknown price shows “—”, never zero.
- Only measured boards
- A board appears only once it has data. The overall ranking appears only when at least two boards are measured, as the mean of normalised board scores.
- Re-run on a schedule
- Nerf Bench repeats a fixed task set on a schedule so a quiet regression shows up as a line going down, not as an anecdote.
Ranked results for every measured board, each with its n.
Open Leaderboard Drift over timeNerf BenchDoes a model get worse over time? A fixed task set, re-run on a schedule.
Open Nerf BenchGRAIN MODELS
Brief a model yourself.
Every model on these boards is one you can run in Grain. Give it a real Mission and judge the result yourself.