GRAIN BENCHMARKS · LEADERBOARD
Models, measured on real work.
Independent results from Grain’s own benchmark harness, on the models Grain runs. No vendor-reported scores, no borrowed numbers: each model gets the same tasks, and each result shows how many runs it rests on.
STATUS · NO RESULTS PUBLISHED YET
First results coming.
Rankings appear here once the first benchmark campaign has run and been graded, not before. Until then there is nothing honest to rank, so this page shows the method instead of numbers.
Read how it’s measured- 01
One board per kind of work
Each board is a fixed task set, such as a coding or research workload, with its own grader. A board goes live only once it has results.
- 02
A score with its sample size
Accept rate or pass rate per model, the number of graded runs behind it, and a 95% interval where there are enough runs.
- 03
What it costs to run
Input and output price per million tokens from Grain’s own model catalog, next to the score, so quality and cost read together.
- 04
An overall view, when earned
Once two or more boards are measured, an overall ranking built from normalised board scores.
HOW IT’S MEASURED
Measured by running the work.
Grain’s benchmark harness hands each model real tasks and grades what it produced, so a score reflects a model doing the work, not answering a quiz about it.
- Grain’s own harness
- Every result comes from Grain’s benchmark harness, run on the models Grain actually serves. We do not import scores from model providers or other leaderboards.
- Fixed task sets
- Each board is a fixed set of tasks with an automated grader. Every model gets the same tasks, the same tools and the same limits.
- Every number has an n
- Each score shows how many graded runs it rests on. Below three runs it is marked “early result”, and a 95% interval is shown where we have one.
- No rounding up
- Scores are truncated, never rounded up, and models with equal scores share a rank. Prices come from Grain’s model catalog; an unknown price shows “—”, never zero.
- Only measured boards
- A board appears only once it has data. The overall ranking appears only when at least two boards are measured, as the mean of normalised board scores.
- Re-run on a schedule
- Nerf Bench repeats a fixed task set on a schedule so a quiet regression shows up as a line going down, not as an anecdote.
Does a model get worse over time? A fixed task set, re-run on a schedule.
Open Nerf Bench One brief, every modelDesign BenchOne design brief, rendered by every model, side by side with cost.
Open Design BenchGRAIN MODELS
Put the models to work.
Every model on these boards is one you can run in Grain. Give it a real Mission and judge the result yourself.