GRAIN BENCHMARKS · NERF BENCH
Does it still do the job?
Models change behind the same name. Nerf Bench re-runs a fixed task set against each model on a schedule, so you can see whether it gets worse over time instead of guessing from a bad afternoon.
STATUS · NO RESULTS PUBLISHED YET
First results coming.
A trend needs history. The first scheduled runs have to finish before there is a line to draw, so this page stays empty until they do. No sample charts, no made-up lines.
Read how it’s measured- 01
The same tasks, every time
A fixed task set that never changes between runs, so a moving score can only mean the model moved.
- 02
On a schedule
Each model is re-run at regular intervals and every run is graded the same way, with its n recorded.
- 03
A trend per model
Pass rate over time on a fixed 0–100% axis, so a small wobble looks small and a real drop looks like one.
- 04
Alerts when it drops
When a model’s pass rate falls between two runs, the drop is flagged with the dates and the size of the change.
HOW IT’S MEASURED
Measured by running the work.
Grain’s benchmark harness hands each model real tasks and grades what it produced, so a score reflects a model doing the work, not answering a quiz about it.
- Grain’s own harness
- Every result comes from Grain’s benchmark harness, run on the models Grain actually serves. We do not import scores from model providers or other leaderboards.
- Fixed task sets
- Each board is a fixed set of tasks with an automated grader. Every model gets the same tasks, the same tools and the same limits.
- Every number has an n
- Each score shows how many graded runs it rests on. Below three runs it is marked “early result”, and a 95% interval is shown where we have one.
- No rounding up
- Scores are truncated, never rounded up, and models with equal scores share a rank. Prices come from Grain’s model catalog; an unknown price shows “—”, never zero.
- Only measured boards
- A board appears only once it has data. The overall ranking appears only when at least two boards are measured, as the mean of normalised board scores.
- Re-run on a schedule
- Nerf Bench repeats a fixed task set on a schedule so a quiet regression shows up as a line going down, not as an anecdote.
Ranked results for every measured board, each with its n.
Open Leaderboard One brief, every modelDesign BenchOne design brief, rendered by every model, side by side with cost.
Open Design BenchGRAIN MODELS
Run the models that hold up.
Every model on these boards is one you can run in Grain. Give it a real Mission and judge the result yourself.