GRAIN BENCHMARKS · NERF BENCH

Does it still do the job?

Models change behind the same name. Nerf Bench re-runs a fixed task set against each model on a schedule, so you can see whether it gets worse over time instead of guessing from a bad afternoon.

First results comingNothing is published until a full campaign finishes

STATUS · NO RESULTS PUBLISHED YET

First results coming.

A trend needs history. The first scheduled runs have to finish before there is a line to draw, so this page stays empty until they do. No sample charts, no made-up lines.

Read how it’s measured
  1. 01

    The same tasks, every time

    A fixed task set that never changes between runs, so a moving score can only mean the model moved.

  2. 02

    On a schedule

    Each model is re-run at regular intervals and every run is graded the same way, with its n recorded.

  3. 03

    A trend per model

    Pass rate over time on a fixed 0–100% axis, so a small wobble looks small and a real drop looks like one.

  4. 04

    Alerts when it drops

    When a model’s pass rate falls between two runs, the drop is flagged with the dates and the size of the change.

HOW IT’S MEASURED

Measured by running the work.

Grain’s benchmark harness hands each model real tasks and grades what it produced, so a score reflects a model doing the work, not answering a quiz about it.

Grain’s own harness
Every result comes from Grain’s benchmark harness, run on the models Grain actually serves. We do not import scores from model providers or other leaderboards.
Fixed task sets
Each board is a fixed set of tasks with an automated grader. Every model gets the same tasks, the same tools and the same limits.
Every number has an n
Each score shows how many graded runs it rests on. Below three runs it is marked “early result”, and a 95% interval is shown where we have one.
No rounding up
Scores are truncated, never rounded up, and models with equal scores share a rank. Prices come from Grain’s model catalog; an unknown price shows “—”, never zero.
Only measured boards
A board appears only once it has data. The overall ranking appears only when at least two boards are measured, as the mean of normalised board scores.
Re-run on a schedule
Nerf Bench repeats a fixed task set on a schedule so a quiet regression shows up as a line going down, not as an anecdote.

GRAIN MODELS

Run the models that hold up.

Every model on these boards is one you can run in Grain. Give it a real Mission and judge the result yourself.

EN

Language