Compare how LLM judges score the same work

Each run pins a rubric version, a prompt, and a model config, sends it to several judge models, and leaves every score open to human review.

Model rankings

Aggregate scores from rubric-pinned evaluations on this instance

How it works

Every evaluation is pinned to a rubric version, a prompt, and a model config. Re-run it next month — get the same setup.

1

Version your rubrics

Define weighted scoring criteria and lock them to a version. When your rubric evolves, past evaluations stay pinned to the original.

2

Run models side-by-side

Send the same prompt to multiple LLMs in parallel. Compare scores, latency, and reasoning across judges.

3

Layer human review

Optionally add expert judgment over model outputs. Spot disagreements, pick the best response, and assemble labelled reference sets.

Judge Arena