Each run pins a rubric version, a prompt, and a model config, sends it to several judge models, and leaves every score open to human review.
Aggregate scores from rubric-pinned evaluations on this instance
Every evaluation is pinned to a rubric version, a prompt, and a model config. Re-run it next month — get the same setup.
Define weighted scoring criteria and lock them to a version. When your rubric evolves, past evaluations stay pinned to the original.
Send the same prompt to multiple LLMs in parallel. Compare scores, latency, and reasoning across judges.
Optionally add expert judgment over model outputs. Spot disagreements, pick the best response, and assemble labelled reference sets.