Assay
07 / 11

Lesson 07 · Judgement · 12 min

Checking your judge

A judge is a system, so it needs an evaluation of its own. The method is the same one you already know: build a small dataset, compare against expected answers, and score the agreement.

Label twenty replies yourself, pass or fail, before you look at what the judge said. Then compare. The number you care about is how often the two of you agree, and — more usefully — how the disagreements split.

False passThe judge passed a reply you would have failed. These are the expensive ones: they let real problems through.
False failThe judge failed a reply you would have passed. Noisy and annoying, but visible.

Around ninety percent agreement is usually workable. Below eighty, the judge is measuring something other than what you meant, and the fix is almost always in the rubric rather than in the model. Read the reasons attached to the disagreements — they will normally show that the judge is applying a rule you did not intend to give it.

Re-check agreement whenever you change the rubric or the judge model. A rubric tuned against one model can behave quite differently against another.

Checkpoint

Reach at least 85% agreement between your labels and the judge's.

Label each reply yourself first, then reveal the judge's verdict and compare. Agreement is calculated as you go.

Nothing has run yet. Pick a model and press Run — results appear row by row.