Lesson 07 · Judgement · 12 min
Checking your judge
A judge is a system, so it needs an evaluation of its own. The method is the same one you already know: build a small dataset, compare against expected answers, and score the agreement.
Label twenty replies yourself, pass or fail, before you look at what the judge said. Then compare. The number you care about is how often the two of you agree, and — more usefully — how the disagreements split.
Around ninety percent agreement is usually workable. Below eighty, the judge is measuring something other than what you meant, and the fix is almost always in the rubric rather than in the model. Read the reasons attached to the disagreements — they will normally show that the judge is applying a rule you did not intend to give it.
Re-check agreement whenever you change the rubric or the judge model. A rubric tuned against one model can behave quite differently against another.
Checkpoint
Reach at least 85% agreement between your labels and the judge's.