Assay
06 / 11

Lesson 06 · Judgement · 14 min

Writing a judge rubric

The customer-facing reply has no single correct wording, so no comparison against a stored answer will work. This is where a model-graded check earns its place: another model reads the reply and decides whether it meets a written standard.

That written standard is the rubric, and its quality determines everything. A rubric that says "is this reply good?" produces scores that drift between runs, because the judge is free to invent its own definition of good each time.

  • Ask one question at a time. A judge asked about tone and accuracy together will blur them into a single impression.
  • Make it binary where you can. Pass or fail is easier to check than a score out of five.
  • State what does not matter, so the judge stops grading it.
  • Require a reason. A verdict without a reason cannot be audited.
Does the reply acknowledge the specific problem the customer
described, without promising any outcome (refund, replacement,
delivery date) that the assistant cannot guarantee?

Judge only this. Ignore tone, length and formatting.

Reach for a judge only when the thing you are checking genuinely has no right answer. Every model-graded check costs a call, takes time, and introduces a second system that can be wrong — which is why the next lesson is about checking whether it is.

Checkpoint

Find a reply where you disagree with the judge, and change the rubric so it decides your way.

Edit the rubric, then run it over ten replies. Read the judge's reasons, not just its verdicts.

Nothing has run yet. Pick a model and press Run — results appear row by row.