Assay

102 models · $0.00 inference · NVIDIA Build API

Learn to evaluate
AI agents.

An eleven-lesson course on building evaluations, from a first dataset to a release gate. Every lesson ends in a lab you run against real models — no install, no notebook, and no bill.

Start lesson 0111 lessons · about 2.1 hours

Ticket triage · accuracy by category

example from lesson 08

shippingbillingreturnsaccountescalate
llama-3.3-70b0.960.940.700.920.90
nemotron-super-49b0.940.920.680.900.88
qwen2.5-7b0.900.860.620.840.80
mistral-small-24b0.880.840.580.820.78

Every model scores lowest on returns. When models from four different labs share a weakness, the cause is usually the labels rather than the models — lesson 04 covers how to tell.

What you build

One worked example throughout: a support-triage agent, 50 labelled tickets, 13 of them deliberately debatable.