The referee's report — measuring accuracy.
You can't trust — or improve — what you never measure. So score the calls against the truth.
After the match, the review team does not ask the VAR how the calls felt. It gathers a set of incidents whose true verdict is already settled — the replays leave no doubt — has the VAR rule each one again, and counts how many it got right. That count is the VAR's report card, and it is called an eval.
Swap the VAR for a model and it is the same idea. You take a set of questions whose right answers you already know — the ground truth — have the model answer each, and compare its answers against the truth. The fraction it gets right is its accuracy: the simplest score there is.
Why bother? Because “it seems good” is not a number. A real score lets you compare two prompts, catch a change that quietly made things worse, and decide whether the thing is fit to be trusted at all.
Run the referee's report over a set of appeals and watch the VAR's score come in.
A number you can trust and track.
The VAR got four of six right — a 67% report. Now that is something concrete: compare it to last week, to a different VAR, or to a new set of instructions, and you can see improvement or slippage instead of guessing at it.
An eval is a test set with known answers, the model's answers, and a score — and accuracy is the first and simplest one. The catch: a score is only as honest as the cases you put in the test, which is the next thing to get right.