An LLM eval rubric is a set of concrete pass/fail criteria you run an AI model’s outputs against, so you can measure whether a prompt or system is actually working instead of relying on a gut feeling from a handful of manual tests. If you’ve ever tweaked a prompt, thought it seemed better, and had no real way to confirm that beyond “it feels right,” an eval is the fix — a repeatable, scored test suite for AI output, the same way unit tests are a repeatable test suite for code.
What Makes an Eval Task Actually Good?
Anthropic’s own engineering guidance on building evals for AI agents sets a high bar for clarity: “a good task is one where two domain experts would independently reach the same pass/fail verdict.” If your grading criteria are vague enough that two reasonable people could disagree on whether an output passed, the eval isn’t measuring anything reliable — it’s just adding noise. A good eval set is also balanced, testing cases where the desired behavior should happen and cases where it shouldn’t, since a one-sided eval can accidentally train a system toward the wrong extreme — for example, testing only “should the agent search the web” cases can push a system toward searching indiscriminately, even when it shouldn’t.
What Common Mistakes Make Evals Useless?
A few patterns show up constantly. Grading the exact sequence of steps an agent took is brittle, since a competent agent regularly finds a valid alternative path the person writing the eval didn’t anticipate — grade the outcome, not the route. Grading against conditions that were never stated in the task description sets the model up for unfair failures, like expecting a specific file path the prompt never mentioned. Shared state between test runs — leftover files, cached data from a previous trial — introduces failures that have nothing to do with how good the model actually is. And if a known-good reference solution can’t pass your own eval, the task specification is broken, not the model.
How Do You Build Your First Eval Set From Scratch?
Start from real failures instead of inventing hypothetical ones: pull 20-50 cases from actual bug reports, support tickets, or manual testing where you know what the correct output should look like. For each one, write a reference solution — a known-good output — both to prove the task is actually solvable and to sanity-check that your grader works correctly before you trust its verdicts on anything else. From there, run the eval regularly as you change prompts or models, and read a sample of the raw transcripts yourself every so often rather than only looking at the pass/fail score, since that’s how you catch a grading bug before it quietly skews weeks of results.
Code-Based Grader or LLM Judge — Which Should You Use?
Match the grader to the task. For objective, checkable outputs — did the SQL query return the right rows, is the JSON valid, does the code pass its tests — a deterministic, code-based grader is faster, cheaper, and more consistent than asking another model to judge it. For subjective quality — is this summary well-written, does this response match the intended tone — an LLM-based grader with a clear rubric is usually the more practical option, similar to how you’d approach AI code review prompts that check for both hard rules (does it compile) and softer judgment calls (is this readable). Whichever grader you use, have a human periodically review its verdicts against your own judgment — LLM judges can hallucinate a justification for the wrong verdict just like any other model output.
Frequently Asked Questions
Do I need dozens of test cases before an eval is useful?
No — even 20 well-chosen, clearly-specified cases pulled from real failures beat 200 vague or overlapping ones. Start small, confirm the grading is trustworthy, then expand coverage over time.
Should evals give partial credit or just pass/fail?
For tasks with multiple components, partial credit gives a more useful signal than binary pass/fail, since it lets you see which specific piece of a multi-step task is failing instead of just knowing the whole thing didn’t succeed.
How does this relate to catching AI hallucinations?
An eval is one of the most reliable ways to measure whether your approach to reducing AI hallucinations is actually working — instead of assuming a prompt change helped, you can run it against a fixed set of known-answer questions and see the pass rate move (or not) in a measurable way.
For a deeper walkthrough of building evals for agentic systems specifically, see Anthropic’s engineering guide, “Demystifying evals for AI agents.”



