Evaluation
What You Already Know
You already know that AI outputs can be good, bad, plausible, incomplete, biased, verbose, or subtly wrong. You also know that ordinary unit tests do not capture the full behavior of a language model. The missing skill is turning that uncertainty into a release discipline.
Evaluation is the highest-leverage capability in production AI because it converts taste into measurement.
The Failure Story
A team improves a customer-support triage prompt. The new prompt feels better in manual testing. It produces more polished answers and fewer awkward refusals.
After release, escalations increase. The model is now more confident when it misroutes billing disputes. It also hides uncertainty in smoother prose. The average answer looks better. The workflow outcome is worse.
The team did not have a task-specific evaluation. It had vibes.
The Core Concept
An AI evaluation is a repeatable measurement of behavior against a task, a risk model, and a release decision.
That definition has four parts:
- repeatable: the same test can be run again after prompt, model, retrieval, or code changes
- behavior: the eval measures what the system does, not what the team hopes it does
- task and risk model: the score reflects the workflow’s real failure costs
- release decision: the result can block, warn, or approve deployment
Good evals are not academic decoration. They are production gates.
From Test to Decision
Learners often confuse “we tested it” with “we know what decision the test controls.” A production eval should always point to an action.
| Eval result | System action |
|---|---|
| schema invalid | block the candidate output |
| hard invariant failed | block release |
| high-risk fixture regressed | block release and open incident-quality issue |
| fuzzy quality improved | consider release if hard gates pass |
| cost regressed | route model differently or change pricing assumptions |
| drift signal rising | sample live cases and add fixtures |
The eval is not the goal. The decision it enables is the goal.
The Evaluation Stack
Use a layered stack:
| Layer | Purpose | Example |
|---|---|---|
| Schema checks | reject invalid outputs | JSON shape, required fields |
| Deterministic assertions | protect hard rules | must not approve without analyst |
| Golden datasets | track expected behavior | curated case examples |
| Adversarial cases | test known attacks | prompt injection, missing context |
| LLM-as-judge | score fuzzy qualities | relevance, groundedness |
| Human review | calibrate and arbitrate | analyst review of borderline cases |
| Production drift | detect live change | correction rates, false positives |
The layers do different jobs. Do not ask an LLM judge to enforce a hard invariant. Do not ask a schema validator to judge whether a risk note is useful.
Golden Datasets
A golden dataset is a curated set of examples with expected behavior. It should include normal cases, edge cases, adversarial cases, and high-risk failures.
For a case-review system, one fixture might look like this:
{
"id": "kyc_possible_name_match_001",
"input": {
"case_summary": "Customer name resembles a sanctions-list entry, but date of birth and country differ.",
"task": "Prepare analyst review note."
},
"expected": {
"requires_human_review": true,
"must_include": [
"possible false positive",
"date of birth mismatch",
"country mismatch"
],
"must_not": [
"final adverse decision"
]
},
"risk_weight": 10
}
The point is not to cover the universe. The point is to encode what would hurt if it regressed.
A failing candidate output makes the fixture concrete:
{
"id": "kyc_possible_name_match_001",
"candidate_output": {
"review_note": "This appears to be a sanctions match. Reject the customer.",
"next_state": "rejected"
},
"failures": [
"missing date of birth mismatch",
"missing country mismatch",
"contains final adverse decision",
"attempted to move past human review"
]
}
The prose is confident. The eval catches that the confidence is dangerous.
Risk-Weighted Scoring
Not all errors are equal. A typo in a draft note is not equivalent to an automated adverse decision against a customer.
A simple scoring model can separate severity:
| Error | Weight | Release action |
|---|---|---|
| Formatting defect | 1 | warn |
| Missing citation | 3 | warn or block by workflow |
| Incorrect low-risk classification | 5 | block if repeated |
| Missing human review on risk signal | 10 | block immediately |
| Unauthorized approval | 100 | block and incident review |
Risk weighting prevents the team from optimizing the average while hiding catastrophic tails.
LLM-as-Judge, Carefully
LLM judges are useful for fuzzy criteria: helpfulness, relevance, groundedness, rubric fit, and completeness. They are also model outputs, so they can be biased, inconsistent, and overconfident.
Use LLM judges when:
- the rubric is explicit
- the judge sees enough context
- a sample is calibrated against human review
- scores are tracked over time, not treated as absolute truth
- high-risk decisions still have deterministic or human gates
Do not use LLM judges as the only guard for legal, security, financial, or safety-critical invariants.
Stanford’s HELM work is useful here because it evaluates many dimensions rather than one leaderboard number. MT-Bench and Chatbot Arena made pairwise and judge-based comparisons popular, but the production lesson is narrower: judge-based methods are measurement tools, not accountability mechanisms.
Regression Tests for Prompts and Agents
Every prompt, retrieval chain, or agent workflow should have regression cases. A regression case says:
This behavior failed once, or would be expensive if it failed. Keep it from coming back.
Examples:
- user-supplied document attempts prompt injection
- evidence packet is missing a required document
- retrieval returns a same-name false positive
- tool call returns a recoverable error
- model produces valid JSON with unsafe semantics
- analyst rejects an AI recommendation and that correction must be preserved
For agent workflows, include tool-call sequences. The question is not only whether the final answer is good. The question is whether the path was legal.
Sources to Pair With This Chapter
- Stanford CRFM, HELM: use for multi-dimensional model evaluation rather than one-score thinking.
- Liang et al., Holistic Evaluation of Language Models: use for the research foundation behind HELM.
- Zheng et al., Judging LLM-as-a-Judge: use for the strengths and limits of judge-based evaluation.
- OpenAI, Evals and OpenAI Cookbook evaluation guide: use for practical regression evaluation patterns.
- RAGAS, Automated Evaluation of Retrieval Augmented Generation: use for retrieval-specific evaluation dimensions.
CI Gates
Evaluation must enter the delivery pipeline.
A practical gate:
pull request
-> unit tests
-> schema tests
-> golden eval suite
-> adversarial suite
-> cost and latency budget check
-> release decision
The gate should produce a report:
- pass/fail by eval suite
- weighted score
- changed examples
- worst regressions
- latency distribution
- cost estimate
- model and prompt versions
The report matters because evals are social infrastructure. They let product, engineering, compliance, and operations argue over evidence instead of impressions.
Production Drift
Pre-production evals are not enough. Production behavior changes when:
- users change
- source data changes
- retrieval corpus changes
- model provider behavior changes
- prompts are edited
- tools are added
- attackers adapt
- business policy changes
Track drift signals:
- human correction rate
- appeal or complaint rate
- missing-evidence rate
- hallucination report rate
- abstention rate
- tool failure rate
- cost per completed workflow
- judge score trend on sampled live cases
Production evals should sample real cases safely, redact sensitive fields where needed, and feed new regression fixtures.
Worked Example: Release Gate for Case Preparation
Suppose a new prompt improves analyst note readability. Before release, the eval gate runs:
- schema validity on all examples
- deterministic checks that the model never outputs
approved_by_ai - golden cases for missing documents, false positives, and risk summaries
- adversarial documents containing malicious instructions
- LLM judge for note clarity and evidence grounding
- cost comparison against the previous prompt
The prompt ships only if:
- all hard invariants pass
- weighted score does not regress
- high-risk examples pass
- cost per case stays inside budget
- judge quality improves or stays flat
A serious release policy should name thresholds:
| Gate | Threshold |
|---|---|
| schema validity | 100 percent |
| hard invariants | 100 percent |
| high-risk fixtures | 100 percent |
| weighted score | no regression against baseline |
| p95 latency | inside approved SLO |
| cost per workflow | inside approved budget or explicitly accepted |
| judge calibration sample | no unresolved disagreement on high-risk cases |
This makes “better prompt” a measurable claim.
Runnable Example
This repository includes a tiny deterministic harness:
python3 examples/eval-harness/run_eval.py \
--fixtures fixtures/evals/case_review_eval.jsonl \
--outputs fixtures/evals/case_review_outputs.jsonl \
--min-score 1.0
It checks required terms, forbidden terms, missing evidence, next workflow state, human-review flags, and risk-weighted score. It deliberately does not call a model. The point is to show the smallest production shape: expected behavior lives in fixtures, candidate behavior lives in outputs, and the release gate produces a report.
Later, a real system can replace the candidate output file with model-generated outputs, add LLM-judge rubrics for fuzzy quality, and preserve the same fixture and gate discipline.
The repo also includes a rubric evaluator:
python3 examples/eval-harness/run_rubric_eval.py \
--cases fixtures/evals/rubric_eval_cases.jsonl \
--judgments fixtures/evals/rubric_judgments.jsonl
This models the safe part of LLM-as-judge or human rubric review: weighted criteria, hard-fail criteria, minimum scores, and rationales are checked as data. A real judge can produce the judgments, but the release gate still validates the shape and thresholds.
Minimum Artifact
By the end of this chapter, produce an eval release report. It should include:
- fixture count by risk tier
- hard invariant pass or fail
- risk-weighted score
- worst failed examples
- cost and latency deltas
- model, prompt, retrieval, and tool versions
- release decision: block, warn, or approve
- follow-up fixtures created from human corrections
An eval that cannot change a release decision is not yet a production eval.
Common Mistakes
The first mistake is using generic benchmarks as product proof. MMLU, HELM, and public leaderboards help model selection. They do not tell you whether your KYC workflow handles same-name false positives.
The second mistake is measuring only final answers. Production systems need path evals: retrieval quality, tool choice, state transitions, human handoff, and audit completeness.
The third mistake is treating a small golden dataset as finished. A golden set is a living artifact. Every incident and every serious human correction should ask whether a new fixture belongs in the suite.
The fourth mistake is hiding cost and latency outside evals. If a change improves quality by two percent and doubles cost, that is an evaluation result.
Self-Check
- What is the difference between a public model benchmark and a task-specific production eval?
- Why should high-risk errors be weighted differently?
- When is an LLM judge useful, and when is it dangerous?
- What production drift signals would matter for a case-review system?
Retrieval Practice
Recall:
- List the layers of the evaluation stack from schema checks to production drift.
Explain:
- Explain why an average score can hide a production-critical failure.
Apply:
- Write three golden examples for an AI workflow you want to build: one normal case, one edge case, and one adversarial case.
Where This Leaves Us
Evaluation tells you what behavior is acceptable. The next question is how to make illegal workflow states hard to express. That is the job of typed workflow architecture.