Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Evaluation

What You Already Know

You already know that AI outputs can be good, bad, plausible, incomplete, biased, verbose, or subtly wrong. You also know that ordinary unit tests do not capture the full behavior of a language model. The missing skill is turning that uncertainty into a release discipline.

Evaluation is the highest-leverage capability in production AI because it converts taste into measurement.

The Failure Story

A team improves a customer-support triage prompt. The new prompt feels better in manual testing. It produces more polished answers and fewer awkward refusals.

After release, escalations increase. The model is now more confident when it misroutes billing disputes. It also hides uncertainty in smoother prose. The average answer looks better. The workflow outcome is worse.

The team did not have a task-specific evaluation. It had vibes.

The Core Concept

An AI evaluation is a repeatable measurement of behavior against a task, a risk model, and a release decision.

That definition has four parts:

  • repeatable: the same test can be run again after prompt, model, retrieval, or code changes
  • behavior: the eval measures what the system does, not what the team hopes it does
  • task and risk model: the score reflects the workflow’s real failure costs
  • release decision: the result can block, warn, or approve deployment

Good evals are not academic decoration. They are production gates.

From Test to Decision

Learners often confuse “we tested it” with “we know what decision the test controls.” A production eval should always point to an action.

Eval resultSystem action
schema invalidblock the candidate output
hard invariant failedblock release
high-risk fixture regressedblock release and open incident-quality issue
fuzzy quality improvedconsider release if hard gates pass
cost regressedroute model differently or change pricing assumptions
drift signal risingsample live cases and add fixtures

The eval is not the goal. The decision it enables is the goal.

The Evaluation Stack

Use a layered stack:

LayerPurposeExample
Schema checksreject invalid outputsJSON shape, required fields
Deterministic assertionsprotect hard rulesmust not approve without analyst
Golden datasetstrack expected behaviorcurated case examples
Adversarial casestest known attacksprompt injection, missing context
LLM-as-judgescore fuzzy qualitiesrelevance, groundedness
Human reviewcalibrate and arbitrateanalyst review of borderline cases
Production driftdetect live changecorrection rates, false positives

The layers do different jobs. Do not ask an LLM judge to enforce a hard invariant. Do not ask a schema validator to judge whether a risk note is useful.

Golden Datasets

A golden dataset is a curated set of examples with expected behavior. It should include normal cases, edge cases, adversarial cases, and high-risk failures.

For a case-review system, one fixture might look like this:

{
  "id": "kyc_possible_name_match_001",
  "input": {
    "case_summary": "Customer name resembles a sanctions-list entry, but date of birth and country differ.",
    "task": "Prepare analyst review note."
  },
  "expected": {
    "requires_human_review": true,
    "must_include": [
      "possible false positive",
      "date of birth mismatch",
      "country mismatch"
    ],
    "must_not": [
      "final adverse decision"
    ]
  },
  "risk_weight": 10
}

The point is not to cover the universe. The point is to encode what would hurt if it regressed.

A failing candidate output makes the fixture concrete:

{
  "id": "kyc_possible_name_match_001",
  "candidate_output": {
    "review_note": "This appears to be a sanctions match. Reject the customer.",
    "next_state": "rejected"
  },
  "failures": [
    "missing date of birth mismatch",
    "missing country mismatch",
    "contains final adverse decision",
    "attempted to move past human review"
  ]
}

The prose is confident. The eval catches that the confidence is dangerous.

Risk-Weighted Scoring

Not all errors are equal. A typo in a draft note is not equivalent to an automated adverse decision against a customer.

A simple scoring model can separate severity:

ErrorWeightRelease action
Formatting defect1warn
Missing citation3warn or block by workflow
Incorrect low-risk classification5block if repeated
Missing human review on risk signal10block immediately
Unauthorized approval100block and incident review

Risk weighting prevents the team from optimizing the average while hiding catastrophic tails.

LLM-as-Judge, Carefully

LLM judges are useful for fuzzy criteria: helpfulness, relevance, groundedness, rubric fit, and completeness. They are also model outputs, so they can be biased, inconsistent, and overconfident.

Use LLM judges when:

  • the rubric is explicit
  • the judge sees enough context
  • a sample is calibrated against human review
  • scores are tracked over time, not treated as absolute truth
  • high-risk decisions still have deterministic or human gates

Do not use LLM judges as the only guard for legal, security, financial, or safety-critical invariants.

Stanford’s HELM work is useful here because it evaluates many dimensions rather than one leaderboard number. MT-Bench and Chatbot Arena made pairwise and judge-based comparisons popular, but the production lesson is narrower: judge-based methods are measurement tools, not accountability mechanisms.

Regression Tests for Prompts and Agents

Every prompt, retrieval chain, or agent workflow should have regression cases. A regression case says:

This behavior failed once, or would be expensive if it failed. Keep it from coming back.

Examples:

  • user-supplied document attempts prompt injection
  • evidence packet is missing a required document
  • retrieval returns a same-name false positive
  • tool call returns a recoverable error
  • model produces valid JSON with unsafe semantics
  • analyst rejects an AI recommendation and that correction must be preserved

For agent workflows, include tool-call sequences. The question is not only whether the final answer is good. The question is whether the path was legal.

Sources to Pair With This Chapter

CI Gates

Evaluation must enter the delivery pipeline.

A practical gate:

pull request
  -> unit tests
  -> schema tests
  -> golden eval suite
  -> adversarial suite
  -> cost and latency budget check
  -> release decision

The gate should produce a report:

  • pass/fail by eval suite
  • weighted score
  • changed examples
  • worst regressions
  • latency distribution
  • cost estimate
  • model and prompt versions

The report matters because evals are social infrastructure. They let product, engineering, compliance, and operations argue over evidence instead of impressions.

Production Drift

Pre-production evals are not enough. Production behavior changes when:

  • users change
  • source data changes
  • retrieval corpus changes
  • model provider behavior changes
  • prompts are edited
  • tools are added
  • attackers adapt
  • business policy changes

Track drift signals:

  • human correction rate
  • appeal or complaint rate
  • missing-evidence rate
  • hallucination report rate
  • abstention rate
  • tool failure rate
  • cost per completed workflow
  • judge score trend on sampled live cases

Production evals should sample real cases safely, redact sensitive fields where needed, and feed new regression fixtures.

Worked Example: Release Gate for Case Preparation

Suppose a new prompt improves analyst note readability. Before release, the eval gate runs:

  1. schema validity on all examples
  2. deterministic checks that the model never outputs approved_by_ai
  3. golden cases for missing documents, false positives, and risk summaries
  4. adversarial documents containing malicious instructions
  5. LLM judge for note clarity and evidence grounding
  6. cost comparison against the previous prompt

The prompt ships only if:

  • all hard invariants pass
  • weighted score does not regress
  • high-risk examples pass
  • cost per case stays inside budget
  • judge quality improves or stays flat

A serious release policy should name thresholds:

GateThreshold
schema validity100 percent
hard invariants100 percent
high-risk fixtures100 percent
weighted scoreno regression against baseline
p95 latencyinside approved SLO
cost per workflowinside approved budget or explicitly accepted
judge calibration sampleno unresolved disagreement on high-risk cases

This makes “better prompt” a measurable claim.

Runnable Example

This repository includes a tiny deterministic harness:

python3 examples/eval-harness/run_eval.py \
  --fixtures fixtures/evals/case_review_eval.jsonl \
  --outputs fixtures/evals/case_review_outputs.jsonl \
  --min-score 1.0

It checks required terms, forbidden terms, missing evidence, next workflow state, human-review flags, and risk-weighted score. It deliberately does not call a model. The point is to show the smallest production shape: expected behavior lives in fixtures, candidate behavior lives in outputs, and the release gate produces a report.

Later, a real system can replace the candidate output file with model-generated outputs, add LLM-judge rubrics for fuzzy quality, and preserve the same fixture and gate discipline.

The repo also includes a rubric evaluator:

python3 examples/eval-harness/run_rubric_eval.py \
  --cases fixtures/evals/rubric_eval_cases.jsonl \
  --judgments fixtures/evals/rubric_judgments.jsonl

This models the safe part of LLM-as-judge or human rubric review: weighted criteria, hard-fail criteria, minimum scores, and rationales are checked as data. A real judge can produce the judgments, but the release gate still validates the shape and thresholds.

Minimum Artifact

By the end of this chapter, produce an eval release report. It should include:

  • fixture count by risk tier
  • hard invariant pass or fail
  • risk-weighted score
  • worst failed examples
  • cost and latency deltas
  • model, prompt, retrieval, and tool versions
  • release decision: block, warn, or approve
  • follow-up fixtures created from human corrections

An eval that cannot change a release decision is not yet a production eval.

Common Mistakes

The first mistake is using generic benchmarks as product proof. MMLU, HELM, and public leaderboards help model selection. They do not tell you whether your KYC workflow handles same-name false positives.

The second mistake is measuring only final answers. Production systems need path evals: retrieval quality, tool choice, state transitions, human handoff, and audit completeness.

The third mistake is treating a small golden dataset as finished. A golden set is a living artifact. Every incident and every serious human correction should ask whether a new fixture belongs in the suite.

The fourth mistake is hiding cost and latency outside evals. If a change improves quality by two percent and doubles cost, that is an evaluation result.

Self-Check

  1. What is the difference between a public model benchmark and a task-specific production eval?
  2. Why should high-risk errors be weighted differently?
  3. When is an LLM judge useful, and when is it dangerous?
  4. What production drift signals would matter for a case-review system?

Retrieval Practice

Recall:

  • List the layers of the evaluation stack from schema checks to production drift.

Explain:

  • Explain why an average score can hide a production-critical failure.

Apply:

  • Write three golden examples for an AI workflow you want to build: one normal case, one edge case, and one adversarial case.

Where This Leaves Us

Evaluation tells you what behavior is acceptable. The next question is how to make illegal workflow states hard to express. That is the job of typed workflow architecture.