Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

AI Observability

What You Already Know

You already know how ordinary software observability works: logs, metrics, traces, errors, dashboards, and alerts. AI systems need all of that, but they also need to explain behavior that is semantic, probabilistic, and workflow-dependent.

AI observability answers:

What did the system believe it was doing, what evidence did it use, what did it produce, what did it cost, and who approved it?

The Failure Story

A customer disputes a case outcome. The team opens the logs and finds:

POST /api/cases/123/review 200
model=gpt-x latency=4.8s tokens=8290

That is not enough.

The team needs to know which prompt version ran, which documents were retrieved, whether the retrieved evidence belonged to the same person, what the model output looked like before post-processing, whether a human changed it, which analyst approved it, whether the same case would pass today’s evals, and why cost spiked.

Ordinary logs say the request happened. AI observability must make the workflow explainable.

The Core Concept

AI observability has four overlapping records:

RecordPurpose
Debug logshelp engineers diagnose failures
Traces and spansshow cross-service execution flow
Semantic eventsrecord AI-specific meaning and evidence
Audit recordspreserve accountable business decisions

Do not collapse these into one log stream. They have different audiences, retention policies, and privacy constraints.

The Trace-to-Decision Ladder

A trace becomes valuable when it can answer a decision question. Build observability upward:

LayerQuestion it answers
request tracewhere did time go?
model spanwhich model, prompt, and token path ran?
evidence spanwhat context was retrieved and used?
semantic eventwhat workflow meaning did the AI output have?
human actionwho accepted, rejected, or corrected it?
audit recordwhat can be proven later?

This ladder keeps observability from becoming a pile of logs. Each layer exists because a future operator, engineer, analyst, or auditor will need a different answer.

Trace the Workflow, Not Only the Request

OpenTelemetry’s generative AI semantic conventions are useful because they name model, operation, prompts, completions, usage, and system attributes. Use them as a starting point, but adapt the trace to the workflow.

A case review trace might look like:

case.review
  case.load
  evidence.retrieve
  prompt.render
  model.generate_risk_note
  output.validate_schema
  eval.run_inline_checks
  analyst.queue_publish
  analyst.review
  audit.packet_finalize

Each span should carry low-cardinality attributes:

  • tenant ID hash
  • workflow type
  • case risk tier
  • prompt version
  • model provider
  • model name
  • tool name
  • outcome class
  • retry count

Avoid high-cardinality or sensitive attributes in metrics. Store sensitive evidence in controlled audit records, not in every trace attribute.

Sources to Pair With This Chapter

  • OpenTelemetry, Generative AI semantic conventions: use for model-operation attributes and telemetry naming.
  • OpenTelemetry, Traces: use for cross-service execution modeling.
  • OpenAI Agents SDK, Tracing: use for agent workflow trace concepts.
  • Arize Phoenix, LLM tracing and evaluation: use as practitioner tooling reference for traces and evals.
  • Reddit practitioner discussion, LLM observability fields: use only as anecdotal signal that teams want prompt, cost, latency, and step-level metadata in real traffic.

Semantic Events

AI-specific events should record meaning:

{
  "event_type": "ai.risk_note.generated",
  "case_id": "case-001",
  "workflow_version": "case-review-v3",
  "prompt_version": "risk-note-2026-05-11",
  "model": "frontier-model",
  "evidence_packet_id": "evidence-889",
  "output_schema_version": "risk-note-v2",
  "grounding_score": 0.82,
  "policy_flags": ["possible_false_positive"],
  "cost_usd": 0.048,
  "latency_ms": 2140
}

The exact fields will vary, but the principle is stable: record the semantic unit you will later debug, evaluate, audit, and price.

Prompt and Model Versioning

If you cannot answer which prompt and model produced an output, you cannot run a serious incident review.

Version:

  • system prompt
  • task prompt
  • retrieval template
  • tool definitions
  • output schema
  • model provider
  • model name
  • model settings
  • evaluation suite

Treat prompt changes like code changes. They need review, release notes, eval results, and rollback.

Cost and Latency Observability

AI observability must include economics:

  • input tokens
  • output tokens
  • cached tokens
  • model price tier
  • tool-call cost
  • retrieval cost
  • total workflow cost
  • cost by tenant
  • cost by feature
  • cost by successful workflow

Latency also needs workflow-level visibility:

  • time to first useful output
  • end-to-end workflow completion time
  • queue delay
  • model latency
  • retrieval latency
  • human wait time

Average latency can hide bad tails. Track percentiles.

Failure Classification

Classify failures in a way that helps design:

ClassExampleOwner
Input failuremissing required documentworkflow or user
Retrieval failurewrong evidence returneddata or search layer
Model failurehallucinated unsupported factprompt or model layer
Tool failureprovider timeoutintegration layer
Policy failureunsafe recommendationgovernance layer
Human workflow failurereviewer rubber-stampedoperations layer
Cost failureexpensive route overusedarchitecture layer

Failure classification keeps teams from turning every incident into “the model was bad”.

Retention and Access Model

Every observable record needs a retention and access policy. Otherwise “observability” quietly becomes a privacy and governance liability.

RecordKeep forAccess
debug logshort operational windowengineering on-call
trace spanperformance and failure diagnosisengineering and platform owners
semantic eventeval, drift, and product operationsengineering, product, and authorized operations
prompt and completion payloadonly when policy allowsrestricted incident or eval reviewers
audit recordregulatory or contractual periodcompliance, auditors, and approved business owners
human correctioneval improvement and quality reviewreview leads and eval maintainers

This table should be decided before launch. It is much harder to make sensitive telemetry safe after it has already spread through logs, dashboards, and vendor tools.

Worked Example: Correction Loop

An analyst rejects an AI risk note because the model treated a same-name match as the same person.

The system should record:

  1. the evidence packet
  2. the AI output
  3. the analyst correction
  4. the reason code: false_positive_identity_match
  5. the fields that disambiguated the subject
  6. whether this case exists in the eval suite
  7. whether retrieval ranking contributed to the error

Then the system should decide:

  • add a golden fixture?
  • adjust retrieval?
  • adjust prompt?
  • add deterministic identity checks?
  • update reviewer guidance?

Observability is not only seeing. It is feeding correction into the architecture.

Runnable Example

This repository includes a small semantic-event validator:

python3 examples/observability/validate_events.py \
  fixtures/observability/case_review_events.jsonl

The fixture records an AI risk-note event and an analyst review event. The validator checks that each event has workflow identity, case identity, prompt version, model identity, evidence packet linkage, outcome classification, cost, latency, token usage, and policy flags.

This is intentionally smaller than a full OpenTelemetry deployment. It teaches the contract first: every AI event should preserve enough meaning that a later engineer, analyst, or auditor can understand what happened.

Minimum Artifact

By the end of this chapter, produce a semantic trace schema. It should include:

  • workflow identity
  • tenant or scope identity
  • evidence packet ID
  • prompt and model version
  • tool calls and outcomes
  • output schema version
  • cost and token fields
  • latency fields
  • human action linkage
  • audit record linkage

If the schema cannot explain a disputed case, it is instrumentation, not observability.

Common Mistakes

The first mistake is logging prompts and completions everywhere. That can leak sensitive data and create retention problems. Store full payloads only where policy allows, with access control and redaction.

The second mistake is tracking model latency but not workflow latency. A fast model inside a slow review queue may still produce a bad user experience.

The third mistake is having no prompt version. Without it, you cannot reproduce behavior.

The fourth mistake is treating human corrections as support tickets instead of training and evaluation signals.

Self-Check

  1. What is semantic observability?
  2. Why are debug logs, traces, semantic events, and audit records different?
  3. What should be versioned in an AI workflow?
  4. How can human corrections feed evaluation?

Retrieval Practice

Recall:

  • Name five fields an AI semantic event should record.

Explain:

  • Explain why workflow cost is an observability concern, not only a finance concern.

Apply:

  • Design a trace for one AI workflow. Include retrieval, prompt rendering, model call, validation, human review, and audit finalization.

Where This Leaves Us

Observability makes behavior inspectable. The next question is what the system must prevent: data leakage, prompt injection, tool misuse, tenant crossover, weak governance, and unsafe autonomy. That is security and governance.