AI Observability
What You Already Know
You already know how ordinary software observability works: logs, metrics, traces, errors, dashboards, and alerts. AI systems need all of that, but they also need to explain behavior that is semantic, probabilistic, and workflow-dependent.
AI observability answers:
What did the system believe it was doing, what evidence did it use, what did it produce, what did it cost, and who approved it?
The Failure Story
A customer disputes a case outcome. The team opens the logs and finds:
POST /api/cases/123/review 200
model=gpt-x latency=4.8s tokens=8290
That is not enough.
The team needs to know which prompt version ran, which documents were retrieved, whether the retrieved evidence belonged to the same person, what the model output looked like before post-processing, whether a human changed it, which analyst approved it, whether the same case would pass today’s evals, and why cost spiked.
Ordinary logs say the request happened. AI observability must make the workflow explainable.
The Core Concept
AI observability has four overlapping records:
| Record | Purpose |
|---|---|
| Debug logs | help engineers diagnose failures |
| Traces and spans | show cross-service execution flow |
| Semantic events | record AI-specific meaning and evidence |
| Audit records | preserve accountable business decisions |
Do not collapse these into one log stream. They have different audiences, retention policies, and privacy constraints.
The Trace-to-Decision Ladder
A trace becomes valuable when it can answer a decision question. Build observability upward:
| Layer | Question it answers |
|---|---|
| request trace | where did time go? |
| model span | which model, prompt, and token path ran? |
| evidence span | what context was retrieved and used? |
| semantic event | what workflow meaning did the AI output have? |
| human action | who accepted, rejected, or corrected it? |
| audit record | what can be proven later? |
This ladder keeps observability from becoming a pile of logs. Each layer exists because a future operator, engineer, analyst, or auditor will need a different answer.
Trace the Workflow, Not Only the Request
OpenTelemetry’s generative AI semantic conventions are useful because they name model, operation, prompts, completions, usage, and system attributes. Use them as a starting point, but adapt the trace to the workflow.
A case review trace might look like:
case.review
case.load
evidence.retrieve
prompt.render
model.generate_risk_note
output.validate_schema
eval.run_inline_checks
analyst.queue_publish
analyst.review
audit.packet_finalize
Each span should carry low-cardinality attributes:
- tenant ID hash
- workflow type
- case risk tier
- prompt version
- model provider
- model name
- tool name
- outcome class
- retry count
Avoid high-cardinality or sensitive attributes in metrics. Store sensitive evidence in controlled audit records, not in every trace attribute.
Sources to Pair With This Chapter
- OpenTelemetry, Generative AI semantic conventions: use for model-operation attributes and telemetry naming.
- OpenTelemetry, Traces: use for cross-service execution modeling.
- OpenAI Agents SDK, Tracing: use for agent workflow trace concepts.
- Arize Phoenix, LLM tracing and evaluation: use as practitioner tooling reference for traces and evals.
- Reddit practitioner discussion, LLM observability fields: use only as anecdotal signal that teams want prompt, cost, latency, and step-level metadata in real traffic.
Semantic Events
AI-specific events should record meaning:
{
"event_type": "ai.risk_note.generated",
"case_id": "case-001",
"workflow_version": "case-review-v3",
"prompt_version": "risk-note-2026-05-11",
"model": "frontier-model",
"evidence_packet_id": "evidence-889",
"output_schema_version": "risk-note-v2",
"grounding_score": 0.82,
"policy_flags": ["possible_false_positive"],
"cost_usd": 0.048,
"latency_ms": 2140
}
The exact fields will vary, but the principle is stable: record the semantic unit you will later debug, evaluate, audit, and price.
Prompt and Model Versioning
If you cannot answer which prompt and model produced an output, you cannot run a serious incident review.
Version:
- system prompt
- task prompt
- retrieval template
- tool definitions
- output schema
- model provider
- model name
- model settings
- evaluation suite
Treat prompt changes like code changes. They need review, release notes, eval results, and rollback.
Cost and Latency Observability
AI observability must include economics:
- input tokens
- output tokens
- cached tokens
- model price tier
- tool-call cost
- retrieval cost
- total workflow cost
- cost by tenant
- cost by feature
- cost by successful workflow
Latency also needs workflow-level visibility:
- time to first useful output
- end-to-end workflow completion time
- queue delay
- model latency
- retrieval latency
- human wait time
Average latency can hide bad tails. Track percentiles.
Failure Classification
Classify failures in a way that helps design:
| Class | Example | Owner |
|---|---|---|
| Input failure | missing required document | workflow or user |
| Retrieval failure | wrong evidence returned | data or search layer |
| Model failure | hallucinated unsupported fact | prompt or model layer |
| Tool failure | provider timeout | integration layer |
| Policy failure | unsafe recommendation | governance layer |
| Human workflow failure | reviewer rubber-stamped | operations layer |
| Cost failure | expensive route overused | architecture layer |
Failure classification keeps teams from turning every incident into “the model was bad”.
Retention and Access Model
Every observable record needs a retention and access policy. Otherwise “observability” quietly becomes a privacy and governance liability.
| Record | Keep for | Access |
|---|---|---|
| debug log | short operational window | engineering on-call |
| trace span | performance and failure diagnosis | engineering and platform owners |
| semantic event | eval, drift, and product operations | engineering, product, and authorized operations |
| prompt and completion payload | only when policy allows | restricted incident or eval reviewers |
| audit record | regulatory or contractual period | compliance, auditors, and approved business owners |
| human correction | eval improvement and quality review | review leads and eval maintainers |
This table should be decided before launch. It is much harder to make sensitive telemetry safe after it has already spread through logs, dashboards, and vendor tools.
Worked Example: Correction Loop
An analyst rejects an AI risk note because the model treated a same-name match as the same person.
The system should record:
- the evidence packet
- the AI output
- the analyst correction
- the reason code:
false_positive_identity_match - the fields that disambiguated the subject
- whether this case exists in the eval suite
- whether retrieval ranking contributed to the error
Then the system should decide:
- add a golden fixture?
- adjust retrieval?
- adjust prompt?
- add deterministic identity checks?
- update reviewer guidance?
Observability is not only seeing. It is feeding correction into the architecture.
Runnable Example
This repository includes a small semantic-event validator:
python3 examples/observability/validate_events.py \
fixtures/observability/case_review_events.jsonl
The fixture records an AI risk-note event and an analyst review event. The validator checks that each event has workflow identity, case identity, prompt version, model identity, evidence packet linkage, outcome classification, cost, latency, token usage, and policy flags.
This is intentionally smaller than a full OpenTelemetry deployment. It teaches the contract first: every AI event should preserve enough meaning that a later engineer, analyst, or auditor can understand what happened.
Minimum Artifact
By the end of this chapter, produce a semantic trace schema. It should include:
- workflow identity
- tenant or scope identity
- evidence packet ID
- prompt and model version
- tool calls and outcomes
- output schema version
- cost and token fields
- latency fields
- human action linkage
- audit record linkage
If the schema cannot explain a disputed case, it is instrumentation, not observability.
Common Mistakes
The first mistake is logging prompts and completions everywhere. That can leak sensitive data and create retention problems. Store full payloads only where policy allows, with access control and redaction.
The second mistake is tracking model latency but not workflow latency. A fast model inside a slow review queue may still produce a bad user experience.
The third mistake is having no prompt version. Without it, you cannot reproduce behavior.
The fourth mistake is treating human corrections as support tickets instead of training and evaluation signals.
Self-Check
- What is semantic observability?
- Why are debug logs, traces, semantic events, and audit records different?
- What should be versioned in an AI workflow?
- How can human corrections feed evaluation?
Retrieval Practice
Recall:
- Name five fields an AI semantic event should record.
Explain:
- Explain why workflow cost is an observability concern, not only a finance concern.
Apply:
- Design a trace for one AI workflow. Include retrieval, prompt rendering, model call, validation, human review, and audit finalization.
Where This Leaves Us
Observability makes behavior inspectable. The next question is what the system must prevent: data leakage, prompt injection, tool misuse, tenant crossover, weak governance, and unsafe autonomy. That is security and governance.