Capstone: An Auditable Case System
What You Already Know
You now have the seven pillars:
- evaluation
- typed workflows
- human-in-the-loop design
- observability
- security and governance
- AI economics
- distribution systems
The capstone shows how they fit together in one architecture.
The System Goal
Build a case-preparation system for a regulated workflow such as KYC, compliance review, public-benefit eligibility, or civic evidence analysis.
The system does not promise that AI makes final decisions. It promises:
The AI prepares an evidence packet. The analyst validates the decision. The audit trail proves what happened.
How to Use This Capstone
Do not read the capstone as a single product spec. Read it as a compression test for the architecture. A good production AI design should survive being described through the same layers:
- domain model
- evaluation plan
- typed workflow
- human review
- observability
- security
- economics
- distribution
- reference contracts
If one layer is missing, the design may still demo well, but it is not yet ready for serious operational trust.
System Context
customer or operator
-> case intake
-> document storage
-> extraction worker
-> evidence packet builder
-> AI drafting and risk signal
-> eval and policy checks
-> analyst review queue
-> audit packet
-> monitoring and improvement loop
The model participates in extraction assistance, summary drafting, missing-evidence detection, and risk-note preparation. It does not own final case approval.
Sources to Pair With This Chapter
- NIST, AI RMF: use as the cross-cutting risk-management frame.
- OWASP GenAI Security Project, LLM Top 10 2025: use for threat and control coverage.
- OpenTelemetry, Generative AI semantic conventions: use for trace and event naming.
- Microsoft Research, Human-AI Interaction Guidelines: use for human oversight design.
- Chris Richardson, Transactional Outbox: use for reliable async workflow publication.
- OpenAI, Evals: use for release-gate and regression-eval structure.
Domain Model
Core entities:
Tenant
Case
Document
EvidencePacket
ExtractionRun
RiskSignal
AnalystReview
AuditPacket
EvaluationRun
Core statuses:
Open
WaitingForDocuments
ExtractionRunning
ReadyForAnalystReview
ApprovedByHuman
RejectedByHuman
Escalated
Closed
Core events:
CaseOpened
DocumentUploaded
ExtractionRequested
ExtractionSucceeded
EvidencePacketBuilt
RiskSignalGenerated
MissingDataDetected
AnalystReviewRequested
AnalystApproved
AnalystRejected
AnalystEscalated
AuditPacketFinalized
Every sensitive transition names an actor, timestamp, prior state, new state, and reason.
Evaluation Plan
The eval suite contains:
- schema validation for model outputs
- golden case fixtures
- prompt-injection cases
- same-name false-positive cases
- missing-document cases
- low-context abstention cases
- tool failure cases
- LLM judge scoring for note quality and grounding
- human calibration samples
- production drift metrics
Release gate:
hard invariants pass
risk-weighted score does not regress
high-risk fixtures pass
cost per case inside budget
latency inside SLO
no new unreviewed tool permissions
The eval report becomes part of the release artifact.
Typed Workflow Plan
Use typed transitions:
enum CaseStatus {
Open,
WaitingForDocuments,
ReadyForAnalystReview,
ApprovedByHuman,
RejectedByHuman,
}
enum CaseEvent {
DocumentUploaded(DocumentId),
ExtractionSucceeded(DocumentId),
MissingDataDetected,
AnalystApproved(AnalystId),
AnalystRejected(AnalystId),
}
Database constraints mirror the domain:
- constrained status values
- unique idempotency keys
- tenant-scoped foreign keys
- append-only audit events
- outbox rows for async jobs
- immutable evidence packet versions
The system separates operation lifecycle from audit lifecycle. Extraction jobs may retry. Case approval does not.
Human Review Plan
The analyst review queue shows:
- evidence packet
- original source documents
- extracted fields
- missing evidence
- risk signals
- AI draft note
- policy checks
- uncertainty indicators
- previous corrections
The analyst can:
- approve
- reject
- request more evidence
- mark false positive
- escalate
Each action records:
- analyst ID
- reason code
- rationale
- evidence packet version
- model and prompt version
- timestamp
High-risk decisions can require second review.
Observability Plan
Trace:
case.intake
document.store
extraction.run
evidence.build
ai.risk_note.generate
eval.inline_checks
review.queue
analyst.action
audit.finalize
Semantic events:
case.document_uploadedai.extraction_completedai.risk_note_generatedeval.case_failedanalyst.review_completedaudit.packet_finalized
Metrics:
- p50 and p95 workflow latency
- model latency
- queue delay
- review time
- cost per case
- correction rate
- false-positive rate
- missing-evidence rate
- prompt-injection detection rate
- evaluation regression count
Audit records are separate from debug logs and have stricter retention and access controls.
Security and Governance Plan
Controls:
- tenant-scoped auth
- tenant-scoped retrieval
- document text labeled as untrusted evidence
- no model-owned final approval tool
- least-privilege tool set per workflow state
- prompt-injection eval suite
- redacted observability payloads
- model and provider inventory
- data retention policy
- human oversight policy
- incident process
Tool boundary:
read_case_evidence: allowed with tenant and case scope
draft_review_note: allowed
request_more_documents: allowed only through policy workflow
approve_case: forbidden to model
reject_case: forbidden to model
Economics Plan
Cost routing:
| Step | Default |
|---|---|
| document parsing | deterministic and OCR provider |
| extraction | medium model or specialized extractor |
| missing evidence | deterministic policy rules |
| risk note | frontier model for high-risk cases only |
| audit packet | deterministic assembly |
| final decision | human review |
Cost SLOs:
- p50 cost per case
- p95 cost per case
- frontier-model percentage
- retry spend cap
- cost per analyst-approved case
The system should make cost visible before pricing.
Distribution Plan
Trust artifacts:
- architecture diagram
- eval methodology
- sample audit packet
- security and tenant-isolation note
- human oversight policy
- cost model
- buyer-specific demo
- implementation essay
- public reference example where possible
The demo should map to budget:
Before: analyst manually assembles case packet in 45 minutes.
After: AI prepares packet in 3 minutes, analyst validates in 10 minutes, audit trail is automatic.
The product story is not “AI chat for compliance”. It is “auditable case preparation with human validation”.
Capstone Variants
The same architecture should not be copied blindly into every domain. It should be translated. The invariant is stable:
AI prepares evidence.
Policy and typed workflow constrain the state transition.
An accountable human validates sensitive outcomes.
The audit trail proves the path.
What changes is the risk model, authority model, review burden, and adoption artifact.
Variant 1: Public-Sector Eligibility Review
A public agency wants to reduce backlog for benefit eligibility, permit review, grant triage, or civic evidence analysis. The expensive failure is not only a wrong answer. It is an opaque denial, inaccessible explanation, biased triage, missing appeal evidence, or a public-record retention failure.
| Layer | Adaptation |
|---|---|
| Domain model | ApplicantId, ProgramId, EligibilityCase, RequiredEvidence, CaseworkerDecision, AppealPacket |
| Evaluation emphasis | missing-evidence detection, multilingual comprehension, disparate-error analysis, appeal reversals, accessibility of explanations |
| Human authority | AI may prepare eligibility notes; a caseworker owns eligibility decisions and adverse-action rationale |
| Observability | record evidence source, policy version, translation path, caseworker override, appeal outcome |
| Security and governance | strict PII handling, retention schedule, public-record boundaries, role-based access, explanation policy |
| Economics | optimize for backlog reduction, review-time reduction, appeal rework reduction, and service-level equity |
| Trust artifact | public methodology note, appeal packet sample, bias/equity eval summary, retention and access-control note |
Forbidden transition:
model_output -> deny_benefit
Allowed transition:
model_output -> evidence_summary -> caseworker_decision -> appealable_audit_packet
The learner mistake is to treat public-sector AI as a faster classifier. The architectural job is to make the system reviewable by applicants, supervisors, auditors, courts, journalists, and future maintainers.
Variant 2: Fintech KYC and LCB-FT Review
A regulated financial institution wants faster onboarding, sanctions triage, beneficial-owner review, transaction-risk summaries, or alert investigation. The expensive failure is regulatory: missed high-risk customers, false positives that overwhelm analysts, unexplained model reliance, weak vendor controls, or evidence that cannot support an audit.
| Layer | Adaptation |
|---|---|
| Domain model | CustomerId, BeneficialOwnerId, ScreeningHit, FalsePositiveReason, RiskRating, ComplianceApproval |
| Evaluation emphasis | same-name false positives, entity disambiguation, missing beneficial-owner evidence, threshold calibration, high-risk fixture recall |
| Human authority | AI may assemble packets and draft risk notes; analysts or compliance officers own onboarding, rejection, escalation, and suspicious-activity processes |
| Observability | trace screening provider, model route, source documents, risk-score inputs, analyst override, final reason code |
| Security and governance | tenant isolation, provider inventory, least-privilege screening tools, vendor-risk review, redacted logs |
| Economics | reduce false-positive handling cost, bound frontier-model usage, measure cost per approved case and cost per escalated alert |
| Trust artifact | regulator-ready audit packet, model/provider inventory, eval report, tool-permission matrix, human-oversight policy |
Forbidden transition:
model_output -> approve_customer
model_output -> reject_customer
model_output -> freeze_account
Allowed transition:
model_output -> risk_note -> analyst_review -> compliance_decision -> immutable_audit_event
The learner mistake is to see KYC as document extraction. The production system is really evidence lifecycle, decision authority, and audit defensibility.
Variant 3: Internal Enterprise Workflow Agent
An enterprise wants an internal agent to answer policy questions, prepare tickets, update systems, draft legal or procurement summaries, or coordinate operational workflows. The expensive failure is quiet privilege misuse: wrong access, incorrect policy interpretation, duplicated work, bad system mutation, or an answer that looks official without owning the authority to be official.
| Layer | Adaptation |
|---|---|
| Domain model | EmployeeId, PolicySourceId, TicketId, ToolPermission, ApprovalRequest, SystemChange |
| Evaluation emphasis | policy-grounding accuracy, stale-source detection, tool-result validation, escalation correctness, refusal for unsupported requests |
| Human authority | AI may draft, route, and prepare changes; system owners approve privileged actions and irreversible mutations |
| Observability | trace policy source, retrieval timestamp, tool call, approval chain, system mutation, rollback link |
| Security and governance | least-privilege tools, just-in-time access, secret redaction, tenant and department scope, approval gates |
| Economics | reduce ticket handling time, avoid unnecessary tool calls, route low-risk FAQs to cheaper models, measure cost per resolved workflow |
| Trust artifact | tool-permission catalog, system-owner approval policy, eval report for policy-grounding, incident rollback runbook |
Forbidden transition:
model_output -> grant_access
model_output -> change_production_config
model_output -> sign_contract
Allowed transition:
model_output -> prepared_action -> owner_approval -> idempotent_execution -> audit_event
The learner mistake is to call this an employee replacement. The safer frame is a workflow assistant that prepares action under typed permissions and human-owned authority.
Variant Design Checklist
For any new capstone variant, fill this before writing product copy:
- What is the sensitive state transition?
- Which human role owns that transition?
- Which model actions are explicitly forbidden?
- Which evidence packet proves the system had enough context?
- Which eval fixtures represent unacceptable harm?
- Which observability fields let an auditor replay the decision path?
- Which cost metric would destroy the business case if ignored?
- Which trust artifact maps to the buyer’s risk?
If the variant cannot answer these questions, it is not yet an architecture. It is still a feature idea.
End-to-End Walkthrough
- A customer opens a case.
- The user uploads identity and address documents.
- The upload uses an idempotency key.
- The database stores the document and an outbox row for extraction.
- The extraction worker reads the outbox and creates an extraction run.
- The evidence builder creates a versioned evidence packet.
- The AI drafts a risk note using only scoped evidence.
- Inline eval checks reject invalid output.
- A prompt-injection detector flags suspicious document instructions.
- The review queue presents evidence before recommendation.
- The analyst marks a same-name sanctions hit as a false positive.
- The correction is logged and added to eval candidate review.
- The analyst approves the case.
- The audit packet is finalized.
- Metrics update cost, latency, correction, and drift dashboards.
Every step is designed so that a later reviewer can ask what happened and get an answer.
Reading the Contract Artifacts
The reference fixture behind this capstone maps the prose into endpoint metadata, event schemas, and variant schemas. Read those artifacts as a teaching object:
- paths show which business actions exist
x-risk-levelshows which actions need stricter controlx-approval-requiredshows where human authority enters the contractDomainEventschemas show which events preserve audit evidenceCapstoneVariantschemas show how public-sector, fintech, and enterprise workflows preserve the same invariant
The artifact layer is deliberately small. Its job is not to replace real API design. Its job is to prove that the architecture can become contracts instead of staying as prose.
The same contract also has a local API smoke test:
python3 examples/reference-architecture/smoke_contract_api.py
That smoke test checks the boundary behavior the capstone cares about: required fields, path/body consistency, and explicit human approval for the critical case-approval endpoint.
Common Mistakes
The first mistake is making the capstone a chat product. The serious workflow is case preparation and validation, not conversation.
The second mistake is letting the AI own final approval because it is convenient for the demo.
The third mistake is storing only the final summary. Audit requires evidence, versions, transitions, and human reasons.
The fourth mistake is forgetting economics. A beautiful workflow that loses money at scale is not production-ready.
Self-Check
- Which capstone components correspond to the seven pillars?
- Why are evidence packets versioned?
- Which tools are forbidden to the model?
- How does a human correction become an evaluation improvement?
Retrieval Practice
Recall:
- List the end-to-end case flow from document upload to audit packet.
Explain:
- Explain why the system can be AI-powered without letting the AI make final decisions.
Apply:
- Choose one of your product ideas and map it to the same seven-pillar architecture. Where is the weakest pillar today?
Where This Leaves Us
This capstone is the architecture pattern the whole book has been building toward. The next passes should deepen each artifact: richer eval harnesses, workflow services, observability schemas, security checklists, cost calculators, and buyer-facing trust packets.
The capstone is not a finished product. It is the control surface for building one without losing evaluation, auditability, human authority, security, economics, or trust.