Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Capstone: An Auditable Case System

What You Already Know

You now have the seven pillars:

  • evaluation
  • typed workflows
  • human-in-the-loop design
  • observability
  • security and governance
  • AI economics
  • distribution systems

The capstone shows how they fit together in one architecture.

The System Goal

Build a case-preparation system for a regulated workflow such as KYC, compliance review, public-benefit eligibility, or civic evidence analysis.

The system does not promise that AI makes final decisions. It promises:

The AI prepares an evidence packet. The analyst validates the decision. The audit trail proves what happened.

How to Use This Capstone

Do not read the capstone as a single product spec. Read it as a compression test for the architecture. A good production AI design should survive being described through the same layers:

  • domain model
  • evaluation plan
  • typed workflow
  • human review
  • observability
  • security
  • economics
  • distribution
  • reference contracts

If one layer is missing, the design may still demo well, but it is not yet ready for serious operational trust.

System Context

customer or operator
  -> case intake
  -> document storage
  -> extraction worker
  -> evidence packet builder
  -> AI drafting and risk signal
  -> eval and policy checks
  -> analyst review queue
  -> audit packet
  -> monitoring and improvement loop

The model participates in extraction assistance, summary drafting, missing-evidence detection, and risk-note preparation. It does not own final case approval.

Sources to Pair With This Chapter

Domain Model

Core entities:

Tenant
Case
Document
EvidencePacket
ExtractionRun
RiskSignal
AnalystReview
AuditPacket
EvaluationRun

Core statuses:

Open
WaitingForDocuments
ExtractionRunning
ReadyForAnalystReview
ApprovedByHuman
RejectedByHuman
Escalated
Closed

Core events:

CaseOpened
DocumentUploaded
ExtractionRequested
ExtractionSucceeded
EvidencePacketBuilt
RiskSignalGenerated
MissingDataDetected
AnalystReviewRequested
AnalystApproved
AnalystRejected
AnalystEscalated
AuditPacketFinalized

Every sensitive transition names an actor, timestamp, prior state, new state, and reason.

Evaluation Plan

The eval suite contains:

  • schema validation for model outputs
  • golden case fixtures
  • prompt-injection cases
  • same-name false-positive cases
  • missing-document cases
  • low-context abstention cases
  • tool failure cases
  • LLM judge scoring for note quality and grounding
  • human calibration samples
  • production drift metrics

Release gate:

hard invariants pass
risk-weighted score does not regress
high-risk fixtures pass
cost per case inside budget
latency inside SLO
no new unreviewed tool permissions

The eval report becomes part of the release artifact.

Typed Workflow Plan

Use typed transitions:

enum CaseStatus {
    Open,
    WaitingForDocuments,
    ReadyForAnalystReview,
    ApprovedByHuman,
    RejectedByHuman,
}

enum CaseEvent {
    DocumentUploaded(DocumentId),
    ExtractionSucceeded(DocumentId),
    MissingDataDetected,
    AnalystApproved(AnalystId),
    AnalystRejected(AnalystId),
}

Database constraints mirror the domain:

  • constrained status values
  • unique idempotency keys
  • tenant-scoped foreign keys
  • append-only audit events
  • outbox rows for async jobs
  • immutable evidence packet versions

The system separates operation lifecycle from audit lifecycle. Extraction jobs may retry. Case approval does not.

Human Review Plan

The analyst review queue shows:

  • evidence packet
  • original source documents
  • extracted fields
  • missing evidence
  • risk signals
  • AI draft note
  • policy checks
  • uncertainty indicators
  • previous corrections

The analyst can:

  • approve
  • reject
  • request more evidence
  • mark false positive
  • escalate

Each action records:

  • analyst ID
  • reason code
  • rationale
  • evidence packet version
  • model and prompt version
  • timestamp

High-risk decisions can require second review.

Observability Plan

Trace:

case.intake
  document.store
  extraction.run
  evidence.build
  ai.risk_note.generate
  eval.inline_checks
  review.queue
  analyst.action
  audit.finalize

Semantic events:

  • case.document_uploaded
  • ai.extraction_completed
  • ai.risk_note_generated
  • eval.case_failed
  • analyst.review_completed
  • audit.packet_finalized

Metrics:

  • p50 and p95 workflow latency
  • model latency
  • queue delay
  • review time
  • cost per case
  • correction rate
  • false-positive rate
  • missing-evidence rate
  • prompt-injection detection rate
  • evaluation regression count

Audit records are separate from debug logs and have stricter retention and access controls.

Security and Governance Plan

Controls:

  • tenant-scoped auth
  • tenant-scoped retrieval
  • document text labeled as untrusted evidence
  • no model-owned final approval tool
  • least-privilege tool set per workflow state
  • prompt-injection eval suite
  • redacted observability payloads
  • model and provider inventory
  • data retention policy
  • human oversight policy
  • incident process

Tool boundary:

read_case_evidence: allowed with tenant and case scope
draft_review_note: allowed
request_more_documents: allowed only through policy workflow
approve_case: forbidden to model
reject_case: forbidden to model

Economics Plan

Cost routing:

StepDefault
document parsingdeterministic and OCR provider
extractionmedium model or specialized extractor
missing evidencedeterministic policy rules
risk notefrontier model for high-risk cases only
audit packetdeterministic assembly
final decisionhuman review

Cost SLOs:

  • p50 cost per case
  • p95 cost per case
  • frontier-model percentage
  • retry spend cap
  • cost per analyst-approved case

The system should make cost visible before pricing.

Distribution Plan

Trust artifacts:

  • architecture diagram
  • eval methodology
  • sample audit packet
  • security and tenant-isolation note
  • human oversight policy
  • cost model
  • buyer-specific demo
  • implementation essay
  • public reference example where possible

The demo should map to budget:

Before: analyst manually assembles case packet in 45 minutes.
After: AI prepares packet in 3 minutes, analyst validates in 10 minutes, audit trail is automatic.

The product story is not “AI chat for compliance”. It is “auditable case preparation with human validation”.

Capstone Variants

The same architecture should not be copied blindly into every domain. It should be translated. The invariant is stable:

AI prepares evidence.
Policy and typed workflow constrain the state transition.
An accountable human validates sensitive outcomes.
The audit trail proves the path.

What changes is the risk model, authority model, review burden, and adoption artifact.

Variant 1: Public-Sector Eligibility Review

A public agency wants to reduce backlog for benefit eligibility, permit review, grant triage, or civic evidence analysis. The expensive failure is not only a wrong answer. It is an opaque denial, inaccessible explanation, biased triage, missing appeal evidence, or a public-record retention failure.

LayerAdaptation
Domain modelApplicantId, ProgramId, EligibilityCase, RequiredEvidence, CaseworkerDecision, AppealPacket
Evaluation emphasismissing-evidence detection, multilingual comprehension, disparate-error analysis, appeal reversals, accessibility of explanations
Human authorityAI may prepare eligibility notes; a caseworker owns eligibility decisions and adverse-action rationale
Observabilityrecord evidence source, policy version, translation path, caseworker override, appeal outcome
Security and governancestrict PII handling, retention schedule, public-record boundaries, role-based access, explanation policy
Economicsoptimize for backlog reduction, review-time reduction, appeal rework reduction, and service-level equity
Trust artifactpublic methodology note, appeal packet sample, bias/equity eval summary, retention and access-control note

Forbidden transition:

model_output -> deny_benefit

Allowed transition:

model_output -> evidence_summary -> caseworker_decision -> appealable_audit_packet

The learner mistake is to treat public-sector AI as a faster classifier. The architectural job is to make the system reviewable by applicants, supervisors, auditors, courts, journalists, and future maintainers.

Variant 2: Fintech KYC and LCB-FT Review

A regulated financial institution wants faster onboarding, sanctions triage, beneficial-owner review, transaction-risk summaries, or alert investigation. The expensive failure is regulatory: missed high-risk customers, false positives that overwhelm analysts, unexplained model reliance, weak vendor controls, or evidence that cannot support an audit.

LayerAdaptation
Domain modelCustomerId, BeneficialOwnerId, ScreeningHit, FalsePositiveReason, RiskRating, ComplianceApproval
Evaluation emphasissame-name false positives, entity disambiguation, missing beneficial-owner evidence, threshold calibration, high-risk fixture recall
Human authorityAI may assemble packets and draft risk notes; analysts or compliance officers own onboarding, rejection, escalation, and suspicious-activity processes
Observabilitytrace screening provider, model route, source documents, risk-score inputs, analyst override, final reason code
Security and governancetenant isolation, provider inventory, least-privilege screening tools, vendor-risk review, redacted logs
Economicsreduce false-positive handling cost, bound frontier-model usage, measure cost per approved case and cost per escalated alert
Trust artifactregulator-ready audit packet, model/provider inventory, eval report, tool-permission matrix, human-oversight policy

Forbidden transition:

model_output -> approve_customer
model_output -> reject_customer
model_output -> freeze_account

Allowed transition:

model_output -> risk_note -> analyst_review -> compliance_decision -> immutable_audit_event

The learner mistake is to see KYC as document extraction. The production system is really evidence lifecycle, decision authority, and audit defensibility.

Variant 3: Internal Enterprise Workflow Agent

An enterprise wants an internal agent to answer policy questions, prepare tickets, update systems, draft legal or procurement summaries, or coordinate operational workflows. The expensive failure is quiet privilege misuse: wrong access, incorrect policy interpretation, duplicated work, bad system mutation, or an answer that looks official without owning the authority to be official.

LayerAdaptation
Domain modelEmployeeId, PolicySourceId, TicketId, ToolPermission, ApprovalRequest, SystemChange
Evaluation emphasispolicy-grounding accuracy, stale-source detection, tool-result validation, escalation correctness, refusal for unsupported requests
Human authorityAI may draft, route, and prepare changes; system owners approve privileged actions and irreversible mutations
Observabilitytrace policy source, retrieval timestamp, tool call, approval chain, system mutation, rollback link
Security and governanceleast-privilege tools, just-in-time access, secret redaction, tenant and department scope, approval gates
Economicsreduce ticket handling time, avoid unnecessary tool calls, route low-risk FAQs to cheaper models, measure cost per resolved workflow
Trust artifacttool-permission catalog, system-owner approval policy, eval report for policy-grounding, incident rollback runbook

Forbidden transition:

model_output -> grant_access
model_output -> change_production_config
model_output -> sign_contract

Allowed transition:

model_output -> prepared_action -> owner_approval -> idempotent_execution -> audit_event

The learner mistake is to call this an employee replacement. The safer frame is a workflow assistant that prepares action under typed permissions and human-owned authority.

Variant Design Checklist

For any new capstone variant, fill this before writing product copy:

  • What is the sensitive state transition?
  • Which human role owns that transition?
  • Which model actions are explicitly forbidden?
  • Which evidence packet proves the system had enough context?
  • Which eval fixtures represent unacceptable harm?
  • Which observability fields let an auditor replay the decision path?
  • Which cost metric would destroy the business case if ignored?
  • Which trust artifact maps to the buyer’s risk?

If the variant cannot answer these questions, it is not yet an architecture. It is still a feature idea.

End-to-End Walkthrough

  1. A customer opens a case.
  2. The user uploads identity and address documents.
  3. The upload uses an idempotency key.
  4. The database stores the document and an outbox row for extraction.
  5. The extraction worker reads the outbox and creates an extraction run.
  6. The evidence builder creates a versioned evidence packet.
  7. The AI drafts a risk note using only scoped evidence.
  8. Inline eval checks reject invalid output.
  9. A prompt-injection detector flags suspicious document instructions.
  10. The review queue presents evidence before recommendation.
  11. The analyst marks a same-name sanctions hit as a false positive.
  12. The correction is logged and added to eval candidate review.
  13. The analyst approves the case.
  14. The audit packet is finalized.
  15. Metrics update cost, latency, correction, and drift dashboards.

Every step is designed so that a later reviewer can ask what happened and get an answer.

Reading the Contract Artifacts

The reference fixture behind this capstone maps the prose into endpoint metadata, event schemas, and variant schemas. Read those artifacts as a teaching object:

  • paths show which business actions exist
  • x-risk-level shows which actions need stricter control
  • x-approval-required shows where human authority enters the contract
  • DomainEvent schemas show which events preserve audit evidence
  • CapstoneVariant schemas show how public-sector, fintech, and enterprise workflows preserve the same invariant

The artifact layer is deliberately small. Its job is not to replace real API design. Its job is to prove that the architecture can become contracts instead of staying as prose.

The same contract also has a local API smoke test:

python3 examples/reference-architecture/smoke_contract_api.py

That smoke test checks the boundary behavior the capstone cares about: required fields, path/body consistency, and explicit human approval for the critical case-approval endpoint.

Common Mistakes

The first mistake is making the capstone a chat product. The serious workflow is case preparation and validation, not conversation.

The second mistake is letting the AI own final approval because it is convenient for the demo.

The third mistake is storing only the final summary. Audit requires evidence, versions, transitions, and human reasons.

The fourth mistake is forgetting economics. A beautiful workflow that loses money at scale is not production-ready.

Self-Check

  1. Which capstone components correspond to the seven pillars?
  2. Why are evidence packets versioned?
  3. Which tools are forbidden to the model?
  4. How does a human correction become an evaluation improvement?

Retrieval Practice

Recall:

  • List the end-to-end case flow from document upload to audit packet.

Explain:

  • Explain why the system can be AI-powered without letting the AI make final decisions.

Apply:

  • Choose one of your product ideas and map it to the same seven-pillar architecture. Where is the weakest pillar today?

Where This Leaves Us

This capstone is the architecture pattern the whole book has been building toward. The next passes should deepen each artifact: richer eval harnesses, workflow services, observability schemas, security checklists, cost calculators, and buyer-facing trust packets.

The capstone is not a finished product. It is the control surface for building one without losing evaluation, auditability, human authority, security, economics, or trust.