Research Synthesis Notes
This page explains how the book uses its sources. It is not a neutral bibliography. It is a map from outside work to architectural decisions.
Standards and Governance
NIST AI RMF
NIST’s AI Risk Management Framework is the book’s main risk-management scaffold. The useful move is its separation of governance, mapping, measurement, and management. That prevents a narrow “model quality” frame.
Architectural use:
- governance becomes the policy and ownership layer
- mapping becomes workflow and context modeling
- measurement becomes evals and observability
- management becomes release gates, incident handling, and continuous improvement
EU AI Act
The EU AI Act is used as a regulatory pressure model, especially for high-risk systems. The book does not treat it as a coding checklist. It uses it to keep documentation, human oversight, accuracy, robustness, cybersecurity, and monitoring visible in the architecture.
Architectural use:
- human oversight must be designed, not assumed
- audit evidence must survive after the model call
- post-deployment monitoring is part of the system lifecycle
OWASP LLM Top 10
OWASP gives the security vocabulary for LLM-specific failure modes. Its main contribution to the book is the idea that LLM risk is not just bad output. It includes prompt injection, data leakage, tool misuse, excessive agency, insecure output handling, and supply chain exposure.
Architectural use:
- separate evidence from instruction
- scope tools by workflow state
- make prompt-injection tests part of evals
- keep high-risk writes approval-gated
Evaluation Research
HELM
HELM is useful because it resists one-number evaluation. It evaluates across scenarios, metrics, and dimensions. The production lesson is that evaluation should match the workflow’s real risk, not a public leaderboard.
Architectural use:
- evaluate many dimensions
- separate capability from suitability
- report trade-offs instead of hiding them behind a single average
MT-Bench and LLM-as-Judge
MT-Bench popularized judge-based comparison for conversational quality. The book uses it carefully: LLM judges can help score fuzzy qualities, but they should not enforce hard safety, legal, or workflow invariants.
Architectural use:
- use judges for rubric-based quality
- calibrate against humans
- keep deterministic checks for forbidden behavior
RAGAS
RAGAS is useful for retrieval-augmented systems because it separates answer quality from retrieval quality. That distinction matters when a model gives a plausible answer using the wrong evidence.
Architectural use:
- evaluate context relevance
- evaluate groundedness
- inspect retrieval failures separately from generation failures
Workflow Architecture
Domain Events and Event Sourcing
Martin Fowler’s domain-event and event-sourcing writing gives the book its language for business-significant events. For AI systems, this matters because a model output is not the same as a business event.
Architectural use:
RiskSignalGeneratedis notCaseApproved- state transitions should be explicit
- audit events should preserve actor, reason, and prior state
Transactional Outbox and Idempotent Consumer
The outbox and idempotent-consumer patterns are the book’s reliability foundation for async AI jobs. Model calls, document extraction, and review queue publication all fail in ordinary distributed-systems ways.
Architectural use:
- write state changes and outbox rows together
- publish asynchronously with retries
- make repeated worker delivery safe
Temporal Durable Workflows
Temporal is used as a reference point for durable execution and workflow determinism. The book does not require Temporal, but it borrows the discipline: workflows must survive retries, restarts, and long-running waits.
Architectural use:
- separate durable workflow state from ephemeral process state
- make retries explicit
- avoid hidden nondeterminism in workflow logic
Human-AI Interaction
Microsoft Human-AI Interaction Guidelines
The Microsoft guidelines help turn “human in the loop” into concrete interaction requirements: timing, user control, feedback, correction, and expectation-setting.
Architectural use:
- show evidence before recommendation in high-risk flows
- expose uncertainty and correction paths
- monitor whether humans actually override the system
Google People + AI Guidebook
Google PAIR contributes the human-centered product lens. The book uses it to keep AI assistance aligned with user goals, feedback loops, and graceful failure.
Architectural use:
- design review queues around analyst work, not model vanity
- keep user correction as a first-class data source
- make failure states understandable
Observability
OpenTelemetry
OpenTelemetry contributes the tracing model and vocabulary for spans, attributes, and distributed context. The GenAI semantic conventions add useful names for model operations, prompts, completions, usage, and system attributes.
Architectural use:
- trace workflow stages, not only HTTP requests
- record prompt and model versions
- keep low-cardinality metrics separate from sensitive evidence
LLM Observability Tooling
Tools such as LangSmith and Phoenix show how practitioners trace model calls, retrieval, tool use, and evals. The book treats them as examples of a broader pattern rather than as mandatory dependencies.
Architectural use:
- record step-level behavior
- connect evals to traces
- preserve cost and latency per workflow
Security and Tooling
Model Context Protocol
MCP is useful because it makes tools and context explicit integration boundaries. That gives the book a concrete way to discuss tool contracts, authorization, and least privilege.
Architectural use:
- define tool scope
- separate read and write tools
- require authorization outside the model
- log tool calls as production actions
Prompt Injection Writing
Practical prompt-injection writing, especially by Simon Willison and the broader security community, shapes the book’s authority-separation model. The key lesson is that retrieved or user-provided text can be adversarial even when it looks like normal content.
Architectural use:
- label untrusted content
- block privileged tool paths
- avoid putting secrets in prompts
- test injection as a production regression case
Economics
FrugalGPT
FrugalGPT contributes the idea of cascades and model routing for cost-quality trade-offs. The book generalizes that idea into risk-aware routing.
Architectural use:
- use cheap deterministic paths first
- route ambiguous or high-risk cases upward
- measure quality and cost together
Provider Cost Documentation
OpenAI cost optimization, prompt caching, and batch API documentation show concrete provider mechanisms. The book uses them as examples, but keeps prices in fixtures because provider prices change.
Architectural use:
- make cost models data-driven
- separate architecture from current price sheets
- use caching and batching only when safe for the workflow
Distribution and Trust
Developer Documentation as Product
Stripe’s documentation is a benchmark for making complex technical products feel trustworthy. GitHub and GitLab materials provide useful patterns for repository trust, product positioning, and buyer communication.
Architectural use:
- turn eval reports into trust artifacts
- turn architecture diagrams into sales enablement
- make demos map to budget and risk
Category Design
Category-design writing is used cautiously. It is not technical evidence. It helps frame why “production AI systems architecture” is a better market category than generic AI automation.
Architectural use:
- define the old way and why it fails
- name the new evaluation criteria
- teach buyers how to recognize serious systems
Practitioner Pain Signals
Practitioner discussions are anecdotal. They are useful for discovering pain, not for proving claims.
Recurring signals:
- teams struggle to define useful LLM observability fields
- prompt injection becomes concrete once tools or private data enter the system
- RAG failures are hard to debug without retrieval traces
- cost uncertainty appears early and compounds with usage
- teams want evals but often lack a workflow-specific fixture discipline
Architectural use:
- prioritize executable examples
- label anecdotal material clearly
- connect pain signals back to authoritative standards and tested artifacts