Research and References
This page organizes source material by chapter. It is not decorative bibliography; it is the source map for the book.
Cross-Cutting Governance and Risk
- NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0)
- NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile
- European Union, Regulation (EU) 2024/1689, the AI Act
- ISO, ISO/IEC 42001 AI management system
- OWASP GenAI Security Project, Top 10 Risk and Mitigations for LLMs and Gen AI Apps 2025
Evaluation
- Stanford CRFM, Holistic Evaluation of Language Models (HELM)
- Liang et al., Holistic Evaluation of Language Models
- Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- OpenAI, Evals repository
- OpenAI Cookbook, Evaluate model performance
- RAGAS, RAGAS: Automated Evaluation of Retrieval Augmented Generation
Typed Workflow Architecture
- Martin Fowler, Domain Event
- Martin Fowler, Event Sourcing
- Chris Richardson, Transactional Outbox pattern
- Chris Richardson, Idempotent Consumer pattern
- Temporal, Durable execution
- Rust Design Patterns, Newtype
- Cliff L. Biffle, The Typestate Pattern in Rust
Human-in-the-Loop
- Microsoft Research, Guidelines for Human-AI Interaction
- Amershi et al., Guidelines for Human-AI Interaction
- Google PAIR, People + AI Guidebook
- European Union, AI Act Article 14: human oversight
- NIST AI RMF, Govern, Map, Measure, Manage functions
Observability
- OpenTelemetry, Semantic conventions for generative AI systems
- OpenTelemetry, Traces
- LangSmith, LLM application observability
- Arize Phoenix, LLM tracing and evaluation
- OpenAI Agents SDK, Tracing
Security and Governance
- OWASP GenAI Security Project, LLM Top 10 2025
- NIST, AI RMF Generative AI Profile
- Model Context Protocol, Specification 2025-11-25
- Model Context Protocol, Authorization 2025-11-25
- Simon Willison, Prompt injection writing and examples
- OpenAI, Safety best practices
AI Economics
- OpenAI, Cost optimization
- OpenAI, Prompt caching
- OpenAI, Batch API
- Chen et al., FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- AWS, Optimizing costs of generative AI applications
Agents, Tools, and Workflow Actors
- Anthropic, Building effective agents
- OpenAI, Agents SDK documentation
- Model Context Protocol, Tools 2025-11-25
- Lilian Weng, LLM-powered autonomous agents
Distribution and Trust
- GitHub, Open source guides
- GitLab, Handbook: product marketing
- Stripe, Docs as product experience
- Common Room, Developer relations and community-led growth writing
- Category Pirates, Category design writing
Practitioner Pain Signals
The book also uses practitioner pain from recurring public discussions: unreliable demos, unclear evals, hidden AI costs, prompt injection, RAG hallucinations, brittle agent loops, missing auditability, and stakeholder mistrust. Treat these as design pressure, not as primary evidence. Primary claims should rest on the sources above.
- Reddit, LLM observability fields in real traffic: anecdotal signal that teams need prompt, output, token, cost, latency, and step metadata.
- Reddit, LLM observability platform suggestions: anecdotal signal around cost uncertainty and in-house observability.
- Reddit, Prompt injection in self-hosted LLM deployment: anecdotal signal that prompt injection becomes a production blocker.
- Reddit, RAG observability discussion: anecdotal signal that retrieval failure diagnosis is a recurring production pain.
- Reddit, Production LLM service pain: anecdotal signal around deterministic preprocessing, evaluation, observability, cost, and latency.
- Reddit, LLM system evals discussion: anecdotal signal that teams struggle to balance controlled evals, observability, and production KPIs.
Use these discussions as evidence of pain, vocabulary, and field pressure. Do not use them as authoritative proof for safety, legal, or architectural claims.