Typed Workflow Architecture
What You Already Know
You already know that AI systems are full of states: draft, extracted, reviewed, approved, rejected, retried, failed, escalated. You also know that production bugs often happen when the system permits a state that should not exist.
Typed workflow architecture is the discipline of making invalid states and transitions difficult to represent.
Before thinking about Rust syntax, think about language. A workflow already has nouns and verbs: case, document, analyst, evidence packet, upload, extract, approve, reject. Typed architecture asks which of those words are important enough that the system should not confuse them.
If two values would be dangerous to swap, they deserve different types. If a transition would be dangerous to perform casually, it deserves a named event. If a state would be embarrassing to explain to an auditor, it should probably be impossible or explicit.
The Failure Story
A case system stores this field:
status: string
The frontend sends approved. The backend sometimes writes approved_by_ai. A worker writes done. A migration adds ready_for_review, but one dashboard still filters for review_ready. The audit export maps unknown values to completed.
Nobody intended to weaken the compliance boundary. The type model allowed it.
The Core Concept
A production workflow needs two things:
- a state model that defines legal states
- an event model that defines legal transitions
In Rust-like pseudocode:
struct CaseId(String);
struct AnalystId(String);
struct DocumentId(String);
enum CaseStatus {
Open,
WaitingForDocuments,
ReadyForAnalystReview,
ApprovedByHuman,
RejectedByHuman,
}
enum CaseEvent {
DocumentUploaded(DocumentId),
ExtractionSucceeded(DocumentId),
MissingDataDetected,
AnalystApproved(AnalystId),
AnalystRejected(AnalystId),
}
The important part is not syntax. The important part is meaning. A case cannot be “kind of approved”. An approval event must name the analyst. The AI has no event variant that approves the case.
Transition Tables
A workflow should be explainable as a transition table:
| Current state | Event | Next state |
|---|---|---|
| Open | DocumentUploaded | WaitingForDocuments |
| WaitingForDocuments | ExtractionSucceeded | ReadyForAnalystReview |
| WaitingForDocuments | MissingDataDetected | WaitingForDocuments |
| ReadyForAnalystReview | AnalystApproved | ApprovedByHuman |
| ReadyForAnalystReview | AnalystRejected | RejectedByHuman |
Everything not in the table is illegal.
This is regulatory armor. When someone asks whether the AI can approve a case, you can point to the model: there is no legal transition for it.
In implementation, the table becomes a transition function:
fn next_status(
current: &CaseStatus,
event: &CaseEvent,
) -> Result<CaseStatus, DomainError> {
match (current, event) {
(CaseStatus::Open, CaseEvent::DocumentUploaded(_)) => {
Ok(CaseStatus::WaitingForDocuments)
}
(CaseStatus::ReadyForAnalystReview, CaseEvent::AnalystApproved(_)) => {
Ok(CaseStatus::ApprovedByHuman)
}
_ => Err(DomainError::IllegalTransition),
}
}
The exact code can vary. The invariant cannot: illegal transitions return typed errors instead of disappearing into fallback branches.
Newtypes
Newtypes protect meaning. String is not a case ID, analyst ID, document ID, model name, tenant ID, prompt version, or idempotency key. It is a storage representation.
Bad boundary:
fn approve(case_id: String, actor_id: String) {}
Better boundary:
fn approve(case_id: CaseId, analyst_id: AnalystId) {}
The second function prevents accidental swaps and makes the domain visible. If the type has validation rules, use a smart constructor:
impl CaseId {
pub fn new(value: impl Into<String>) -> Result<Self, CaseIdError> {
let value = value.into();
if value.trim().is_empty() {
return Err(CaseIdError::Empty);
}
Ok(Self(value))
}
}
The invariant belongs at construction time, not scattered through every call site.
Operation Lifecycle vs Audit Lifecycle
AI systems often confuse two lifecycles.
The operation lifecycle is about work execution:
queued -> running -> retrying -> succeeded -> failed
The audit lifecycle is about business meaning:
open -> waiting_for_documents -> ready_for_review -> approved_by_human
A model extraction job can fail and retry without changing the case’s business status. A human approval changes the audit lifecycle. Keep those separate.
This separation prevents accidental designs where “worker succeeded” becomes “case approved”.
Idempotency
Production systems retry. Networks fail. Webhooks repeat. Workers crash after writing to the database but before acknowledging a queue message.
Idempotency means the same intended operation can be safely applied more than once without duplicating business effects.
For a document upload:
idempotency_key =
stable_hash(tenant_id, case_id, document_hash, operation_type)
If the same upload request arrives twice, the system should return the existing document result, not create two evidence records and two extraction jobs.
Idempotency is not an optimization. It is part of correctness.
Build idempotency keys from canonical typed fields with stable encoding. Then enforce uniqueness where the business effect is stored, usually with a database constraint such as unique(tenant_id, idempotency_key).
Outbox Pattern
When a state change must publish an event, do not write the database and publish to the queue as unrelated actions. If the database commit succeeds and the queue publish fails, the system is split.
The transactional outbox pattern solves this:
- update the business state
- write an outbox row in the same transaction
- a publisher process reads unsent outbox rows
- publish with retries
- mark the outbox row as sent
This gives the system a durable record of work that must happen.
For AI workflows, outbox events often include:
ExtractionRequestedEvidencePacketReadyAnalystReviewRequestedEvaluationFailedAuditPacketFinalized
Sources to Pair With This Chapter
- Martin Fowler, Domain Event: use for business-significant events.
- Martin Fowler, Event Sourcing: use for append-only state reconstruction and auditability.
- Chris Richardson, Transactional Outbox: use for reliable state-change publication.
- Chris Richardson, Idempotent Consumer: use for retry-safe consumers.
- Temporal, Workflow documentation: use for durable workflow and determinism constraints.
- Rust Design Patterns, Newtype: use for type-level domain meaning.
Worked Example: Illegal Approval
The example crate in examples/rust-workflows encodes a small case workflow. The key test is:
#[test]
fn ai_cannot_approve_without_human_review_state() -> Result<(), DomainError> {
let case_id = CaseId::new("case-002")?;
let analyst_id = AnalystId::new("analyst-001")?;
let mut case = Case::open(case_id);
match case.apply(CaseEvent::AnalystApproved(analyst_id)) {
Err(DomainError::IllegalTransition {
from: CaseStatus::Open,
event: CaseEvent::AnalystApproved(_),
}) => {}
other => panic!("expected illegal transition, got {other:?}"),
}
Ok(())
}
The domain does not need a prompt that says “do not approve from Open”. The transition function rejects the state change.
This is the deeper design move: use types and transitions to protect what the prompt should never own.
Persistence Boundary
Typed workflow architecture does not mean the database disappears. The database should enforce the same truth:
- constrained status values
- foreign keys for real relationships
- unique idempotency keys
- append-only audit records where required
- timestamps for state changes
- tenant-scoped indexes
Do not let the application say one thing and the database permit another.
A minimal relational mirror might include:
create table cases (
tenant_id text not null,
case_id text not null,
status text not null check (
status in ('open', 'waiting_for_documents', 'ready_for_review',
'approved_by_human', 'rejected_by_human')
),
primary key (tenant_id, case_id)
);
create table case_events (
tenant_id text not null,
case_id text not null,
event_id text not null,
event_type text not null,
occurred_at timestamptz not null,
primary key (tenant_id, event_id)
);
The database does not replace the domain model. It prevents a second truth from forming underneath it.
Minimum Artifact
By the end of this chapter, produce a transition specification. It should include:
- domain entities and their newtypes
- legal statuses
- legal events
- a transition table
- forbidden transitions
- actor required for each sensitive transition
- idempotency keys for retryable operations
- database constraints that mirror the domain model
If the transition cannot be checked in code or schema, it is still only policy language.
Common Mistakes
The first mistake is representing meaningful states as loose strings. Loose strings are convenient until every integration invents a synonym.
The second mistake is using booleans for lifecycle state. approved: true cannot say who approved, from what prior state, with what evidence, or after which review.
The third mistake is letting external DTOs leak into domain logic. Provider responses, HTTP payloads, database rows, and domain models should be separate. Convert at boundaries.
The fourth mistake is adding types without enforcing transitions. Newtypes help, but the workflow also needs a transition model.
Self-Check
- Why is
status: stringdangerous in regulated workflows? - What is the difference between operation lifecycle and audit lifecycle?
- Why does idempotency matter for AI jobs?
- How does the outbox pattern reduce split-brain workflow failures?
Retrieval Practice
Recall:
- Name three domain concepts that should not cross boundaries as raw strings.
Explain:
- Explain why a worker job succeeding should not automatically mean a case is approved.
Apply:
- Draw a transition table for an AI workflow you want to build. Mark every transition that requires a human actor.
Where This Leaves Us
Typed workflows tell the system what states are legal. The next question is who is allowed to move the workflow through sensitive transitions. That is human-in-the-loop design.