Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Typed Workflow Architecture

What You Already Know

You already know that AI systems are full of states: draft, extracted, reviewed, approved, rejected, retried, failed, escalated. You also know that production bugs often happen when the system permits a state that should not exist.

Typed workflow architecture is the discipline of making invalid states and transitions difficult to represent.

Before thinking about Rust syntax, think about language. A workflow already has nouns and verbs: case, document, analyst, evidence packet, upload, extract, approve, reject. Typed architecture asks which of those words are important enough that the system should not confuse them.

If two values would be dangerous to swap, they deserve different types. If a transition would be dangerous to perform casually, it deserves a named event. If a state would be embarrassing to explain to an auditor, it should probably be impossible or explicit.

The Failure Story

A case system stores this field:

status: string

The frontend sends approved. The backend sometimes writes approved_by_ai. A worker writes done. A migration adds ready_for_review, but one dashboard still filters for review_ready. The audit export maps unknown values to completed.

Nobody intended to weaken the compliance boundary. The type model allowed it.

The Core Concept

A production workflow needs two things:

  • a state model that defines legal states
  • an event model that defines legal transitions

In Rust-like pseudocode:

struct CaseId(String);
struct AnalystId(String);
struct DocumentId(String);

enum CaseStatus {
    Open,
    WaitingForDocuments,
    ReadyForAnalystReview,
    ApprovedByHuman,
    RejectedByHuman,
}

enum CaseEvent {
    DocumentUploaded(DocumentId),
    ExtractionSucceeded(DocumentId),
    MissingDataDetected,
    AnalystApproved(AnalystId),
    AnalystRejected(AnalystId),
}

The important part is not syntax. The important part is meaning. A case cannot be “kind of approved”. An approval event must name the analyst. The AI has no event variant that approves the case.

Transition Tables

A workflow should be explainable as a transition table:

Current stateEventNext state
OpenDocumentUploadedWaitingForDocuments
WaitingForDocumentsExtractionSucceededReadyForAnalystReview
WaitingForDocumentsMissingDataDetectedWaitingForDocuments
ReadyForAnalystReviewAnalystApprovedApprovedByHuman
ReadyForAnalystReviewAnalystRejectedRejectedByHuman

Everything not in the table is illegal.

This is regulatory armor. When someone asks whether the AI can approve a case, you can point to the model: there is no legal transition for it.

In implementation, the table becomes a transition function:

fn next_status(
    current: &CaseStatus,
    event: &CaseEvent,
) -> Result<CaseStatus, DomainError> {
    match (current, event) {
        (CaseStatus::Open, CaseEvent::DocumentUploaded(_)) => {
            Ok(CaseStatus::WaitingForDocuments)
        }
        (CaseStatus::ReadyForAnalystReview, CaseEvent::AnalystApproved(_)) => {
            Ok(CaseStatus::ApprovedByHuman)
        }
        _ => Err(DomainError::IllegalTransition),
    }
}

The exact code can vary. The invariant cannot: illegal transitions return typed errors instead of disappearing into fallback branches.

Newtypes

Newtypes protect meaning. String is not a case ID, analyst ID, document ID, model name, tenant ID, prompt version, or idempotency key. It is a storage representation.

Bad boundary:

fn approve(case_id: String, actor_id: String) {}

Better boundary:

fn approve(case_id: CaseId, analyst_id: AnalystId) {}

The second function prevents accidental swaps and makes the domain visible. If the type has validation rules, use a smart constructor:

impl CaseId {
    pub fn new(value: impl Into<String>) -> Result<Self, CaseIdError> {
        let value = value.into();
        if value.trim().is_empty() {
            return Err(CaseIdError::Empty);
        }
        Ok(Self(value))
    }
}

The invariant belongs at construction time, not scattered through every call site.

Operation Lifecycle vs Audit Lifecycle

AI systems often confuse two lifecycles.

The operation lifecycle is about work execution:

queued -> running -> retrying -> succeeded -> failed

The audit lifecycle is about business meaning:

open -> waiting_for_documents -> ready_for_review -> approved_by_human

A model extraction job can fail and retry without changing the case’s business status. A human approval changes the audit lifecycle. Keep those separate.

This separation prevents accidental designs where “worker succeeded” becomes “case approved”.

Idempotency

Production systems retry. Networks fail. Webhooks repeat. Workers crash after writing to the database but before acknowledging a queue message.

Idempotency means the same intended operation can be safely applied more than once without duplicating business effects.

For a document upload:

idempotency_key =
  stable_hash(tenant_id, case_id, document_hash, operation_type)

If the same upload request arrives twice, the system should return the existing document result, not create two evidence records and two extraction jobs.

Idempotency is not an optimization. It is part of correctness.

Build idempotency keys from canonical typed fields with stable encoding. Then enforce uniqueness where the business effect is stored, usually with a database constraint such as unique(tenant_id, idempotency_key).

Outbox Pattern

When a state change must publish an event, do not write the database and publish to the queue as unrelated actions. If the database commit succeeds and the queue publish fails, the system is split.

The transactional outbox pattern solves this:

  1. update the business state
  2. write an outbox row in the same transaction
  3. a publisher process reads unsent outbox rows
  4. publish with retries
  5. mark the outbox row as sent

This gives the system a durable record of work that must happen.

For AI workflows, outbox events often include:

  • ExtractionRequested
  • EvidencePacketReady
  • AnalystReviewRequested
  • EvaluationFailed
  • AuditPacketFinalized

Sources to Pair With This Chapter

Worked Example: Illegal Approval

The example crate in examples/rust-workflows encodes a small case workflow. The key test is:

#[test]
fn ai_cannot_approve_without_human_review_state() -> Result<(), DomainError> {
    let case_id = CaseId::new("case-002")?;
    let analyst_id = AnalystId::new("analyst-001")?;
    let mut case = Case::open(case_id);

    match case.apply(CaseEvent::AnalystApproved(analyst_id)) {
        Err(DomainError::IllegalTransition {
            from: CaseStatus::Open,
            event: CaseEvent::AnalystApproved(_),
        }) => {}
        other => panic!("expected illegal transition, got {other:?}"),
    }
    Ok(())
}

The domain does not need a prompt that says “do not approve from Open”. The transition function rejects the state change.

This is the deeper design move: use types and transitions to protect what the prompt should never own.

Persistence Boundary

Typed workflow architecture does not mean the database disappears. The database should enforce the same truth:

  • constrained status values
  • foreign keys for real relationships
  • unique idempotency keys
  • append-only audit records where required
  • timestamps for state changes
  • tenant-scoped indexes

Do not let the application say one thing and the database permit another.

A minimal relational mirror might include:

create table cases (
  tenant_id text not null,
  case_id text not null,
  status text not null check (
    status in ('open', 'waiting_for_documents', 'ready_for_review',
               'approved_by_human', 'rejected_by_human')
  ),
  primary key (tenant_id, case_id)
);

create table case_events (
  tenant_id text not null,
  case_id text not null,
  event_id text not null,
  event_type text not null,
  occurred_at timestamptz not null,
  primary key (tenant_id, event_id)
);

The database does not replace the domain model. It prevents a second truth from forming underneath it.

Minimum Artifact

By the end of this chapter, produce a transition specification. It should include:

  • domain entities and their newtypes
  • legal statuses
  • legal events
  • a transition table
  • forbidden transitions
  • actor required for each sensitive transition
  • idempotency keys for retryable operations
  • database constraints that mirror the domain model

If the transition cannot be checked in code or schema, it is still only policy language.

Common Mistakes

The first mistake is representing meaningful states as loose strings. Loose strings are convenient until every integration invents a synonym.

The second mistake is using booleans for lifecycle state. approved: true cannot say who approved, from what prior state, with what evidence, or after which review.

The third mistake is letting external DTOs leak into domain logic. Provider responses, HTTP payloads, database rows, and domain models should be separate. Convert at boundaries.

The fourth mistake is adding types without enforcing transitions. Newtypes help, but the workflow also needs a transition model.

Self-Check

  1. Why is status: string dangerous in regulated workflows?
  2. What is the difference between operation lifecycle and audit lifecycle?
  3. Why does idempotency matter for AI jobs?
  4. How does the outbox pattern reduce split-brain workflow failures?

Retrieval Practice

Recall:

  • Name three domain concepts that should not cross boundaries as raw strings.

Explain:

  • Explain why a worker job succeeding should not automatically mean a case is approved.

Apply:

  • Draw a transition table for an AI workflow you want to build. Mark every transition that requires a human actor.

Where This Leaves Us

Typed workflows tell the system what states are legal. The next question is who is allowed to move the workflow through sensitive transitions. That is human-in-the-loop design.