Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

AI Economics

What You Already Know

You already know that model calls cost money. The production skill is deeper: cost must be modeled per workflow, not per API call.

AI economics answers:

Can usage grow while quality, latency, and gross margin remain acceptable?

The Failure Story

A startup automates document review. Early customers love it. Usage grows. The system sends every document chunk to the most expensive model, retries failures without limits, stores no cache, performs synchronous analysis even for non-urgent tasks, and routes easy cases through the same pipeline as hard cases.

Revenue grows. Gross margin collapses.

The team did not build an AI company. It built a pass-through payment mechanism for inference.

The Core Concept

Model cost is an architectural constraint.

Track cost at these levels:

LevelQuestion
API callwhat did this model request cost?
taskwhat did extraction or classification cost?
workflowwhat did one completed case cost?
tenantwhich customer drives cost?
featurewhich product surface consumes margin?
outcomewhat is the cost per successful business result?

The important unit is usually cost per successful workflow.

Margin Is a Design Constraint

Economics is not a spreadsheet after launch. It constrains architecture while you design.

If a workflow costs more when it succeeds than the customer pays for the successful outcome, quality improvements can make the business worse. If high-risk cases cost more because they require more review, that may be correct. The goal is not to minimize every cost. The goal is to spend money where risk and value justify it.

Ask:

What cost should increase with risk?
What cost should decrease with scale?
What cost should disappear because deterministic code can do the job?

This is the difference between cheap architecture and economically coherent architecture.

Cost Formula

A rough workflow cost:

workflow_cost =
  retrieval_cost
  + model_input_tokens * input_price
  + model_output_tokens * output_price
  + tool_costs
  + storage_costs
  + human_review_minutes * labor_cost
  + retry_cost

Then compare it to workflow value:

gross_margin_per_workflow =
  revenue_per_workflow - workflow_cost

If the model saves one hour of analyst time but adds five minutes of review and a few cents of inference, it may be excellent. If it automates a low-value task with expensive frontier calls, it may be structurally bad.

Sensitivity Analysis

The cost model should say which variables can break the business case:

VariableWhat can go wrongDesign response
input tokenslong documents dominate costchunk, summarize, cache policy text, and extract deterministically where possible
output tokensverbose answers waste marginuse structured outputs and bounded note templates
frontier-model shareevery case takes the expensive routeroute by risk, confidence, and value
retry ratetool loops multiply spendcap retries and classify retry causes
review minutesAI creates extra human workmeasure review time, correction rate, and rework
tenant mixone customer drives lossmonitor tenant-level cost and price high-complexity workflows explicitly

This is where economics becomes architecture. If one variable can destroy margin, the system needs a control, not a hope.

Sources to Pair With This Chapter

Model Routing

Do not route every task to the strongest model.

Use a routing matrix:

TaskDefault modelEscalate when
schema cleanupsmall model or deterministic codeschema conflict
classificationsmall or medium modellow confidence or high risk
risk note draftingmedium modelcomplex evidence
legal-sensitive summaryfrontier modelalways, plus review
final decisionnot model-ownedhuman transition

FrugalGPT’s core lesson is practical: cascades and model selection can reduce cost while preserving or improving performance. The production version is not only “use cheaper models”. It is “route by risk, uncertainty, and value”.

Cache Strategy

Caching can save money, but unsafe caching can leak data or preserve stale reasoning.

Cache targetGood candidate?Risk
static policy textyesversion invalidation
public reference materialyessource freshness
tenant-specific evidencesometimestenant leakage
model output for case decisionrarelystale or unaudited state
embeddingsyes with versioningmodel and corpus drift
prompt prefixyesprompt version mismatch

OpenAI’s prompt caching and batch APIs show a broader point: provider features can change the economics, but architecture must decide when those features are safe.

Batch and Async Processing

Not every AI task needs a synchronous answer.

Use synchronous processing when:

  • a human is waiting
  • interaction quality depends on immediacy
  • the task is small and bounded

Use asynchronous processing when:

  • documents are large
  • retries are likely
  • the result feeds a later review
  • batching reduces cost
  • the workflow has a natural queue

Many enterprise AI workflows are better as durable jobs than chat-style request-response interactions.

When Not to Use an LLM

The cheapest, safest model call is the one you do not make.

Do not use an LLM for:

  • exact arithmetic
  • stable rule checks
  • schema validation
  • permission decisions
  • deterministic transformations
  • simple keyword filters
  • final authority decisions in regulated workflows

Use code, database constraints, search, rules, or smaller models when they are enough.

This is not anti-AI. It is systems discipline.

Worked Example: Case Cost

A case-preparation workflow has:

  • document OCR: fixed provider cost
  • extraction: medium model
  • missing-document check: deterministic rules
  • risk note: frontier model only for high-risk cases
  • review: analyst minutes
  • audit packet: deterministic assembly

Low-risk case:

OCR + medium extraction + deterministic checks + sampled review

High-risk case:

OCR + medium extraction + frontier risk note + mandatory analyst review + second review

The high-risk case costs more because it should. Cost follows risk.

Runnable Example

This repository includes a workflow cost calculator:

python3 examples/economics/cost_model.py \
  fixtures/economics/case_review_cost_model.json

The fixture does not pretend provider prices are permanent. It keeps prices in data and validates the architecture-level budget: p50 cost, p95 cost, retry budget, and frontier-model cost share. This lets a team change model prices without changing the calculator.

The lesson is operational: every workflow should have a cost model before adoption makes cost variance painful.

Cost SLOs

Add cost objectives:

  • p50 cost per case
  • p95 cost per case
  • maximum retry spend per workflow
  • maximum frontier-model percentage
  • cost per successful extraction
  • cost per analyst-approved packet
  • tenant spend anomaly threshold

Cost SLOs belong beside latency and reliability SLOs.

Minimum Artifact

By the end of this chapter, produce a workflow cost model. It should include:

  • cost per task
  • cost per completed workflow
  • p50 and p95 cost
  • retry budget
  • frontier-model share
  • human review cost
  • cache and batch assumptions
  • revenue or value per workflow
  • margin sensitivity when usage grows

If cost is only measured per API call, the architecture cannot yet defend product margin.

Common Mistakes

The first mistake is measuring token cost without human review cost. Human oversight is part of the workflow economics.

The second mistake is optimizing for the cheapest model before measuring risk. Cheap wrong decisions are expensive.

The third mistake is failing to cap retries. A broken tool loop can become a cost incident.

The fourth mistake is pricing the product before understanding cost variance. High-risk customers may generate higher review and inference cost.

Self-Check

  1. Why is cost per workflow more useful than cost per API call?
  2. How should risk influence model routing?
  3. When is caching dangerous?
  4. Which parts of your AI workflow should be deterministic instead of model-driven?

Retrieval Practice

Recall:

  • Write the rough workflow cost formula.

Explain:

  • Explain why usage growth can make an AI product worse as a business.

Apply:

  • Pick one workflow. Divide every step into deterministic code, small model, frontier model, human review, or async batch.

Where This Leaves Us

AI economics keeps the system viable. The final pillar asks how serious architecture becomes visible to the market. For technical products, distribution is not separate from engineering. It is how trust compounds.