AI Economics
What You Already Know
You already know that model calls cost money. The production skill is deeper: cost must be modeled per workflow, not per API call.
AI economics answers:
Can usage grow while quality, latency, and gross margin remain acceptable?
The Failure Story
A startup automates document review. Early customers love it. Usage grows. The system sends every document chunk to the most expensive model, retries failures without limits, stores no cache, performs synchronous analysis even for non-urgent tasks, and routes easy cases through the same pipeline as hard cases.
Revenue grows. Gross margin collapses.
The team did not build an AI company. It built a pass-through payment mechanism for inference.
The Core Concept
Model cost is an architectural constraint.
Track cost at these levels:
| Level | Question |
|---|---|
| API call | what did this model request cost? |
| task | what did extraction or classification cost? |
| workflow | what did one completed case cost? |
| tenant | which customer drives cost? |
| feature | which product surface consumes margin? |
| outcome | what is the cost per successful business result? |
The important unit is usually cost per successful workflow.
Margin Is a Design Constraint
Economics is not a spreadsheet after launch. It constrains architecture while you design.
If a workflow costs more when it succeeds than the customer pays for the successful outcome, quality improvements can make the business worse. If high-risk cases cost more because they require more review, that may be correct. The goal is not to minimize every cost. The goal is to spend money where risk and value justify it.
Ask:
What cost should increase with risk?
What cost should decrease with scale?
What cost should disappear because deterministic code can do the job?
This is the difference between cheap architecture and economically coherent architecture.
Cost Formula
A rough workflow cost:
workflow_cost =
retrieval_cost
+ model_input_tokens * input_price
+ model_output_tokens * output_price
+ tool_costs
+ storage_costs
+ human_review_minutes * labor_cost
+ retry_cost
Then compare it to workflow value:
gross_margin_per_workflow =
revenue_per_workflow - workflow_cost
If the model saves one hour of analyst time but adds five minutes of review and a few cents of inference, it may be excellent. If it automates a low-value task with expensive frontier calls, it may be structurally bad.
Sensitivity Analysis
The cost model should say which variables can break the business case:
| Variable | What can go wrong | Design response |
|---|---|---|
| input tokens | long documents dominate cost | chunk, summarize, cache policy text, and extract deterministically where possible |
| output tokens | verbose answers waste margin | use structured outputs and bounded note templates |
| frontier-model share | every case takes the expensive route | route by risk, confidence, and value |
| retry rate | tool loops multiply spend | cap retries and classify retry causes |
| review minutes | AI creates extra human work | measure review time, correction rate, and rework |
| tenant mix | one customer drives loss | monitor tenant-level cost and price high-complexity workflows explicitly |
This is where economics becomes architecture. If one variable can destroy margin, the system needs a control, not a hope.
Sources to Pair With This Chapter
- OpenAI, Cost optimization: use for provider-level cost control tactics.
- OpenAI, Prompt caching: use for caching economics and constraints.
- OpenAI, Batch API: use for async batch cost patterns.
- Chen et al., FrugalGPT: use for model cascades and cost-quality trade-offs.
- AWS, Optimizing costs of generative AI applications: use for cloud-level cost architecture.
- Reddit practitioner discussion, production LLM service pain: use only as anecdotal signal around deterministic preprocessing, evals, observability, cost, and latency.
Model Routing
Do not route every task to the strongest model.
Use a routing matrix:
| Task | Default model | Escalate when |
|---|---|---|
| schema cleanup | small model or deterministic code | schema conflict |
| classification | small or medium model | low confidence or high risk |
| risk note drafting | medium model | complex evidence |
| legal-sensitive summary | frontier model | always, plus review |
| final decision | not model-owned | human transition |
FrugalGPT’s core lesson is practical: cascades and model selection can reduce cost while preserving or improving performance. The production version is not only “use cheaper models”. It is “route by risk, uncertainty, and value”.
Cache Strategy
Caching can save money, but unsafe caching can leak data or preserve stale reasoning.
| Cache target | Good candidate? | Risk |
|---|---|---|
| static policy text | yes | version invalidation |
| public reference material | yes | source freshness |
| tenant-specific evidence | sometimes | tenant leakage |
| model output for case decision | rarely | stale or unaudited state |
| embeddings | yes with versioning | model and corpus drift |
| prompt prefix | yes | prompt version mismatch |
OpenAI’s prompt caching and batch APIs show a broader point: provider features can change the economics, but architecture must decide when those features are safe.
Batch and Async Processing
Not every AI task needs a synchronous answer.
Use synchronous processing when:
- a human is waiting
- interaction quality depends on immediacy
- the task is small and bounded
Use asynchronous processing when:
- documents are large
- retries are likely
- the result feeds a later review
- batching reduces cost
- the workflow has a natural queue
Many enterprise AI workflows are better as durable jobs than chat-style request-response interactions.
When Not to Use an LLM
The cheapest, safest model call is the one you do not make.
Do not use an LLM for:
- exact arithmetic
- stable rule checks
- schema validation
- permission decisions
- deterministic transformations
- simple keyword filters
- final authority decisions in regulated workflows
Use code, database constraints, search, rules, or smaller models when they are enough.
This is not anti-AI. It is systems discipline.
Worked Example: Case Cost
A case-preparation workflow has:
- document OCR: fixed provider cost
- extraction: medium model
- missing-document check: deterministic rules
- risk note: frontier model only for high-risk cases
- review: analyst minutes
- audit packet: deterministic assembly
Low-risk case:
OCR + medium extraction + deterministic checks + sampled review
High-risk case:
OCR + medium extraction + frontier risk note + mandatory analyst review + second review
The high-risk case costs more because it should. Cost follows risk.
Runnable Example
This repository includes a workflow cost calculator:
python3 examples/economics/cost_model.py \
fixtures/economics/case_review_cost_model.json
The fixture does not pretend provider prices are permanent. It keeps prices in data and validates the architecture-level budget: p50 cost, p95 cost, retry budget, and frontier-model cost share. This lets a team change model prices without changing the calculator.
The lesson is operational: every workflow should have a cost model before adoption makes cost variance painful.
Cost SLOs
Add cost objectives:
- p50 cost per case
- p95 cost per case
- maximum retry spend per workflow
- maximum frontier-model percentage
- cost per successful extraction
- cost per analyst-approved packet
- tenant spend anomaly threshold
Cost SLOs belong beside latency and reliability SLOs.
Minimum Artifact
By the end of this chapter, produce a workflow cost model. It should include:
- cost per task
- cost per completed workflow
- p50 and p95 cost
- retry budget
- frontier-model share
- human review cost
- cache and batch assumptions
- revenue or value per workflow
- margin sensitivity when usage grows
If cost is only measured per API call, the architecture cannot yet defend product margin.
Common Mistakes
The first mistake is measuring token cost without human review cost. Human oversight is part of the workflow economics.
The second mistake is optimizing for the cheapest model before measuring risk. Cheap wrong decisions are expensive.
The third mistake is failing to cap retries. A broken tool loop can become a cost incident.
The fourth mistake is pricing the product before understanding cost variance. High-risk customers may generate higher review and inference cost.
Self-Check
- Why is cost per workflow more useful than cost per API call?
- How should risk influence model routing?
- When is caching dangerous?
- Which parts of your AI workflow should be deterministic instead of model-driven?
Retrieval Practice
Recall:
- Write the rough workflow cost formula.
Explain:
- Explain why usage growth can make an AI product worse as a business.
Apply:
- Pick one workflow. Divide every step into deterministic code, small model, frontier model, human review, or async batch.
Where This Leaves Us
AI economics keeps the system viable. The final pillar asks how serious architecture becomes visible to the market. For technical products, distribution is not separate from engineering. It is how trust compounds.