Stop Paying the Orchestration Tax (2026)

Matt Payne··Updated ·8 min read
Key Takeaway

Route high-volume AI agent jobs to Haiku 5.5 at $0.10 per million input tokens instead of Sonnet at $2.00. A 1M-job pipeline drops from $13,000 to $2,752 per month. Build a small-model decision layer, cache repeated prompts, and escalate only on evidence.

Stop Paying the Orchestration Tax

Most agent builders use one expensive model for every task.

That's like asking your attorney to sort the mail.

Anthropic built Claude Haiku 5.5 for classification, support, summaries, database queries, browser work, and coding subagents. Anthropic says it costs 75% less than Haiku 4.5 across typical workloads.

Build your pipeline around a small-model decision layer.

Step 1: Find Where You're Paying Frontier Rates for Clerk Work

Start with 7 to 14 days of agent logs.

Track every model call by task, model, input tokens, output tokens, cache status, latency, and result. If your system can't show those fields, fix that before you change routing.

Use five task buckets:

1. Classification 2. Extraction 3. Generation 4. Tool use 5. Complex reasoning

Lead qualification usually starts with classification and extraction. Support triage is classification. Content QA often uses a structured checklist.

None of those jobs needs Opus by default.

Anthropic prices Haiku 5.5 requests under 100,000 tokens at:

Charge per 1M tokensHaiku 5.5Sonnet 5.5
Input$0.10$2.00
Output$0.50$10.00
Cache read$0.01$0.10
Cache write$0.125$2.50

Haiku input costs 95% less than Sonnet input at that tier. Haiku output also costs 95% less.

That gap is your orchestration tax.

Your first report should show:

  • Cost per completed task
  • Cost per failed task
  • Cost by task type
  • Escalation rate
  • Cache hit rate
  • P50 and P95 latency
  • Human correction rate
  • Tool-call failure rate

Don't measure cost per API call alone.

A cheap call that produces junk is expensive. A $0.02 call that resolves a ticket beats a $0.002 call that creates three retries.

Step 2: Put Haiku in Front of Your Expensive Models

The small-model decision layer has one job.

It decides what happens next.

For a sales agent, Haiku can classify a lead as qualified, unqualified, duplicate, missing data, or human review. For support, it can select billing, technical, cancellation, account access, or escalation.

Keep the output boring.

{ "route": "billing", "complexity": 2, "confidence": 0.94, "requires_human": false, "contains_sensitive_data": false, "reason_code": "invoice_question" }

Don't ask the router to write the final answer. Don't let it call 20 tools. Don't give it a personality.

That turns simple routing into another expensive agent.

A useful routing policy looks like this:

if contains_sensitive_data: route = "human_review"

elif confidence < 0.80: route = "sonnet"

elif complexity >= 4: route = "sonnet"

elif task_type in ["classification", "extraction", "triage"]: route = "haiku"

else: route = "haiku_with_validation"

The model should return a fixed schema. Your code should reject anything outside that schema.

Put limits around Haiku's decisions.

Cloudflare's Auto Router scores requests by complexity, stakes, prior context, and task type. Its internal tests put routed work at 35% of Claude Opus 5.5's cost with similar daily-work performance.

Factory takes the same approach. Its router sends routine jobs to cheaper models and escalates harder work.

Factory reported 63% lower inference costs across production sessions. Routed runs reached 99% of Claude Opus 4.7's pass rate on Terminal-Bench 2.

Routing beats loyalty to one model.

Step 3: Cache the Parts That Don't Change

Cache-aware model routing prices the entire path before choosing a model.

That includes prompt writes, prompt reads, context length, and the cost of switching models.

Claude prompt caching works on repeated prompt prefixes. The stable content must appear first.

Put these items at the front:

1. System instructions 2. Policy rules 3. Tool definitions 4. Output schema 5. Product or support documentation 6. Examples

Put changing content last:

1. Lead record 2. Support ticket 3. User message 4. Tool result 5. Current date or account state

Changing one sentence near the top can kill the cache match.

Use two cache layers.

Layer one is provider prompt caching. Anthropic stores the processed prompt prefix. A later request with the same prefix pays the cache-read rate.

Layer two is your application cache. Redis can store completed results for repeatable jobs, such as duplicate checks or standard policy questions.

A practical cache key looks like this:

sha256( tenant_id + task_type + policy_version + schema_version + normalized_input + model_family )

Never share a key across customers. Include the policy version so old rules don't survive a policy change.

TTL depends on the job:

Cached itemSuggested TTL
Lead duplicate check24 hours
Product classification7 days
Support policy answer1 hour
Account status1–5 minutes
Agent prompt prefixProvider TTL
Human approval stateDon't cache

Anthropic's standard prompt-cache window has historically needed close attention to idle gaps. Research on agent keepalives found that a request around four minutes can preserve a five-minute cache without wasteful pings every 30 seconds.

Don't blindly add keepalives.

Use them when tool calls or human approvals regularly cross the cache window. A quiet agent shouldn't spend money pretending to be busy.

Step 4: Escalate on Evidence, Not Vibes

A router needs a clear escape hatch.

Send the task to Sonnet or Opus when Haiku sees high stakes, low confidence, conflicting records, long context, or an unsupported tool request.

The flow should look like this:

Request | v Result Cache | +-- Hit --> Return | v Haiku Decision Layer | +-- Simple --> Haiku Worker --> Validator | +-- Hard ----> Sonnet or Opus --> Validator | +-- Risky ---> Human Review | v Action + Audit Log

The validator matters more than the router.

For lead qualification, compare the output against CRM fields. For support, check the cited policy section. For content QA, run fixed checks for links, dates, banned phrases, and required sections.

People say AI hallucinates when they skip validation.

Your billing system also returns bad results when you skip validation. Nobody says Stripe had a dream.

Set hard escalation rules:

  • Confidence below 0.80
  • Two source records disagree
  • Required field is missing
  • Tool call fails twice
  • Output breaks the JSON schema
  • Refund exceeds a fixed dollar limit
  • Account has a legal or security flag

Run 500 to 1,000 historical tasks before changing live traffic. Compare Haiku's route with the known outcome.

Track false negatives separately.

A false positive support escalation costs time. A false negative fraud classification can cost far more.

Your pass mark must match the job.

Content-tag classification might ship at 95% agreement. Refund approval may require 100% rule compliance and human sign-off.

Step 5: Prove the Savings With Real Cost Math

Assume an agent handles 1 million jobs each month.

Each job uses 4,000 input tokens and 500 output tokens. Running all jobs on Sonnet 5.5 would cost:

  • Input: 4 billion tokens × $2 per million = $8,000
  • Output: 500 million tokens × $10 per million = $5,000
  • Total: $13,000 per month

Now route 85% to Haiku and 15% to Sonnet.

Haiku handles 850,000 jobs:

  • Input: 3.4 billion tokens × $0.10 = $340
  • Output: 425 million tokens × $0.50 = $212.50

Sonnet handles 150,000 jobs:

  • Input: 600 million tokens × $2 = $1,200
  • Output: 75 million tokens × $10 = $750

Add a Haiku routing call using 2,000 input tokens and 100 output tokens per job:

  • Router input: 2 billion tokens × $0.10 = $200
  • Router output: 100 million tokens × $0.50 = $50

The routed total is $2,752.50.

That's 78.8% less than the $13,000 Sonnet-only pipeline.

Prompt caching can cut it further.

If 2,000 of each job's 4,000 input tokens come from a repeated Haiku prefix, those tokens cost $0.01 per million on a cache hit instead of $0.10. Sonnet cache reads cost $0.10 instead of $2.00.

Anthropic also cut Sonnet 5.5 cache reads from $0.20 to $0.10 per million tokens. Anthropic says that lowers typical Sonnet agent costs by about 20%.

Your real result depends on cache hits, escalation rate, and retry volume.

Track cost per correct completed action.

Don't Skip the Boring Controls

Cheap agents can create expensive mistakes faster.

Use separate API keys for development and production. Set daily token limits by workflow, not just by account.

AWS offers Haiku 5.5 through Amazon Bedrock with IAM, CloudTrail, CloudWatch, regional data controls, and Bedrock Guardrails. Those controls matter when your agent touches customer records or support history.

Log these fields for every call:

  • Request ID
  • Customer ID
  • Workflow version
  • Prompt version
  • Selected model
  • Route reason
  • Input and output tokens
  • Cache read and write tokens
  • Latency
  • Validation result
  • Final action
  • Human override

Redact passwords, API keys, payment data, and private customer fields before logging.

Set a weekly routing review.

Look for rising escalation rates, falling cache hits, schema failures, and one workflow eating the token budget. Model updates can change results without changes to your code.

V1 will hit 60% to 70% of the target.

That's normal. Treat the router like a new hire.

Give it examples. Review mistakes. Tighten the rules. Retest after every model or prompt change.

The best AI systems don't need to impress anyone in a demo.

They stop sending $0.10 jobs to a $10 model.

FAQ

What is cache-aware model routing?

Cache-aware model routing chooses an AI model based on task difficulty and the full request cost. That cost includes input tokens, output tokens, prompt-cache reads, cache writes, and the cost of switching models during a long session.

How does model routing for AI agents reduce LLM costs?

A small model classifies each job and sends routine work to the cheapest model that can complete it. Factory reported 63% lower production inference costs. The example in this guide cuts monthly model spend from $13,000 to $2,752.50.

How do I set up agent routing with Claude Haiku 5.5?

Send each request to Haiku 5.5 with a fixed JSON schema for route, complexity, confidence, and risk. Use code-based rules to send low-confidence or high-stakes tasks to Sonnet, Opus, or a human reviewer.

How much can prompt caching cut AI agent costs?

Claude Haiku 5.5 cache reads cost $0.01 per million tokens, compared with $0.10 for standard input under 100,000 tokens. Sonnet 5.5 cache reads cost $0.10, compared with $2.00 for standard input.

Should every AI agent use a small-model decision layer?

High-volume agents that handle mixed work should use one. Lead qualification, support triage, content QA, summaries, and database lookups are good candidates because most requests are narrow and repeatable.

Related Reading

AI Answer

How much cheaper is Claude Haiku 5.5 than Sonnet 5.5?

Haiku 5.5 input tokens cost $0.10 per million, compared to $2.00 for Sonnet 5.5. That is a 95% price difference. Cache reads drop Haiku input to $0.01 per million tokens.

AI Answer

How much money can model routing save on a 1 million job per month AI agent?

Routing 85% of jobs to Haiku and 15% to Sonnet cuts monthly costs from $13,000 to $2,752.50. That is a 78.8% reduction. Adding prompt caching cuts the bill further.

AI Answer

What confidence score should trigger escalation from Haiku to Sonnet?

Send the task to Sonnet when Haiku returns a confidence score below 0.80. Hard escalation rules also apply when required fields are missing, tool calls fail twice, or output breaks the JSON schema.