OpenAI's Decisions API Is a Fast Route for Agents (2026)
OpenAI's Decisions API picks from a fixed answer list in ~150ms, about 10x faster than a standard model call. Put it before expensive AI workflows to route leads, block risky actions, and cut model spend. Test your own latency before promising sub-100ms to stakeholders.
OpenAI's Decisions API Is a Fast Route for Agents
OpenAI announced the Decisions API on September 29, 2026. It accepts text or images and returns one predefined answer.
That sounds boring.
Good. The best AI systems are boring.
Revenue teams don't need GPT-6 Astra to reason through every form submission. They need fast answers to small questions.
Is this lead qualified? Which rep owns it? Can this email be sent?
The Decisions API could answer those questions without writing an essay first.
Step 1: Separate Decisions From Generation
Most AI workflows send every task to one large model.
That's lazy architecture.
A lead enters HubSpot. The model reads it, scores it, writes a summary, picks a rep, and drafts an email.
You just paid a reasoning model to perform five different jobs.
OpenAI's Decisions API changes that pattern. You define a question and provide a closed list of answers.
A lead-routing request might include these choices:
- `sales_priority`
- `sales_standard`
- `nurture`
- `support`
- `spam`
- `manual_review`
The model must select one. It doesn't write free-form text.
OpenAI's DevDay demo showed roughly 150 milliseconds per decision. The regular Responses API took about 1.6 seconds.
That's about a 10x latency difference.
One third-party test reported 19.4-millisecond P50 latency and 38.6-millisecond P99 latency. OpenAI hasn't published an official service guarantee.
Don't sell your boss on "sub-100ms" yet. Test it in your own workflow first.
The historical parallel is the old rules engine.
Software teams didn't ask their main application to reconsider every business rule from scratch. They put a fast decision layer before the expensive work.
The Decisions API follows that pattern but uses model judgment.
It's also more useful than prompting a cheap model. A normal model can return bad JSON, extra commentary, or an invented category.
A constrained decision endpoint should return one allowed choice. Your code can act without cleaning up the answer.
Step 2: Build the Lead Router Before the Agent
Your lead router should sit directly after intake.
That includes Typeform submissions, inbound call transcripts, LinkedIn responses, website chat, and scraped intent signals.
Turn those sources into one record:
```json { "source": "website_form", "company_size": 240, "industry": "logistics", "country": "US", "message": "We need help qualifying 8,000 leads each month.", "consent": true, "existing_customer": false } ```
Run hard rules before any AI call.
Missing consent should trigger a block. Route known customers to account management.
Don't ask a model to decide facts your database already knows.
Send the remaining record to the OpenAI Decisions API. Ask separate questions for separate risks.
| Question | Allowed answers |
|---|---|
| Does this match our ICP? | `high`, `medium`, `low`, `unknown` |
| What's the buying intent? | `active`, `researching`, `none`, `unknown` |
| Is this likely spam? | `yes`, `no`, `uncertain` |
| What happens next? | `book`, `assign`, `enrich`, `nurture`, `review` |
Don't combine all four into one giant classification.
Small decisions are easier to test. They're also easier to fix.
Use n8n to receive the webhook and call the Decisions API. Then send the result to HubSpot, Salesforce, Apollo, Slack, or your calendar tool.
A high-fit lead with active intent can trigger enrichment and rep assignment. Put a medium-fit lead into nurture.
Send uncertain results to review.
That last route matters. A forced choice isn't automatically a correct choice.
Anthus tested regular GPT-6 Luna on 3,600 reasoning problems. When Luna reported 99% confidence, it was correct only 68% of the time.
That wasn't the specialized Decisions API. It still shows why you need to test confidence against real results.
Step 3: Add an AI Compliance Gate Before Action
Lead routing saves money. Stopping bad actions may matter more.
Put a second Decisions call before your agent sends anything.
This compliance gate should review the proposed action, not just the lead.
For an outbound email, use choices such as:
- `approve`
- `remove_claim`
- `remove_sensitive_data`
- `consent_required`
- `manual_review`
- `block`
Feed it the message, contact record, policy version, and channel. Don't send unrelated CRM history.
Keep hard rules outside the model.
Code should block a "do not contact" record. Send an unsupported guarantee, such as "we'll increase revenue by 40%," to review.
The model handles gray areas. Code handles non-negotiable rules.
Use the same pattern for call summaries, CRM updates, discounts, refunds, and campaign launches.
A sales agent shouldn't create a Salesforce opportunity without a QA check. A marketing agent shouldn't send 10,000 emails because one model chose the wrong tool.
Cialdini's work on influence starts with trust. Bad automation can destroy trust faster because it acts faster.
Many AI BDR tools get this backward. They automate sending before they control the decision.
StoryPros builds the control layer first. The writing model comes later.
Your fallback order should be simple:
1. Approve safe actions automatically. 2. Escalate uncertain actions to a stronger model. 3. Send high-risk cases to a person. 4. Fail closed when the API times out.
Never let an unavailable classifier trigger automatic approval.
Step 4: Measure Cost, Speed, and Wrong Decisions
OpenAI hasn't published Decisions API pricing.
Anyone quoting exact savings is guessing.
The supplied research also doesn't include verified Astra pricing. You can't calculate "Astra-level savings" without your own bill.
GPT-6 Luna's regular pricing was reported at $0.10 per million input tokens. Output was reported at $0.50 per million tokens.
OpenAI hasn't said whether Decisions uses those rates.
Use this formula instead:
> Monthly savings = avoided full-model cost − Decisions cost − review cost − monitoring cost
Track it by lead. Monthly API totals can hide bad routes.
Your dashboard should include:
| Metric | Why it matters |
|---|---|
| Routing latency at P50 and P99 | Averages hide slow failures |
| Cost per lead | Shows intake efficiency |
| Cost per qualified lead | Connects AI spend to sales value |
| Manual-review rate | Shows automation coverage |
| False-positive rate | Measures bad leads sent to sales |
| False-negative rate | Measures good leads incorrectly rejected |
| Rep override rate | Exposes routing mistakes |
| Lead-to-meeting rate | Shows business impact |
| Time to first action | Measures speed-to-lead |
| Compliance block rate | Shows prevented risk |
OpenAI's demo compared 150 milliseconds with 1.6 seconds. Four sequential decisions would take about 0.6 seconds instead of 6.4 seconds.
That gap matters for real-time forms and phone calls.
Cost routing also works outside revenue teams. Factory reported 63% lower inference costs from its production router.
Factory's routed tasks reached 99% of Claude Opus 4.7's Terminal-Bench 2 pass rate. Median session savings reached 72.5%.
Those results don't mean your lead router will save 63%.
They show that routing can cut costs when cheaper models handle routine tasks and stronger models handle hard cases.
Step 5: Log Every Decision and Launch Slowly
V1 won't be your final routing policy.
Treat the first release like a new sales hire. Watch it closely and correct it.
Start with one route. New website leads are a good choice.
Run the Decisions API beside your current process. Don't let it take action yet.
Compare its choices with real outcomes for two weeks. Track booked meetings, rejected leads, rep overrides, and compliance flags.
Your audit log needs these fields:
```text decision_id timestamp workflow_name policy_version source input_hash allowed_answers selected_answer confidence_if_available latency_ms downstream_action human_override final_outcome estimated_cost ```
Store an input hash when possible. Don't copy sensitive lead data into every log.
Version the decision policy.
If marketing changes the ICP from companies with 50 employees to 200, record that change. Otherwise, your historical accuracy numbers become useless.
OpenAI hasn't confirmed whether the Decisions API returns confidence scores. Treat reported confidence fields as unverified until OpenAI publishes documentation.
Build your workflow so confidence is optional.
If confidence exists, test it against labeled leads. Don't assume 90% means nine correct decisions out of ten.
Its call booking rate reached 72%. Booking improved by nearly 44 percentage points.
Those gains came from a larger system, not one classifier. Your Decisions API integration still needs to prove its own lift.
Turn on actions only after routing accuracy clears your target. Keep manual review for uncertain and high-risk cases.
FAQ
What is the OpenAI Decisions API?
The OpenAI Decisions API is a constrained classification and routing endpoint. It accepts text or images and selects one answer from a developer-defined list.
OpenAI announced it in limited preview on September 29, 2026. Public pricing and service guarantees weren't available at launch.
How does the Decisions API differ from regular LLM calls?
Regular LLM calls generate open-ended text. The Decisions API chooses from fixed answers such as `sales`, `support`, or `manual_review`.
That removes much of the parsing and retry code. It also makes downstream actions easier to control.
Can the Decisions API handle lead routing?
Yes. Lead routing matches OpenAI's stated use cases for classification and request routing.
A revenue team can classify ICP fit, intent, spam risk, and the next approved action. HubSpot or Salesforce can then receive the selected route through n8n.
Is the OpenAI Decisions API really sub-100ms?
OpenAI's launch materials showed about 150 milliseconds per decision. One third-party test reported a 19.4-millisecond P50 and 38.6-millisecond P99.
Sub-100ms performance may be possible. Don't promise it until your own production test proves it.
Should the Decisions API replace a stronger model?
No. Use it to decide when you need the stronger model.
Use Decisions for routing, gating, and basic classification. Reserve Astra-level reasoning for research, complex qualification, and high-value customer interactions.
Related Reading
How fast is the OpenAI Decisions API compared to a regular model call?
The Decisions API returned answers in roughly 150 milliseconds in OpenAI's launch demo. Regular Responses API calls took about 1.6 seconds. One third-party test reported a 19.4-millisecond P50 latency and 38.6-millisecond P99 latency.
Can I use the OpenAI Decisions API to route leads in HubSpot or Salesforce?
Yes, lead routing matches OpenAI's stated use cases for classification. A revenue team can ask four separate questions covering ICP fit, buying intent, spam risk, and next action. Results pass through n8n to HubSpot, Salesforce, or Apollo.
How much money can routing decisions through a cheaper model save on AI inference costs?
Factory reported 63% lower inference costs after deploying a production router that sent routine tasks to cheaper models. Routed tasks still reached 99% of Claude Opus 4.7's Terminal-Bench 2 pass rate. Median session savings reached 72.5%.