The OpenAI API now includes Decisions in public beta, giving developers a dedicated way to make fast, structured decisions inside applications without asking a general-purpose model to generate a full text response. Powered by GPT-6 Luna, the new API can evaluate text and images and return a probability, a predefined choice or a score that software can immediately use for routing, filtering or automation.
The release could become an important infrastructure layer for AI applications. Instead of sending every request through a large model and then parsing its natural-language response, developers can define the exact decisions an application needs and receive typed answers directly. OpenAI says the Decisions API can make these decisions up to 10 times faster than using GPT-6 Luna through the Responses API.
Source and verification note (7 October 2026): OpenAI’s 6 October public-beta announcement, API changelog and developer guide establish the release date, endpoint, question types and list price. Independent original coverage and product examinations by Dealroom, Creative AI News and WorkerKit corroborate the public beta. The ‘up to 10x faster’ line is OpenAI’s performance claim; no comparable independent benchmark for this release was identified.
Key takeaways
- OpenAI released the Decisions API in public beta on October 6, 2026.
- It runs on GPT-6 Luna, with that model currently the only supported option.
- Developers can send text, images or both as input.
- The API supports three decision types: predicates, choices and scores.
- Applications can use the results to classify content, route requests, select tools or determine an agent’s next action.
- OpenAI says Decisions API is up to 10x faster than using GPT-6 Luna through the Responses API.
- The dedicated endpoint is
POST /v1/decisions. - Pricing for GPT-6 Luna through Decisions API is $0.10 per 1 million input tokens, with no charges for output tokens or cache reads and writes.
- OpenAI expects the API to reach general availability in the coming weeks.
OpenAI API Decisions: what developers receive
The simplest way to understand Decisions API is to think of it as a decision-making layer between raw application data and software actions.
A traditional AI application might send a customer message to a large language model and ask:
“Read this complaint and tell me which department should handle it.”
The model might respond:
“This appears to be a billing-related issue involving a duplicate transaction, so I recommend sending it to the billing department.”
The application then has to interpret that answer and convert it into a machine-readable action.
Decisions API removes much of that intermediate step.
A developer can define a question such as:
“Which department should handle this complaint?”
and provide fixed choices such as:
- Billing
- Technical support
- Shipping
- Account
- Other
The API returns the selected choice along with probabilities and confidence information.
The software can then immediately route the request.
OpenAI describes this as a way to turn text and images into decisions that an application can use directly.
Why OpenAI built a separate decision API
Large language models are designed to perform many different tasks.
That flexibility is useful when an application needs reasoning, generation, coding or conversation. But it can be inefficient when the software needs only a narrow decision.
Consider an online retailer receiving 10 million customer messages.
Most of those messages might need to be classified into a small number of categories.
Running every message through a full conversational workflow and generating text for each one creates unnecessary computation and additional application logic.
A dedicated decision endpoint can instead ask the model to answer a finite set of questions.
That makes the model’s output predictable enough for software to consume directly.
OpenAI says Decisions API is designed specifically for this type of workflow and can operate up to 10 times faster than sending the same type of decision-making task through the Responses API.
Three types of decisions
Decisions API currently supports three main output types.
1. Predicates
A predicate asks whether a particular condition is true.
For example:
Question: Does this product photograph show visible damage?
The API could return a probability such as:
Damaged: 0.95
The application can then establish its own threshold.
For example, a retailer could automatically approve a review when the probability exceeds a particular level and send uncertain cases to a human employee.
This makes predicates useful for classification and automated screening.
2. Choices
A choice question asks the model to select from a predefined list.
For example:
Customer message: “I was charged twice for the same order.”
Possible choices:
- Billing
- Technical support
- Shipping
- Account
- Other
The API selects the most appropriate option and returns probabilities for the available choices, along with a confidence value.
This is particularly useful for application routing.
A company’s customer-support system could use the result to decide which queue should receive a ticket.
3. Scores
The third option is scoring.
Developers can define ordered levels and ask the model to evaluate an input against those levels.
For example, a customer complaint could be assigned:
- Low severity
- Medium severity
- High severity
The API returns a score and probability information associated with the defined levels.
That could allow an application to prioritise cases automatically.
OpenAI’s documentation says score results can fall between defined levels because the final score can be probability-weighted.
Real-time app routing is one of the biggest use cases
The routing capability is particularly important as software becomes more dependent on multiple AI models and agents.
Imagine an application receiving millions of requests.
Some requests may be simple enough for a lightweight model.
Others may require a more capable reasoning model.
Some may need a specialised image model.
Others may need a tool or autonomous agent.
A routing system could use Decisions API to determine which destination is appropriate.
For example:
Incoming request → Decisions API → selected model/tool/agent → final response
The decision itself does not need to generate a conversational answer.
It simply determines what happens next.
OpenAI specifically positions Decisions API for choosing the right model, tool or action in near real time.
It could become an important layer for AI agents
The timing of the release is significant because AI agents are becoming more sophisticated.
Agents increasingly have access to tools, browsers, databases, APIs and other software.
That creates a recurring problem: what should the agent do next?
A general-purpose model can make that decision, but developers may prefer a more constrained mechanism for repetitive routing decisions.
Decisions API could sit inside the agent architecture as a lightweight decision layer.
For example:
User request → classify intent → select agent → choose tool → execute task
The first step could be handled by Decisions API.
An application might ask:
“Does this request require financial data?”
If the answer is yes, the application routes the request to a finance-specific agent.
If not, it sends the request elsewhere.
OpenAI’s DevDay announcement described Decisions API as a way to classify content, route requests and choose an agent’s next action.
Text and images are supported
The API is not restricted to text.
Developers can provide text, images or a combination of both.
That expands the potential applications considerably.
A retail platform could analyse a product photograph and determine whether an item appears damaged.
A manufacturing system could classify an image from an inspection camera.
A support application could combine a customer’s written complaint with an uploaded screenshot.
A moderation system could evaluate both written content and images.
The current API documentation says image inputs must be supplied as inline base64 data URLs, rather than external image URLs or file IDs. A request can contain up to 128 image parts.
Multiple questions can be evaluated together
Developers do not necessarily have to send one request for every decision.
Independent questions can share the same input.
For example, an ecommerce application could send a product image and simultaneously ask:
- Is the product damaged?
- What category does the product belong to?
- How severe is the visible damage?
The API can return answers for the questions in order.
This could reduce the need for multiple separate model calls when decisions are based on the same evidence.
OpenAI recommends keeping independent questions separate while defining clear, observable criteria for each question.
Developers still need to set the thresholds
One important detail is that Decisions API does not automatically determine how an application should act on every probability or confidence value.
Developers still have to decide what constitutes an acceptable threshold.
Suppose an AI system returns:
Fraud probability: 0.82
Should the transaction automatically be blocked?
That depends on the business.
A bank might prefer to send the transaction to a human review team.
Another system might require a probability above 0.95 before taking automatic action.
OpenAI recommends using labelled examples from the application to establish appropriate thresholds and considering the relative costs of false positives and false negatives.
This is an important distinction because a model’s confidence is not the same thing as a guarantee of correctness.
Decisions API could reduce unnecessary LLM work
The broader opportunity is efficiency.
Large language models are increasingly being embedded in software workflows, but many applications do not need a long generated response at every stage.
They need a small structured decision.
For example:
“Is this relevant?”
“Which queue?”
“Which tool?”
“Which model?”
“How severe?”
These are narrow questions.
By returning typed answers rather than generated prose, Decisions API is designed to make these steps faster and easier to integrate into conventional software.
OpenAI’s API changelog specifically describes the endpoint as turning text and images into typed answers and says it is 10 times faster than the Responses API for this use case.
Pricing is designed for high-volume decisions
OpenAI is also positioning the API for workloads where large numbers of small decisions may be required.
For GPT-6 Luna through Decisions API, input costs $0.10 per 1 million tokens.
OpenAI says developers pay only for input tokens, with no separate charges for cache reads, cache writes or output tokens on this endpoint. Regional processing premiums and long-context pricing multipliers can still apply.
That pricing structure could make the API attractive for high-volume classification and routing workloads.
The economics will ultimately depend on how much input each decision consumes and how many decisions an application needs to make.
OpenAI is entering a new AI category
The Decisions API launch also has a competitive dimension.
The product arrives shortly after TypeSafe AI launched Jev, a decision-focused AI model designed specifically for structured decisions rather than conventional text generation.
The Wall Street Journal reported that Jev’s launch has attracted significant interest and prompted other companies, including OpenAI and Databricks, to release similar decision-oriented products.
Databricks has introduced an ai_decide function, while other companies are also exploring models designed around fast, structured decisions.
This suggests a new layer of the AI stack may be emerging.
Instead of treating every AI request as a conversation with a general-purpose LLM, developers could increasingly use specialised decision models for the routing and classification steps around larger AI systems.
Decisions API is not a replacement for the Responses API
The new endpoint should not be viewed as a universal replacement for OpenAI’s main Responses API.
Responses remains the better fit when an application needs generated text, complex reasoning, tool calls or a broader agent workflow.
Decisions API is much narrower.
Its advantage comes from limiting the problem.
Developers define the questions and the possible answers in advance. The model then evaluates the supplied evidence against those requirements.
For applications that need open-ended generation, this constrained design would not be appropriate.
Public beta means the product can still change
Decisions API is currently in public beta.
OpenAI says general availability is expected in the coming weeks.
That means developers can use the API now, but the interface, supported capabilities, pricing or behaviour could still change before the final release.
The API was initially introduced at OpenAI’s DevDay on September 29 in limited preview before being opened to all developers in public beta on October 6.
For more on this OpenAI API update, read our separate reports on new API usage tiers and OpenAI’s mathematics research release. Neither is the same product as the Decisions endpoint.
The bigger picture
The launch points to a subtle shift in how AI applications are being designed.
The first generation of AI software largely revolved around sending a prompt to a large language model and displaying its answer.
The next generation is likely to contain many specialised AI calls around a central application.
One model may generate the answer. Another may classify the request. Another may decide which tool should be used. An agent may then execute the task.
Decisions API is designed for one of those intermediate layers.
If the approach works at scale, developers could treat AI decisions more like conventional software primitives: a structured input goes in, a predictable typed decision comes out, and the rest of the application takes over.
Looking Ahead
The immediate test for Decisions API will be whether developers can use it reliably enough for production routing without introducing costly classification errors. Speed alone will not determine its success; accuracy, calibration, predictable behaviour and the economics of running millions of decisions will matter just as much.
Longer term, the API could become part of the infrastructure connecting models and agents. If AI applications increasingly operate as networks of specialised models and tools, fast decision-making between those components could become a critical layer—and OpenAI is now positioning GPT-6 Luna to provide it.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



