OpenAI Decision API: Fast AI Routing With Typed Answers
OpenAI’s Decision API is a public beta for fast, typed decisions over text and images. Here’s what it does, how it’s priced, where it helps, and where it still has limits.
The Open AI Decision API is OpenAI’s new public beta for turning text and image inputs into typed answers that applications can use for routing, classification, scoring, and other fast decision flows. OpenAI says the API is built for cases where you want a model to answer a narrow question with a predictable structure, and it positions the endpoint as being much faster than the Responses API for these tasks. The official docs also say the Decision API is powered by gpt-6-luna and exposed through the dedicated POST /v1/decisions endpoint. (developers.openai.com)
For developers, this matters because many AI products do not need a long free-form answer. They need a yes/no confidence, a choice among labeled options, or a numeric score. That is exactly the shape of the Decision API. OpenAI describes three answer types: predicates for probability that something is true, choices for selecting among predefined options, and scores for evaluating an input against a range or rubric. (developers.openai.com)
What the Decision API is for
The strongest use cases are the ones where speed and consistency matter more than open-ended generation:
- Content moderation and safety triage.
- Request routing to a human team or a downstream model.
- Lead, ticket, or document classification.
- Fraud or risk scoring when you have labeled decision criteria.
- Review workflows where a model should produce a confidence-bearing decision, not a paragraph.
OpenAI says you can use labeled examples from your own application to choose thresholds for routing, filtering, or review. That is a good reminder that the model output is not the final policy; your product logic still decides what to do with the score or probability. (developers.openai.com)
Key features
1. Typed answers
Unlike a generic text-generation call, the Decisions API returns a structured decision artifact. The docs state that predicate answers return an estimated probability, while choice and score answers also include a confidence field and probability distribution. That makes the output easier to feed into app logic than a raw text completion. (developers.openai.com)
2. Multimodal input
The API can evaluate text, images, or both, which is useful for tasks such as product image checks, document screening, or multimodal policy review. OpenAI’s docs explicitly call out support for text and image inputs. (developers.openai.com)
3. Dedicated low-latency endpoint
OpenAI says the API is designed to be about 10x faster than the Responses API for these decision tasks. I would read that as a product claim about the API’s intended workload rather than a universal benchmark for every app, because real latency depends on network, payload size, and your integration. Still, the direction is clear: this endpoint is optimized for fast decisioning. (developers.openai.com)
4. Practical thresholding
The docs recommend using labeled examples to tune thresholds based on the cost of false positives and false negatives. That is the right mental model for production: do not treat the model score as “truth,” treat it as an input to your own decision boundary. (developers.openai.com)
5. Minimal SDK friction
OpenAI’s guide shows support in the current SDKs, with the docs listing minimum versions for Python, JavaScript, Go, Ruby, and Java. If you are already using one of those SDKs, adding the Decisions API should look familiar. (developers.openai.com)
Pricing
As of the official docs, the Decisions API uses gpt-6-luna and charges $0.10 per 1M input tokens. OpenAI says you pay only for input tokens on /v1/decisions; there are no cache-read, cache-write, or output-token charges for that endpoint. The docs also note that regional processing premiums and long-context multipliers can apply, so your effective price may vary. (developers.openai.com)
OpenAI’s model page for gpt-6-luna shows the broader model pricing context as well, including standard input, cached input, cache writes, and output pricing for the model outside the Decisions endpoint. But for this article, the important part is that Decisions API pricing is specialized and token economics differ from ordinary text generation. (developers.openai.com)
Limitations and caveats
The API is useful, but it is not magic. The official docs and model pages make a few constraints clear:
- Public beta: OpenAI says the Decisions API is in public beta and expects to GA in the coming weeks, so the surface area may still change. (developers.openai.com)
- Only one model today: The docs say
gpt-6-lunais the only model currently available for Decisions. (developers.openai.com) - Not a free-form writing tool: The feature is aimed at narrow decisions, not long-form generation. If you need prose, summaries, or tool orchestration, the Responses API is the better fit. (developers.openai.com)
- Thresholds are your responsibility: OpenAI gives probabilities and confidence signals, but your app still has to decide where to draw the line. (developers.openai.com)
- Pricing complexity: Regional processing and long-context multipliers can affect cost, so you should confirm the current pricing page before budgeting. (developers.openai.com)
There is also an important engineering limitation to remember: the output is only as good as your schema, labels, and decision policy. If your training or examples are noisy, the routing will be noisy too. That is not a limitation unique to OpenAI; it is a general property of classification systems.
Example: routing a support ticket
Below is a simple JavaScript example showing how you might use a decision call to route a ticket to support, billing, or engineering. This is a conceptual example based on the typed-answer pattern in the docs; adapt field names to the current SDK and API reference.
import OpenAI from "openai";
const client = new OpenAI({ apiKey: process.env.OPENAI_API_KEY });
async function routeTicket(ticketText) {
const response = await client.decisions.create({
model: "gpt-6-luna",
input: [
{
role: "user",
content: `Classify this ticket into support, billing, or engineering:\n\n${ticketText}`
}
],
questions: [
{
name: "ticket_route",
type: "choice",
choices: ["support", "billing", "engineering"]
}
]
});
const answer = response.answers.find((a) => a.name === "ticket_route");
return {
route: answer.choice,
confidence: answer.confidence,
distribution: answer.distribution
};
}
(async () => {
const result = await routeTicket(
"I was charged twice for the same plan and need a refund."
);
console.log(result);
})();In production, you would usually combine that result with a threshold. For example, if the model is less than 0.80 confident, you might send the ticket to manual review instead of auto-routing it. That is exactly the kind of workflow the OpenAI docs encourage when they discuss labeled examples and threshold selection. (developers.openai.com)
When to use it, and when not to
Use the Open AI Decision API when the application needs a fast, structured decision and the acceptable outputs are already known. Avoid it when you need detailed reasoning, a conversational back-and-forth, or a long narrative answer. In those cases, the Responses API or structured output patterns are still the better fit. OpenAI’s structured outputs docs also remind developers that if the goal is to structure a model response, Structured Outputs is a separate capability from decisions. (developers.openai.com)
Bottom line
The Open AI Decision API is a focused product for a very common product need: fast AI routing with typed answers. Its appeal is not general intelligence; it is operational simplicity. You get a narrower interface, structured answers, input-only pricing on the decision endpoint, and a design that fits classification and scoring tasks well. If your app lives on top of routing, triage, and review logic, this is a feature worth testing. If you need open-ended generation, keep using the broader Responses API and structured outputs where appropriate. (developers.openai.com)