← All articles

OpenAI Decisions API: Typed Classification, Routing, and Scoring

OpenAI’s Decisions API returns typed answers for classification, routing, and scoring. Here’s how the beta works, what it costs, and where human review still matters.

The OpenAI Decisions API turns text or images into bounded answers your application can act on: a yes-or-no-style predicate, a choice from options you define, or a score against a rubric. That makes it worth evaluating when an app needs a quick, structured judgment—such as routing a support request, checking a product photo, or prioritizing a queue—without asking a general-purpose model to write a full response.

OpenAI introduced the API in beta on October 6, 2026. Its current developer guide documents a dedicated POST /v1/decisions endpoint powered by gpt-6-luna. The video linked here is OpenAI’s short Decisions API introduction; the implementation details below are checked against OpenAI’s written documentation, which is the better reference when you build against the API.

Decisions API at a glance

  • Stage: public beta, according to the current guide.
  • Model: gpt-6-luna is the only model currently documented for the endpoint.
  • Input: shared text evidence or user messages containing text and inline images.
  • Question types: predicate, choice, and score.
  • Endpoint: POST https://api.openai.com/v1/decisions.
  • Best fit: bounded decisions with explicit answer types and a known set of outcomes.

OpenAI says the endpoint returns answers about 10 times faster than the Responses API. That is OpenAI’s comparison, not an independent benchmark. Treat it as a reason to run your own latency test—not as a guarantee for your prompts, traffic, or application.

What the API actually returns

You provide evidence once, then define one or more questions against that evidence. Each question has a type, instructions, and—where relevant—a set of allowed choices or score levels. The response contains an answers array; give questions unique names so your code can identify each result.

Predicate: estimate whether a condition is true

A predicate is useful when your application wants an estimated probability for a condition, such as “Does this photo appear to show a damaged item?” The result is an estimate from 0 to 1. It is not a verified fact or an automatic authorization to take the next step.

Choice: select from options you supply

A choice question selects one of your predefined values, such as billing, technical_support, or human_review. The API also returns probabilities over the discrete options and a confidence field. Keep an explicit fallback option for ambiguous or unsupported cases.

Score: apply an ordered rubric

A score question rates the input against ordered levels—for example, low, medium, and high severity. The API calculates a probability-weighted average of the level indices, so its numeric score can land between named levels. Your application still needs to define what those levels mean and what action, if any, each score should trigger.

A small routing example

Imagine an online service receiving short messages from customers. You could ask the API to choose among a few destinations rather than produce an open-ended category string:

{
  "model": "gpt-6-luna",
  "input": "My invoice has a duplicate charge and I need help.",
  "questions": [
    {
      "type": "choice",
      "name": "queue",
      "instructions": "Choose the best support queue. Use human_review if the message is unclear or does not fit.",
      "choices": [
        { "value": "billing", "description": "Invoices, payments, or charges." },
        { "value": "technical_support", "description": "Product errors or setup problems." },
        { "value": "human_review", "description": "Unclear, sensitive, or unsupported requests." }
      ]
    }
  ]
}

Your service—not the model—would then validate that queue is an allowed destination, record the result if appropriate, and route the message according to your access and escalation rules. For refunds, account changes, or other consequential actions, add policy checks and human approval rather than treating a model answer as permission.

When to use Decisions instead of Responses

Use Decisions when the application needs one of its documented answer types: a predicate estimate, a fixed choice, or a score against ordered levels. The constrained shape is useful for routing and triage because your code can reason about a known set of outcomes.

Use the Responses API with Structured Outputs when you need a custom JSON object—such as several extracted fields plus an explanation. Use function calling when the model should request a tool call with arguments. These are different jobs: Decisions returns an answer for your software to interpret; it does not itself call your business tools or complete a workflow.

Design thresholds, fallbacks, and review

The probability and confidence values are signals to evaluate, not proof that a decision is correct. OpenAI’s guide recommends testing with labeled examples from your application and choosing thresholds based on the cost of false positives and false negatives.

Before routing live traffic, build a small evaluation set that reflects the language, image quality, edge cases, and ambiguous inputs your users actually produce. Measure errors by outcome, not only aggregate accuracy. Decide what happens when confidence is low, an answer is refused, the response is malformed, or the model returns a valid but unexpected option.

For a higher-impact workflow, use a conservative sequence:

  1. Let the API propose a typed answer.
  2. Check that answer against your application’s allowed values and business rules.
  3. Apply normal identity, authorization, and data-validation checks.
  4. Send uncertain or consequential cases to a person.
  5. Monitor outcomes and adjust thresholds using reviewed examples.

This is especially important for access decisions, eligibility, safety reviews, payments, and other contexts where a wrong result has real consequences. A model-generated probability is not a substitute for a policy, a runtime permission check, or a responsible appeal path.

Input and deployment constraints to check

The API reference documents text strings or user messages containing text and inline images. Its image inputs must be supplied as data URLs; external image URLs and file IDs are not supported by this endpoint. The reference also lists limits on supported message roles and content types, so teams should verify these constraints before designing an ingestion pipeline.

OpenAI lists the beta’s gpt-6-luna input price at $0.10 per million tokens. The Decisions guide says there are no cache-read, cache-write, or output-token charges for this endpoint, while regional processing premiums and long-context input multipliers can apply. Check current pricing and your organization’s terms before projecting costs; prices and beta behavior can change.

The guide also says Zero Data Retention and HIPAA use are supported for eligible customers, with data residency and regional processing supported in the United States and Europe (EEA + Switzerland). Eligibility, agreements, and limitations apply. Confirm the current data controls documentation with your compliance team before sending sensitive material.

Is it useful for your application?

The Decisions API is a focused tool for developers who need typed judgments rather than long-form generated text. Its most promising early uses are bounded: classify, select, or score—then let application code enforce the rules and decide what happens next.

The API is still in public beta and currently tied to one model. Start with a low-risk workflow, keep a human-review path, and compare it with your existing classifier or Responses-based implementation on representative examples. OpenAI’s speed claim may matter, but quality, error costs, input handling, privacy requirements, and fallback behavior should determine whether the design fits your product.

If you’re adding AI routing or decision support to a customer-facing product, AI consulting and application development can help define the evaluation set, permission boundaries, and review flow before the model output reaches users. The AI Web Awards platform is an example of FindMilan’s work building an AI-powered web application.

Sources and further reading

FAQ

Frequently asked questions

What is the OpenAI Decisions API?

The Decisions API is a public-beta endpoint that evaluates text and images against developer-defined predicate, choice, or score questions and returns typed answers. OpenAI currently documents gpt-6-luna as its only supported model.

How is the Decisions API different from Structured Outputs?

Decisions is designed for bounded predicate, choice, and scoring questions. OpenAI recommends Structured Outputs with the Responses API when an application needs a custom JSON object, such as extracted fields or a written explanation.

How much does the Decisions API cost?

OpenAI currently lists input pricing of $0.10 per million tokens for gpt-6-luna through /v1/decisions, with regional-processing premiums and long-context multipliers. The guide says there are no cache-read, cache-write, or output-token charges for this endpoint; check the live pricing docs before estimating costs.

Does the Decisions API perform the selected action?

No. It returns an answer for your application to interpret. Your code must validate the result, apply permissions, handle uncertainty, and decide whether to perform an action or ask a person to review it.

Is the Decisions API generally available?

No. OpenAI’s current guide describes it as a public beta and says gpt-6-luna is the only supported model. Availability and terms may change as the API evolves.

Need help with AI consulting and application development?

Turn the idea into a working system.