OpenAI has released the Decisions API in public beta. It turns text and images into typed answers your code can branch on. OpenAI team states the OpenAI Decisions API runs about 10x faster than the Responses API. It targets a common pattern: prompt an LLM, then parse its text into a label.
TL;DR
- Size: GPT-6 Luna parameter count is not disclosed. Its model card lists a 1,050,000-token context window.
- Runs on: OpenAI-hosted API only, via
POST /v1/decisions. No open weights, no self-hosting. - Performance: About 10x faster than the Responses API, per OpenAI.
- Best: $0.10 per 1M input tokens, with no output, cache-read or cache-write charges.
- Bottom line:
- Best: fast, typed decisions with probabilities.
- Worst: 1 model, beta status, no independent evals yet.
- Best: fast, typed decisions with probabilities.
What is the OpenAI Decisions API?
The Decisions API is an OpenAI endpoint that evaluates text, images or both and returns typed answers. It does not generate prose. A request has 3 fields: model, input and questions. The only supported model today is gpt-6-luna. OpenAI expects general availability in the coming weeks.
How does the Decisions API work?
Each request carries shared evidence plus a list of questions. Each question has a unique name, a type and instructions. The response returns an answers array keyed by those names.
What are the 3 question types?
- predicate: checks a condition and returns a
probabilityfrom 0 to 1. Example: does a product photo show a crack, tear or dent? - choice: picks 1 value from options you supply. It also returns per-option
probabilitiesand aconfidencefield. - score: rates input against ordered
levels, indexed from 0. The score is a probability-weighted average of level indices.
OpenAI’s severity example makes the math concrete. Level probabilities of 0.1, 0.7 and 0.2 yield a score of 1.1. That value sits between ‘Workaround available’ and ‘Fully blocked’.
When should you use Structured Outputs instead?
OpenAI draws a clear line. Use Decisions for probabilities, choices or scores. Use Structured Outputs to fill your own JSON schema or write explanations. Use function calling when a model must request a tool call with arguments.
How fast is it, and what is the evidence?
OpenAI’s document claims about 10x faster responses than the Responses API. DevDay coverage put a decision near 150 ms, versus about 1.6 seconds for regular Luna calls. OpenAI has not published accuracy or calibration data for the endpoint. The docs advise setting thresholds with labeled examples from your own application.