OpenAI Decisions API: Luna-Powered Choices in 150ms
thenewstack.io

OpenAI Decisions API: Luna-Powered Choices in 150ms

Tech News
3 min read

Published by AINave Editorial

TL;DROpenAI’s Decisions API is designed to choose from developer-defined answers and return a confidence score, rather than generate chat. OpenAI says results take 150 milliseconds, but pricing and other rollout details remain unclear.

OpenAI’s Decisions API turns a familiar prompt-based task into a more constrained endpoint: developers supply questions, answers and context, and the system returns a predefined choice with a confidence score. Announced at the company’s DevDay and based on Luna, OpenAI’s smallest and most affordable model in its current lineup, the API was available in limited preview when The New Stack reported on it on September 29, 2026. OpenAI said broader rollout was planned for the coming days, but that report does not establish whether it happened. OpenAI’s announcement and preview details

A middle ground between prompting and training

Teams often ask a chat model to choose from a list and try to infer confidence from token probabilities. The New Stack describes those estimates as rough. At the other end, a small classifier can be fast and cheap, but requires labeled data and retraining when the set of labels changes. The reported trade-offs between prompted models and classifiers

A decision model sits between those approaches: developers can provide new labels in the request, while receiving a structured choice and score rather than prose. The distinction matters when the task is bounded, such as sorting content, routing a request or selecting an agent’s next action. OpenAI’s Moderation API also returns scores, but its categories are set by OpenAI; the Decisions API lets developers supply the answers. The proposed uses and comparison with Moderation

Fast, according to OpenAI, but not yet fully specified

OpenAI says the Decisions API returns results in 150 milliseconds, compared with 1.6 seconds for GPT-6 Luna. Those are company-reported figures; the available report does not provide a benchmark methodology, so they should be read as a stated comparison rather than an independently established performance result. OpenAI’s reported latency figures

That speed claim is only one part of the decision for teams evaluating the endpoint. The New Stack reported that per-call pricing, the number of candidate answers a request can handle and whether developers can tune the system on their own data were still unclear. Those details affect whether the API is practical for a particular workflow, especially where the choice set is large or changes frequently. The unresolved pricing, candidate-limit and tuning details

The useful distinction is not simply that this is another model endpoint. It gives developers a way to express a changing set of choices without relying on chat-style output or maintaining a classifier for every label set. Whether that convenience outweighs the still-unknown operating constraints depends on what the broad release includes.

FAQs

It returns a choice from developer-provided answers with a confidence score instead of generating chat-style prose. Developers provide the questions, answers and context. The API’s described input and output

Sources

Latest Tech News