OpenAI Launches Decisions API for Faster GPT-6 Luna Routing
OpenAI has launched its Decisions API in public beta, offering developers up to ten times faster routing and classification powered by its GPT-6 Luna model.

OpenAI has released its Decisions API into public beta, providing developers with a structured endpoint designed for rapid classification, routing, and agent control-flow judgments. Powered by GPT-6 Luna, the lowest-priced model in the GPT-6 family, the new API is built to handle text and image inputs. In a company demonstration routing 10,000 customer requests among Billing, Technical, and Sales categories, the Decisions API returned results in 150 milliseconds per request. This represents a 10.7-fold speed increase over the 1.6 seconds required by the standard Responses API.
The API supports three structured output formats to eliminate the need for parsing raw text or fixing broken JSON. These include predicates for estimating probabilities, choices for selecting predefined options with confidence scores, and scores for numeric evaluations. For developers building autonomous agents, this speedup is critical. Running 10 sequential decision calls would take roughly 1.5 seconds of model time using the Decisions API, compared to 16 seconds through the Responses API, excluding network overhead.
Early evaluations from preview users show how the endpoint performs against specialized alternatives. In an offline replay of computer-use tasks conducted by Every, the Decisions API selected the correct control in 76 of 78 steps, while Jev, a dedicated decision model, got 73 correct. The two systems tied on accuracy in a thread-classification test, though Jev was faster and performed better across Every's broader evaluation set.
While GPT-6 Luna costs $0.10 per million input tokens and $0.50 per million output tokens through standard chat completions, OpenAI has not yet published the specific pricing, candidate-answer limits, or tuning support for the Decisions API. The company has also not detailed the exact serving optimizations or latency percentiles behind the benchmark. However, the endpoint integrates directly with OpenAI's existing software development kits, authentication, and billing systems, making it an easy addition for current customers.
This is our own summary of reporting by AlphaSignal



