OpenAI has opened its Decisions API to public beta and added image inputs that allow applications to make fast decisions from visual content. The API, first announced last week as a likely response to TypeSafe's Jev, is priced at $0.10 per 1 million input tokens, with no output-token charges and no charges for cache reads or writes.
The API runs on GPT-6 Luna and turns text or images into three types of structured outputs: predicates, which gauge how likely something is to be true; choices, which pick the best fit from predefined options; and scores, which place an input on a numeric range.
The image input capabilities are aimed at fast, vision-driven action. OpenAI demonstrates the API navigating a video-game car through traffic, deciding whether to switch lanes or keep driving. The company claims it can make such decisions in a fraction of the time a reasoning model would take.
OpenAI also hints at broader computer-use applications, where the API could take screenshots and quickly select the next action. Its demonstration video features a preview unit of Hugging Face's upcoming Microduck robot, an open-source bipedal developer robot. Paired with GPT-Live, OpenAI's real-time voice model, the Decisions API can analyze the robot's camera frames to decide where it should look. In the example, a command to "follow the apple" prompts the API to identify the fruit and direct the robot toward it, and the pairing can respond to more ambiguous instructions such as "follow the fruit."
OpenAI's model is not the only one in the field. Perplexity, Cloudflare and Amazon all released their own decision models this month, opting for open-weight models that developers can download and run on their own infrastructure.