Microsoft launched the Microsoft-Decision-1 model within the Foundry platform, just three days after OpenAI made its Decisions API available to developers in public preview. Contrary to what the closeness between the two companies might suggest, Microsoft’s model was initially trained on Alibaba’s Qwen3.5-9B, with the company announcing its intention to rebuild it later on its own MAI models and on OpenAI models.
The model is intended to add decision-making capabilities to existing applications, agents, and workflows, such as evaluating outputs, selecting the appropriate model, or determining the next step in an automated process. Microsoft offers it within Foundry at $0.042 per million input tokens, while output tokens are free; this is the same price TypeSafe charges for the Jev model.
Microsoft’s Internal Tests
Microsoft says that the model’s first user is the company itself. Xbox Research tested it to rank more than 10,000 pieces of player feedback, the Copilot team used it to evaluate conversations and agent responses, and on-call engineers used it to extract context during live incidents. Microsoft Discovery also used it to evaluate experiments before an agent replanned its task.
According to the company’s figures, Decision-1 was more than 14 times faster than GPT-6 Sol in an Xbox-related test and achieved 46 times higher consistency in a Discovery test. The material does not explain the methodology behind these comparisons, so the figures cannot be considered a general judgment of the model’s performance outside those scenarios.
Why Does This Matter?
Decision models are moving toward serving as a low-cost control layer between applications and larger generative models. This layer becomes important when an agent needs to select a model, tool, or execution path many times, because the cost of each decision can accumulate alongside the flow of generative requests associated with it.
Microsoft has an additional incentive to develop this type of model internally, particularly with GitHub Copilot’s plans to decide whether to execute certain tasks on the device or send them to cloud models. However, the company has not confirmed whether Decision-1 will actually participate in Copilot decisions, nor has it disclosed what is sent to the cloud.
Compatibility and Media Limitations
The Foundry listing states that Decision-1 accepts up to 32,768 text tokens and returns results in JSON format, but does not support images. This places it in a narrower range than OpenAI’s Decisions API and Cloudflare’s Clef, which uses a visual encoder.
Microsoft has also not announced the availability of the model’s weights, while Cloudflare released the Clef model under the Apache 2.0 license. AWS, Upstage, and Ollama adopted TypeSafe’s System One interface as a shared interface in this field, but Microsoft has not confirmed that Decision-1 is fully compatible with it, even though its code example in Foundry calls an endpoint named /systemone.
Confidence Does Not Mean Resilience to Misinformation
Microsoft says that the model’s probabilities are calibrated; that is, a prediction carrying a probability of 90% should be correct in approximately nine out of every ten cases within representative cases. However, it recommends that customers verify calibration using their own data.
The JevOut study points to why this caveat matters: short, naturally worded additions reversed 312 of 508 initially correct decisions made by the Jev model, with the model assigning a probability of at least 70% to the wrong answer in 229 cases. Microsoft’s test of its model, meanwhile, included eight types of changes, such as reordering the options and rephrasing the description, and changed 1.3% of the answers on average. However, the JevOut study did not test Decision-1, so this result does not establish how the model would handle inputs designed to mislead it.