Microsoft AI Releases Microsoft-Decision-1: A Qwen3.5-9B Decision-Scoring Model


Microsoft has released Microsoft-Decision-1, a decision model for routing, classification, verification and agent control. Microsoft-Decision-1 is a decision-scoring model that returns a calibrated probability for each fixed answer option instead of generated text. It is post-trained from Alibaba’s Qwen3.5-9B and available now in Microsoft Foundry and OpenRouter. .

TL;DR

  • Size: Built on Qwen3.5-9B; exact parameter count not disclosed. 32,768-token context window.
  • Runs on: Hosted API only (Microsoft Foundry, OpenRouter via Azure). No open weights, no quantized variants, hardware not disclosed.
  • Performance: Highest average accuracy in Microsoft’s 36-benchmark comparison, at 85 ms p50 latency.
  • Best: 83.5% average accuracy across 36 benchmarks, ahead of Quyet-1.0-Large at 81.9%.
  • Worst: Calibration of 92.2, second to Quyet-1.0-Large at 93.1. Text-only, no explanations.
  • Bottom line:
    • Best thing: fast, cheap, calibrated decisions at $0.042 per million input tokens with free output.
      • Worst thing: closed weights, and every benchmark is vendor-run.

What is a decision model?

A decision model reads an input and scores a closed set of options. It does not write prose. Microsoft frames decision models as a new AI category, built for outputs software can act on immediately.

How does Microsoft-Decision-1 work?

Microsoft post-trained Qwen3.5-9B for single-pass decision scoring. Given a situation, a question and fixed options, it returns a probability per option in one call. It plans to rebase future versions on MAI and OpenAI models.

The Foundry model card lists the supported formats:

  • yes/no, multiple-choice, rating, classification and rubric questions
  • grading of AI responses and proposed agent actions
  • groundedness checks against supplied evidence
  • explicit abstention options such as “cannot tell”

Training used public datasets under Microsoft’s Open Data process plus synthetic data. Output is JSON. OpenRouter notes that weights update continually while the API shape stays fixed.

How fast and accurate is it?

Microsoft compared 9 systems across 36 benchmarks with 147,137 questions. Benchmarks were kept blind from training. Microsoft-Decision-1 led on average accuracy at 83.5%.

Its p50 latency was 85 ms, with p95 at 125 ms. That is 4.5 times quicker than Quyet-1.0-Large and 35 times quicker than GPT-6 Sol, which took 3.01 s.



Source link

  • Related Posts

    Nace AI Open-Sources Drex 1.5: A 9B Decision Model That Scores Options, Not Text

    Nace.AI has open-sourced Drex 1.5, a 9B decision model for agents and backend workflows. The Drex 1.5 decision model does not write text. It reads a state and typed questions,…

    OpenAI Decisions API Hits Public Beta With 10x Faster Typed Answers

    OpenAI has released the Decisions API in public beta. It turns text and images into typed answers your code can branch on. OpenAI team states…

    Leave a Reply

    Your email address will not be published. Required fields are marked *