Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak


Alibaba’s Qwen team has released Qwen-Audio-3.1, a 5-model audio stack spanning ASR, TTS and realtime interaction. The main model is Qwen-Audio-3.1-Realtime, a full-duplex speech model built for voice agents that call tools. Qwen also cut prices: about 85% on Realtime, about 70% on TTS and up to 95% on ASR.

Is it deployable? Yes, as a managed API. qwen-audio-3.1-realtime-plus is live on QwenCloud over WebSocket. No open weights were announced.

What Ships on QwenCloud

The model page lists text and audio as both input and output. Context is 262K tokens, with 245K max input and 16K max output. Default limits are 60 requests and 100K tokens per minute. Pricing is $6.4 per 1M audio input tokens and $0.8 per 1M text input tokens. Text and audio output costs $24 per 1M tokens, with output text not charged. Key features include function calling, web search, structured outputs, context cache and fine-tuning.

A companion model, Qwen-Audio-3.1-ASR-Flash-Filetrans, targets offline long-audio transcription. It supports hot words, speaker separation, punctuation and multilingual plus Chinese dialect recognition. It costs $0.15 input and $0.47 output per 1M tokens.

Architecture: 2 Models Behind 1 Voice

The system runs 2 models with the same Audio Encoder and LLM design. A full-duplex decision model predicts whether to keep listening, speak, stop or resume. A speech-to-text model writes the response content as text. A context-aware voice renderer then turns that text into streaming speech. It conditions on conversation history, voice cues and acoustic context.

Training is organized into 3 layers: Think, Act, and Speak and Coordinate.

Think: M²-OPD

Core-Cocktail SFT re-anchors the audio model to its source text LLM using million-hour-scale paired data. Multimodality OPD follows. A Text Teacher and a frozen Audio Reference score each token of the student’s own trajectory. This is on-policy distillation, not imitation of pre-written answers. Domain experts for empathy, pragmatic intent and acoustic scenes are then trained with GRPO. Multi-Teacher OPD merges them into 1 deployable model.

Act: Executable Environments

Each training domain bundles a tool pool, a stateful JSON database and a natural-language business policy. Domains are seeded from open-source tool and MCP server definitions. Every task defines 1 of 3 outcomes: a write, a justified refusal, or an unsupported request. Scoring checks terminal state, then permitted writes, then behavioral assertions. A fluent reply cannot rescue a failed state check.

GRPO receives rewards at dialogue, milestone and turn level. Search training penalizes redundant queries with rquery=qmin⁡(1,nrefnpred)r_{\text{query}} = q \min \left( 1, \frac{n_{\text{ref}}}{n_{\text{pred}}} \right). Mean queries per search call fell from 4.37 to 1.05. Trigger F1 slipped from 60.87% to 58.61%.

Speak and Coordinate

This layer decides whether, when and how to speak. On Full-Duplex-Bench v1.5, replies to people talking to someone else fell from 0.13 to 0.03. On v3.0, the filler rate dropped from 0.7590 to 0.2960. There are trade-offs. After interruptions, the unwanted resume rate rose from 0.035 to 0.130. Interruption stop latency is 1.116 seconds, versus 0.383 for GPT-Realtime-2.

Interactive Explainer

Explore the Think, Act, Speak loop, duplex decisions, a scored training episode and the search reward.

‘+bar(‘3.0′,x[1],’f30’,x[1].toFixed(3).replace(/0+$/,”).replace(/\.$/,”))+bar(‘3.1′,x[2],’f31’,x[2].toFixed(4).replace(/0+$/,”).replace(/\.$/,”))}
$(‘qaBars’).innerHTML=h;anim($(‘qaBars’));post()}
function bar(lab,v,cls,txt){return ‘

‘+lab+’‘+(txt===”?’0′:txt)+’

‘}
function anim(root){var f=root.querySelectorAll(‘.fill’);setTimeout(function(){for(var i=0;iSK2859 and Aug 20 fast, almost in one breath.’},
{t:’Model calls search_flights(SFO, NRT, 2026-08-20) to confirm Flex seats remain.’},
{t:’Model reads back flight, date, Flex fare, $990.0 and expiry, then gets explicit confirmation.’,c:1},
{t:’Model calls place_hold(..., fare_class="flex"). Tool writes to the private DB copy.’,w:1},
{t:’Model tells Emily the reference HD482902 and expiry 2026-08-15.’},
{t:’Scorer runs 3 checks in order, using the tool trace, not the dialogue.’,s:1}];
var ei=-1;
var DB0='”HD482901″: {\n “flight_no”: “SK2859”,\n “status”: “active”,\n “expiry_date”: “2026-08-08″ // expired\n}’;
var DB1=DB0+’,\n”HD482902”: {\n “flight_no”: “SK2859”,\n “date”: “2026-08-20”,\n “fare_class”: “flex”,\n “amount”: 990.0,\n “status”: “active”,\n “created_date”: “2026-08-12”,\n “expiry_date”: “2026-08-15″\n}’;
function epRender(){var ul=$(‘qaSteps’),skip=$(‘qaSkip’).checked,h=””;for(var i=0;i‘+(i+1)+’‘+(EP[i].c&&skip&&i<=ei?’Skipped: no read-back, no confirmation.’:EP[i].t)+’‘}ul.innerHTML=h;
$(‘qaDb’).textContent=ei>=3?DB1:DB0;
var s=ei>=5;function set(id,ok){var e=$(id);e.className=s?(ok?’ok’:’bad’):’pend’;e.textContent=s?(ok?’PASS’:’FAIL’):’pending’}
set(‘qaC1’,true);set(‘qaC2’,true);set(‘qaC3’,!skip);var v=$(‘qaVerdict’);
if(s){v.innerHTML=skip?’Episode FAILS. Correct state cannot excuse a broken interaction rule.’:’Episode PASSES. Valid success feeds GRPO.’}else v.textContent=”Verdict appears after scoring”;post()}
$(‘qaEpNext’).onclick=function(){if(ei=r?’ over’:”)+'”>query ‘+(i+1)+”;$(‘qaChips’).innerHTML=h;post()}
$(‘qaQ’).oninput=rw;$(‘qaR’).oninput=rw;$(‘qaP’).oninput=rw;rw();

/* SLIDE 5 */
var M=[
{k:’τ-Voice overall (%)’,g:43.8,a:78.4,b:82.0,d:1,mx:100,low:false,n:’Half-duplex speech-to-text adaptation by the Qwen team; not comparable to official full-duplex τ-Voice results.’},
{k:’Audio MultiChallenge (%)’,g:50.33,a:47.12,b:52.21,d:2,mx:100,low:false,n:’GPT-4o-mini judge; official AMC uses o4-mini, so scores are not directly comparable to the leaderboard.’},
{k:’Multilingual BBA avg (%)’,g:82.1,a:81.7,b:88.1,d:1,mx:100,low:false,n:’In-house 14-language extension of Big Bench Audio. Largest gains on Arabic, Thai and Vietnamese.’},
{k:’Multi-turn attack, zh (%)’,g:42.0,a:80.5,b:26.0,d:1,mx:100,low:true,n:’In-house, 200 sessions. Attack success rate: lower is better.’},
{k:’Multi-party sessions passed’,g:0.07,a:0.08,b:0.96,d:2,mx:1,low:false,n:’In-house Qwen-Audio-FDB, 100 sessions; a session passes only if every point passes.’},
{k:’Human red-team pass (%)’,g:96.0,a:null,b:92.0,d:1,mx:100,low:false,n:’5 testers, 50 sessions, automatic judge. GPT-Realtime-2 leads here.’}];
var curMet=0,mb=$(‘qaMet’);
for(var m=0;m

Qwen-Audio-3.0-Realtimenot reported



Source link

  • Related Posts

    Anthropic Releases Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0 at the Same $2/$10 Price

    Anthropic just released Claude Sonnet 5.5. It is the second model in the Claude 5.5 family, following Claude Opus 5.5. Anthropic positions it as a faster, lower-cost complement to Opus…

    NVIDIA Launches Open Agent Safety Platform: OpenShell Sandboxes Agents on Vera CPUs While Sentry on BlueField-4 Quarantines Them in Milliseconds

    NVIDIA has launched the NVIDIA Open Agent Safety Platform, an open software platform and reference system design for AI agent security. It pairs the OpenShell secure runtime with NVIDIA Sentry,…

    Leave a Reply

    Your email address will not be published. Required fields are marked *