
Qwen-Audio-3.1-Realtime: Tool Use Meets a Turn-Taking Trade-off
Published by AINave Editorial • Reviewed by Ramit
Alibaba’s Qwen team has released Qwen-Audio-3.1-Realtime, a full-duplex voice model designed to let an agent listen, respond and use tools without treating turn-taking as a simple pause-and-reply exchange. The interesting detail is the trade-off: the release reports fewer responses to background speech, but slower stops when a user interrupts. The reported measures make interruption handling a more useful lens than a broad claim that the model is simply more conversational.
Voice interaction is split into decisions and speech
Qwen-Audio-3.1 is a five-model audio stack covering speech recognition, speech synthesis and real-time interaction. For Realtime, the described system uses one model to decide whether to keep listening, speak, stop or resume, and another to produce response text. A voice renderer then turns that text into streaming speech, informed by conversation history and acoustic context. The release describes this division of work.
That separation matters for voice-agent builders because a useful spoken answer depends on more than the answer’s wording. The system must also decide whether the person is addressing it, whether to yield during an interruption and whether a tool call is warranted. Qwen says it trained tool use in executable environments where a task could require a valid database change, a justified refusal or recognition that a request was unsupported. In search training, mean queries per call fell from 4.37 to 1.05, while trigger F1 also fell, from 60.87% to 58.61%. That is a concrete reminder that using fewer searches and choosing when to search are distinct goals. Those figures come from the reported training results.
Fewer background replies, slower interruption stops
On Full-Duplex-Bench v1.5, the reported rate of replies to people talking to someone else fell from 0.13 to 0.03. But after interruptions, the unwanted-resume rate rose from 0.035 to 0.130. The reported interruption-stop latency was 1.116 seconds for Qwen-Audio-3.1-Realtime, compared with 0.383 seconds for GPT-Realtime-2. These are benchmark results, not guarantees for a particular deployed agent, but they expose a practical tension: avoiding irrelevant speech does not automatically make an agent faster to yield when interrupted. The benchmark comparisons and conditions are reported in the release coverage.
API access, with token prices to interpret carefully
The model is available as qwen-audio-3.1-realtime-plus through QwenCloud over WebSocket. Its listing includes function calling and web search, and gives a 262K-token context, with a 245K maximum input and 16K maximum output. No open weights were announced, so the described route is managed API access rather than self-hosting. The deployment details and limits are listed here.
Listed rates are $6.4 per million audio input tokens and $0.8 per million text input tokens. Text and audio output is listed at $24 per million tokens, with output text not charged. Those figures are not a direct cost comparison with another provider: audio tokenization differs, and token rates alone do not establish the total cost of a voice workflow. For teams assessing the model, the trade-off worth understanding is not just price or answer quality, but how its listening and interruption decisions fit the interaction they need. The listed pricing and tokenization caveat are part of that calculation.



















