Alibaba's Qwen team has introduced Qwen-Audio-3.1-Realtime, a real-time voice model designed to do more than answer spoken questions. The model combines audio and text input and output with tool use, web search, and voice cloning, while its system design focuses on reasoning over changing requests and deciding when to speak or act.
Alibaba Cloud Model Studio lists qwen-audio-3.1-realtime-plus as a real-time duplex speech model available in its platform. The service supports streaming conversations through WebSocket, WebRTC, and AOQ, giving developers several ways to build voice applications. The official Qwen-Audio-3.1-Realtime project page publishes the model's benchmark results, while Alibaba Cloud's Model Studio documentation documents the hosted API, limits, and pricing.
Qwen-Audio-3.1-Realtime Treats Voice as an Ongoing Interaction
The main change is the move from a simple speech-in, speech-out exchange toward a voice agent that can manage an interaction over time. Qwen's technical report describes three connected abilities: Think, Act, and Speak & Coordinate.
- Think: reason over evolving spoken requests and use audio-native capabilities alongside language skills.
- Act: use tools, interpret feedback, and complete tasks through executable environments.
- Speak & Coordinate: control how, when, and whether the assistant speaks or acts during an interaction.
That design puts turn-taking and action selection alongside the model's answer generation. In a voice interface, that matters because an assistant has to handle interruptions, incomplete requests, and changing instructions without treating every sound as a new command.
The Model Uses Full-Duplex Voice Interaction
Qwen-Audio-3.1-Realtime supports full-duplex streaming, meaning audio can be sent and received continuously rather than waiting for a complete spoken turn before the system responds. Alibaba's real-time guide supports acoustic VAD, a semantic turn-detection mode, and push-to-talk.
The semantic mode is designed to use both acoustic and semantic information when deciding whether a person has finished speaking. Alibaba says this helps prevent filler sounds such as "uh" or "hmm" from unnecessarily interrupting a conversation.
That approach is different from a basic speech pipeline where an automatic speech recognition system transcribes a finished utterance, a language model generates text, and a text-to-speech system speaks the result. Qwen-Audio handles the real-time speech interaction as a speech-to-speech model.
Early Results Show Gains Over Qwen-Audio-3.0-Realtime
Qwen's published evaluation reports several improvements over its previous real-time model. The headline results cover multi-turn instructions, spoken task execution, multilingual reasoning, and empathetic responses.
| Evaluation | Qwen-Audio-3.1-Realtime | Change vs. 3.0 |
|---|---|---|
| Audio MultiChallenge | 52.21% | +5.09 percentage points |
| τ²-Bench Audio | 82.19% | +3.59 percentage points |
| 14-language QA | 88.09% | +6.40 percentage points |
| EchoMind | 4.03/5 | +0.33 score points |
The technical report also gives a separate comparison on its half-duplex speech-to-text adaptation of τ-Voice, where overall task success rises from 78.4% with Qwen-Audio-3.0-Realtime to 82.0% with 3.1. On Full-Duplex-Bench v1.5, the report says the response rate to background speech drops from 73.0% to 13.0%.
These results come from Qwen's published evaluations and are not a universal ranking of voice models. The persistent voice-agent runtime shown on the Qwen project page is described as a system extension, and the page says matched Qwen-Audio-3.1 measurements for that extension are not yet reported.
Tool Calling, Web Search, and Voice Cloning Are Built In
The hosted qwen-audio-3.1-realtime-plus service accepts both audio and text and can produce audio and text. Alibaba documents support for function calling, web search, and cloned voices.
Function calling lets a voice application connect the model to external actions. Web search provides a way to retrieve current information during an interaction. Voice cloning lets developers create a voice through Alibaba's voice-cloning API and then use the resulting voice ID with the real-time model.
Alibaba notes one important limitation: web search and function calling cannot be enabled together in the same configuration.
The API Is Designed for Existing Real-Time Voice Apps
Developers can connect to Qwen-Audio-3.1-Realtime through WebSocket, WebRTC, or Alibaba's AOQ transport. Alibaba also provides Android, iOS, and HarmonyOS SDK references for its real-time connection options.
For teams already using Qwen-Audio-3.0-Realtime-Plus, Alibaba says the WebSocket event protocol remains unchanged when migrating to 3.1 Plus. The main model change is the qwen-audio-3.1-realtime-plus model identifier, along with the appropriate voice configuration.
| Model detail | Qwen-Audio-3.1-Realtime-Plus |
|---|---|
| Context window | 262,144 tokens |
| Maximum input | 245,760 tokens |
| Maximum output | 16,384 tokens |
| Singapore text input | $0.80 per million tokens |
| Singapore audio input | $6.40 per million tokens |
| Singapore text output | $6.40 per million tokens |
| Singapore audio output | $24 per million tokens |
Alibaba lists separate Beijing pricing and rate limits, so the cost figures above apply specifically to the Singapore endpoint. Pricing for audio and text is also separated, which matters when estimating the cost of a voice-heavy application.
Qwen Is Expanding Its Voice AI Push
Qwen-Audio-3.1-Realtime arrives alongside a broader push into audio and real-time AI. Saganote previously covered Alibaba's Qwen3.8 launch and Alibaba's Qwen3.8-Max model, while JetBrains Junie Local running Qwen3.6 on M5 Macs shows another path for Qwen models outside hosted voice services.
The voice side of the market is also becoming more crowded. Saganote has covered ElevenLabs v4's expression controls and language support, xAI's Grok Voice Transcribe 2.0, and Grok Voice Think Fast 2 for real-time agents.
Google's Gemini 3.8 Live background reasoning and OpenAI's GPT-Live full-duplex voice model provide additional context for the direction of real-time voice systems. Microsoft's Foundry Build 2026 coverage also shows how hosted agents and broader agent infrastructure are becoming part of the same developer conversation. Qwen's approach is to put spoken reasoning, action, and turn-taking into the same interaction loop.
Alibaba's Qwen work also reaches beyond voice. Saganote previously reported that Alibaba's Qwen model will power Apple Intelligence in China, linking the model family to a separate consumer AI deployment.
What Qwen-Audio-3.1-Realtime Means for Developers
The practical change is that a voice application can be built around a model that accepts continuous speech, returns streaming speech and text, and connects to external actions. Developers can choose between automatic acoustic turn detection, semantic turn detection, or push-to-talk depending on the interaction they need.
The benchmark gains suggest progress over Qwen-Audio-3.0-Realtime, especially on multi-turn instruction following, spoken task execution, multilingual reasoning, and handling background speech. They do not by themselves establish how Qwen-Audio-3.1-Realtime performs across every production workload or against every competing real-time model.
For now, Alibaba Cloud Model Studio provides the clearest developer path to the model, with documented real-time APIs, multiple transport options, tool integrations, and regional pricing. The Qwen project also presents longer-running voice-agent behavior as a separate runtime extension rather than as a matched 3.1 model benchmark result.
Frequently Asked Questions
What is Qwen-Audio-3.1-Realtime?
Is Qwen-Audio-3.1-Realtime full duplex?
Can Qwen-Audio-3.1-Realtime use tools?
How does it differ from Qwen-Audio-3.0-Realtime?
Qwen-Audio-3.1-Realtime is now positioned as a hosted real-time voice model that can listen, reason, act, and manage turn-taking inside one streaming interaction. Its published results show measurable gains over Qwen-Audio-3.0-Realtime, while Alibaba's Model Studio documentation provides the APIs and deployment details needed to test the model in real applications.
