Smartphone and microphone representing Qwen-Audio-3.1-Realtime voice agents with real-time audio waves

Alibaba Qwen Releases Qwen-Audio-3.1-Realtime for More Capable Voice Agents

The new duplex voice model combines spoken reasoning, tool use, and turn-taking control in one real-time system.

TL;DR: Qwen-Audio-3.1-Realtime is Alibaba's new real-time duplex voice model, combining spoken reasoning, tool use, and turn-taking control with support for function calling, web search, and voice cloning.

Alibaba's Qwen team has introduced Qwen-Audio-3.1-Realtime, a real-time voice model designed to do more than answer spoken questions. The model combines audio and text input and output with tool use, web search, and voice cloning, while its system design focuses on reasoning over changing requests and deciding when to speak or act.

Alibaba Cloud Model Studio lists qwen-audio-3.1-realtime-plus as a real-time duplex speech model available in its platform. The service supports streaming conversations through WebSocket, WebRTC, and AOQ, giving developers several ways to build voice applications. The official Qwen-Audio-3.1-Realtime project page publishes the model's benchmark results, while Alibaba Cloud's Model Studio documentation documents the hosted API, limits, and pricing.

Qwen-Audio-3.1-Realtime Treats Voice as an Ongoing Interaction

The main change is the move from a simple speech-in, speech-out exchange toward a voice agent that can manage an interaction over time. Qwen's technical report describes three connected abilities: Think, Act, and Speak & Coordinate.

  • Think: reason over evolving spoken requests and use audio-native capabilities alongside language skills.
  • Act: use tools, interpret feedback, and complete tasks through executable environments.
  • Speak & Coordinate: control how, when, and whether the assistant speaks or acts during an interaction.

That design puts turn-taking and action selection alongside the model's answer generation. In a voice interface, that matters because an assistant has to handle interruptions, incomplete requests, and changing instructions without treating every sound as a new command.

The Model Uses Full-Duplex Voice Interaction

Qwen-Audio-3.1-Realtime supports full-duplex streaming, meaning audio can be sent and received continuously rather than waiting for a complete spoken turn before the system responds. Alibaba's real-time guide supports acoustic VAD, a semantic turn-detection mode, and push-to-talk.

The semantic mode is designed to use both acoustic and semantic information when deciding whether a person has finished speaking. Alibaba says this helps prevent filler sounds such as "uh" or "hmm" from unnecessarily interrupting a conversation.

That approach is different from a basic speech pipeline where an automatic speech recognition system transcribes a finished utterance, a language model generates text, and a text-to-speech system speaks the result. Qwen-Audio handles the real-time speech interaction as a speech-to-speech model.

Early Results Show Gains Over Qwen-Audio-3.0-Realtime

Qwen's published evaluation reports several improvements over its previous real-time model. The headline results cover multi-turn instructions, spoken task execution, multilingual reasoning, and empathetic responses.

EvaluationQwen-Audio-3.1-RealtimeChange vs. 3.0
Audio MultiChallenge52.21%+5.09 percentage points
τ²-Bench Audio82.19%+3.59 percentage points
14-language QA88.09%+6.40 percentage points
EchoMind4.03/5+0.33 score points

The technical report also gives a separate comparison on its half-duplex speech-to-text adaptation of τ-Voice, where overall task success rises from 78.4% with Qwen-Audio-3.0-Realtime to 82.0% with 3.1. On Full-Duplex-Bench v1.5, the report says the response rate to background speech drops from 73.0% to 13.0%.

Benchmark context matters

These results come from Qwen's published evaluations and are not a universal ranking of voice models. The persistent voice-agent runtime shown on the Qwen project page is described as a system extension, and the page says matched Qwen-Audio-3.1 measurements for that extension are not yet reported.

Tool Calling, Web Search, and Voice Cloning Are Built In

The hosted qwen-audio-3.1-realtime-plus service accepts both audio and text and can produce audio and text. Alibaba documents support for function calling, web search, and cloned voices.

Function calling lets a voice application connect the model to external actions. Web search provides a way to retrieve current information during an interaction. Voice cloning lets developers create a voice through Alibaba's voice-cloning API and then use the resulting voice ID with the real-time model.

Alibaba notes one important limitation: web search and function calling cannot be enabled together in the same configuration.

The API Is Designed for Existing Real-Time Voice Apps

Developers can connect to Qwen-Audio-3.1-Realtime through WebSocket, WebRTC, or Alibaba's AOQ transport. Alibaba also provides Android, iOS, and HarmonyOS SDK references for its real-time connection options.

For teams already using Qwen-Audio-3.0-Realtime-Plus, Alibaba says the WebSocket event protocol remains unchanged when migrating to 3.1 Plus. The main model change is the qwen-audio-3.1-realtime-plus model identifier, along with the appropriate voice configuration.

Model detailQwen-Audio-3.1-Realtime-Plus
Context window262,144 tokens
Maximum input245,760 tokens
Maximum output16,384 tokens
Singapore text input$0.80 per million tokens
Singapore audio input$6.40 per million tokens
Singapore text output$6.40 per million tokens
Singapore audio output$24 per million tokens

Alibaba lists separate Beijing pricing and rate limits, so the cost figures above apply specifically to the Singapore endpoint. Pricing for audio and text is also separated, which matters when estimating the cost of a voice-heavy application.

Qwen Is Expanding Its Voice AI Push

Qwen-Audio-3.1-Realtime arrives alongside a broader push into audio and real-time AI. Saganote previously covered Alibaba's Qwen3.8 launch and Alibaba's Qwen3.8-Max model, while JetBrains Junie Local running Qwen3.6 on M5 Macs shows another path for Qwen models outside hosted voice services.

The voice side of the market is also becoming more crowded. Saganote has covered ElevenLabs v4's expression controls and language support, xAI's Grok Voice Transcribe 2.0, and Grok Voice Think Fast 2 for real-time agents.

Google's Gemini 3.8 Live background reasoning and OpenAI's GPT-Live full-duplex voice model provide additional context for the direction of real-time voice systems. Microsoft's Foundry Build 2026 coverage also shows how hosted agents and broader agent infrastructure are becoming part of the same developer conversation. Qwen's approach is to put spoken reasoning, action, and turn-taking into the same interaction loop.

Alibaba's Qwen work also reaches beyond voice. Saganote previously reported that Alibaba's Qwen model will power Apple Intelligence in China, linking the model family to a separate consumer AI deployment.

What Qwen-Audio-3.1-Realtime Means for Developers

The practical change is that a voice application can be built around a model that accepts continuous speech, returns streaming speech and text, and connects to external actions. Developers can choose between automatic acoustic turn detection, semantic turn detection, or push-to-talk depending on the interaction they need.

The benchmark gains suggest progress over Qwen-Audio-3.0-Realtime, especially on multi-turn instruction following, spoken task execution, multilingual reasoning, and handling background speech. They do not by themselves establish how Qwen-Audio-3.1-Realtime performs across every production workload or against every competing real-time model.

For now, Alibaba Cloud Model Studio provides the clearest developer path to the model, with documented real-time APIs, multiple transport options, tool integrations, and regional pricing. The Qwen project also presents longer-running voice-agent behavior as a separate runtime extension rather than as a matched 3.1 model benchmark result.


Frequently Asked Questions

What is Qwen-Audio-3.1-Realtime?
It is Alibaba's real-time duplex voice model for streaming audio and text interaction, with support for function calling, web search, and voice cloning.
Is Qwen-Audio-3.1-Realtime full duplex?
Yes. Alibaba describes it as a full-duplex real-time speech model with streaming audio input and output.
Can Qwen-Audio-3.1-Realtime use tools?
Yes. The hosted model supports function calling and web search, although Alibaba says those two features cannot be enabled together.
How does it differ from Qwen-Audio-3.0-Realtime?
Qwen reports higher results on several evaluations, including Audio MultiChallenge, τ²-Bench Audio, 14-language QA, and EchoMind, while its technical report also reports a lower response rate to background speech.

Qwen-Audio-3.1-Realtime is now positioned as a hosted real-time voice model that can listen, reason, act, and manage turn-taking inside one streaming interaction. Its published results show measurable gains over Qwen-Audio-3.0-Realtime, while Alibaba's Model Studio documentation provides the APIs and deployment details needed to test the model in real applications.


Share this
Saganote

About Author

Saganote

Saganote is an independent technology publication covering artificial intelligence, cybersecurity, startups, software, consumer technology, and innovation. Our editorial team researches, writes, and reviews original news, analysis, and explainers to provide accurate, timely, and well-sourced coverage of the technology industry.