AINewsnow

Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak

Alibaba's Qwen team released Qwen-Audio-3.1-Realtime, a full-duplex voice model trained to reason, call tools and decide when to speak. On a τ-Voice adaptation, task success rises to 82.0% from 78.4%. Replies to background speech drop from 73% to 13%. It is available now as an API on QwenCloud. The…

Read the full story at MarkTechPost ↗

Timeline · 1 report

  1. 2026-09-29 04:58 · MarkTechPost
    Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak

More stories

  1. Best current Qwen Flash Next Q4-ish? + worth using? — r/LocalLLaMA
  2. I built an open-source Prompt Engine & Screenplay Studio to solve video diffusion token drops & character drift (MiniMax H3 / Kling / Maestro) — r/PromptEngineering
  3. Layer Extract & Layer Remove Loras For Qwen Image 2.1 — r/StableDiffusion
  4. Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task. — r/LocalLLaMA
  5. Community reports say the first samples of Qwen 4 are already approaching Fable / Opus-level quality. — r/singularity
  6. Qwen-Image 2.1 Inpainting with LanPaint — alpha channel included — r/StableDiffusion
  7. 85 GB DeepSeek-V4-Flash at ~3 tok/s on a 12 GB RTX 3060 + 64 GB DDR5 RAM - Overspill for FreeToken, inspired by Colibri — r/LocalLLaMA
  8. A LoRA I made: AnyAngle LoRA for Qwen Image 2.1. Style-Aligned Arbitrary Camera Angles — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →