AINewsnow

MiMo-V2.6-Flash on vLLM: fixes for "empty responses" with thinking + tools, and a hidden 2,048-token output cap

Some people here say MiMo-V2.6 is bad with tools and are going back to GLM-5.3-Flash. I spent today running MiMo-V2.6-Flash-RL as the backend for an agent harness, on 2× DGX Spark with vLLM, using the tonyd2wild recipe. Most of the "tool problems" I hit turned out to be serving bugs rather than the…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-22 19:28 · r/LocalLLaMA
    MiMo-V2.6-Flash on vLLM: fixes for "empty responses" with thinking + tools, and a hidden 2,048-token output cap

More stories

  1. Jev's calibration was measured. The LLMs won [D] — r/MachineLearning
  2. 299 real user intents tested Jev against production base line. Here is the result. — r/AI_Agents
  3. GLM 5.3 now available in Mistral Vibe Code for Pro, Team and Enterprise — r/ChatGPTCoding
  4. CAISI’s Assessment of Z.ai’s GLM-5.3 Cyber Capabilities — r/ArtificialInteligence
  5. CPU Inference on Dell Poweredge R820 — r/LocalLLM
  6. Z.ai disables coding assistant feature after flaw exposed enterprise code upload risk — InfoWorld AI
  7. Chinese AI firm Z.ai faces reputation hit after users spot unauthorised uploads — South China Morning Post Tech
  8. Zhipu open-sources ZCode after data dispute and plans no-retention controls for MaaS — TechNode

Get the daily brief of stories like this at 6:30 every morning →