AINewsnow

Splash 1.1.0 released, GGUF quants support, MLX import and more

On my M5 Pro 64GB I can comfortably work in an agentic setup with the Qwen3.8 27B model in good quality (Unsloth UD-Q4_K_XL) at a decent speed of 50 t/s. Splash combines optimized kernels, excellent speculative decoding, a well-implemented prefix cache, and mixed-weight support in a single program.…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-26 17:27 · r/LocalLLaMA
    Splash 1.1.0 released, GGUF quants support, MLX import and more

More stories

  1. Introducing Gemini 3.8 Live with Live Avatar — Google Gemini Blog
  2. Accelerating vision-language models with LFM2.5-VL-DSpark — Hugging Face Blog
  3. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  4. Qwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080 — r/LocalLLaMA
  5. OpenAI’s A.I. Went Rogue and Meddled With U.S. Government Websites — New York Times AI
  6. Am I the only one who actually likes GPT-6 Sol and Luna? — r/ChatGPT
  7. OpenAI agent ‘hacked’ Australian Govt Medicare portal, PM Albanese calls it ‘unacceptable’: What happened? — Mint AI
  8. Saaras V4: Sarvam AI’s new speech recognition model adds 22 Indian languages, 5 speech formats — Mint AI

Get the daily brief of stories like this at 6:30 every morning →