GSQHalo.cpp - Another day, another fork. This one is for the memory constraint folks. 2×256K + MTP, 1386 t/s prefill and 44 t/s decode at 128K using GSQ-RCO Quants on a 96GB machine, KV cache storage on SSD
Coverage of "GSQHalo.cpp - Another day, another fork. This one is for the memory constraint folks. 2×256K + MTP, 1386 t/s prefill and 44 t/s decode at 128K using GSQ-RCO Quants on a 96GB machine, KV cache storage on SSD" from 1 source, with a live timeline of who reported what and when.
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-04 00:09 · r/LocalLLM
GSQHalo.cpp - Another day, another fork. This one is for the memory constraint folks. 2×256K + MTP, 1386 t/s prefill and 44 t/s decode at 128K using GSQ-RCO Quants on a 96GB machine, KV cache storage on SSD
More stories
- NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI — NVIDIA Blog
- Guided Vision in Gemini Live: built for accessibility — Google Gemini Blog
- Google tests its plan for AI data centers in space with Project Suncatcher — Scientific American
- A model guide for the GPT-6 family — OpenAI News
- The latest AI news we announced in September 2026 — Google AI Blog
- Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs — Allen Institute for AI (Ai2)
- OpenAI Fires Researchers for Allegedly Sharing Information with AI Safety Group — Wall Street Journal Technology
- Google unveils Gemini 4 Argon: Its most powerful AI model yet, focused on coding and cyber defence — Mint AI
Get the daily brief of stories like this at 6:30 every morning →