AINewsnow

The KV Cache Is the Bottleneck Now β€” 1-Bit Quantization, MoE Stragglers, and Per-Second GPU Billing

This story is from 2026-10-07. It is preserved in the archive; the latest stories are on the live feed.

Welcome to this week's LLM Inference Digest, covering roughly September 30 – October 7, 2026 across arXiv (cs.LG, cs.DC), the vLLM blog, Hugging Face blog, Together AI blog, Modal blog, and Latent Space. πŸ”₯ Highlights MegaFlux: Skew-Resilient MoE Megakernels via Pipelined Expert Replication β€” fixes…

Read the full story at DEV Community β€” AI β†—

Timeline Β· 1 report

  1. 2026-10-07 12:05 Β· DEV Community β€” AI
    The KV Cache Is the Bottleneck Now β€” 1-Bit Quantization, MoE Stragglers, and Per-Second GPU Billing

More stories

  1. Introducing EmbeddingGemma 2: A best-in-class open model for natively multimodal embeddings | Google β€” r/LocalLLaMA
  2. Qwen3.8-Flash-Next-Q8_0 running on a V100 @ 130Watts 32GB Vram and 128GB System Ram β€” r/LocalLLM
  3. Rogue AI or human error? The real story behind the OpenAI-Hugging Face incident β€” Scientific American
  4. World Models: The Simulation Strikes Back β€” r/computervision
  5. ~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5. β€” r/huggingface
  6. A quick Minimax H3 news round-up - 4th October 2026 β€” r/comfyui
  7. Trained a ~20K LM (probably smallest) that can still write stories β€” r/huggingface
  8. FastVideo’s FastH3 now runs on a single consumer machine β€” r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning β†’