AINewsnow

Anyone sitting on a lot of slow system memory and a modest GPU.. try Strata + Qwen3.8 Next.

IQ3_XXS weights are just under 80GB and my slowww DDR4+7900XTX is stabilizing around 45-50t/s (sometimes higher while coding depending on mtp). Looking online I'm seeing similar results for users with 12GB and 16GB cards, and significantly faster numbers for owners of DDR5. (In comparison, Llama CP…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-03 03:12 · r/LocalLLaMA
    Anyone sitting on a lot of slow system memory and a modest GPU.. try Strata + Qwen3.8 Next.

More stories

  1. Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo — r/LocalLLM
  2. add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp — r/LocalLLaMA
  3. Browser FPS with 3D models, textures and SFX generated locally on one GPU, plus a local Qwen 27B for part of the code: my pipeline and what failed — r/LocalLLM
  4. Open source inference engine (like LM Studio or Unsloth Desktop) that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD or nothing but a CPU. — r/LocalLLaMA
  5. Sharing my Qwen3.8-27B at 8-bit on 2x RTX 3090 with vLLM: 115 tok/s decode, ~1,780 tok/s prefill, 262K context (NVLink + DFlash2, full recipe and A/B numbers) — r/LocalLLM
  6. WHIRL: an open-source native Windows inference engine for the Radeon AI PRO R9700 (C++/HIP, no WSL). Qwen3.8-27B fine-tune in MXFP4: up to 2.5× llama.cpp prefill, 107–328 tok/s decode, 3× server throughput — r/LocalLLM
  7. A beginner's guide to the Llama-4-Maverick-Instruct model by Meta on Replicate — DEV Community — Machine Learning
  8. New in llama.cpp: Decision Models — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →