Qwen3.8 flash next ISTA-DASLab GGUF 50t/s TG and 1500t/s PP with 12GB VRAM and 64GB RAM Laptop on 'Strata' engine
I think most people are sleeping on this inference engine. I tried multiple llama.cpp forks and none of them comes close to the inference speed of Strata. Initial version had some bugs with kv cache, cpu throttling and the developer fixed them. Inference engine (only runs on Nvidia for now; AMD sup…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-10-01 20:21 · r/LocalLLaMA
Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant - 2026-09-30 04:01 · r/LocalLLaMA
Qwen3.8 flash next ISTA-DASLab GGUF 50t/s TG and 1500t/s PP with 12GB VRAM and 64GB RAM Laptop on 'Strata' engine
More stories
- I built an open-source tool that tells you why your vLLM server is slow (NVIDIA only for now, Mac support planned) — r/LocalLLM
- Open source inference engine (like LM Studio or Unsloth Desktop) that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD or nothing but a CPU. — r/LocalLLaMA
- Top AI and tech firms sign 'morally binding' accord to 'self-police' development after meeting at White House — Euronews Next
- Optimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell — PyTorch Blog
- Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo — r/LocalLLM
- add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp — r/LocalLLaMA
- Sharing my Qwen3.8-27B at 8-bit on 2x RTX 3090 with vLLM: 115 tok/s decode, ~1,780 tok/s prefill, 262K context (NVLink + DFlash2, full recipe and A/B numbers) — r/LocalLLM
- Browser FPS with 3D models, textures and SFX generated locally on one GPU, plus a local Qwen 27B for part of the code: my pipeline and what failed — r/LocalLLM
Get the daily brief of stories like this at 6:30 every morning →