AINewsnow

Strata Qwen3.8 FN abliterated dual 16GB GPU + 128GB RAM

Have been running Qwen3.8 FN Q4 abliterated with stock llama.cpp on my machine. I get around ~28tps and ~400-600pp. I was actually quite happy with it until I saw people with Strata running it at 60tps running on 12GB VRAM laptops. I prefer using an abliterated becuase I dont want to fight it, when…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-30 10:37 · r/LocalLLM
    Strata Qwen3.8 FN abliterated dual 16GB GPU + 128GB RAM

More stories

  1. Help me plan a Qwen 3.8 Flash Next install on a 5090 + 64gb DDR5 system — r/LocalLLM
  2. Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task. — r/LocalLLaMA
  3. Qwen3.8-Flash-Next on 12GB VRAM - 65 t/s — r/LocalLLM
  4. Qwen3.8-Flash-Next (125B) at 12-15 tok/s on a 2021 32GB M1 Max — r/LocalLLaMA
  5. vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain by fxgsell · Pull Request #29520 · ggml-org/llama.cpp — r/LocalLLaMA
  6. Llama.cpp and new model releases ... is Great is the enemy of Good in the LLM world? — r/LocalLLM
  7. Final-year student in India trying to break into generative-model inference optimization — roadmap feedback? — r/MLQuestions
  8. Sharing my Qwen3.8-27B at 8-bit on 2x RTX 3090 with vLLM: 115 tok/s decode, ~1,780 tok/s prefill, 262K context (NVLink + DFlash2, full recipe and A/B numbers) — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →