AINewsnow

Used Opus 5.5 to optimize llama.cpp inference for Swift Qwen 3.8 27B Q6_K on RTX 5090 - decode 143 tok/s prefill 2840 tok/s

https://preview.redd.it/uflbt6x61tsh1.png?width=472&format=png&auto=webp&s=20499138331923a3d84505e7fa530d02d1a37978 https://preview.redd.it/toxiq4ah1tsh1.png?width=419&format=png&auto=webp&s=5cdcc3244accdeee500ddac8a5cac43e63c54851 https://preview.redd.it/21lkvj782tsh1.png?width=439&format=png&auto…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-01 07:09 · r/LocalLLM
    Used Opus 5.5 to optimize llama.cpp inference for Swift Qwen 3.8 27B Q6_K on RTX 5090 - decode 143 tok/s prefill 2840 tok/s

More stories

  1. Qwen flash next on 12+16gb vram, and 32gb ram viable? — r/LocalLLM
  2. What model sits between Qwen 3.8 27b and Flash next for coding? — r/LocalLLaMA
  3. vulkan: fuse qwen4exp's SCALE -> SIGMOID -> SCALE -> hc_post chain by fxgsell · Pull Request #29520 · ggml-org/llama.cpp — r/LocalLLaMA
  4. Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator — r/LocalLLM
  5. Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers — r/LocalLLM
  6. Qwen 3.8 is a workhorse — r/LocalLLaMA
  7. Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune) — r/LocalLLM
  8. Layer Extract & Layer Remove Loras For Qwen Image 2.1 — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →