Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print.
So everyone has been yelling about how I should be using Gufo instead of halogen because it's open source and it's "just as good or better". Checking in on their GitHub (GitHub.com/gufo-org/gufo) got me immediately .. "Qwen 27B Q4: 70.56 tok/s single user, 123 tok/s with 8 users" on a strix halo de…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-02 05:36 · r/LocalLLM
Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print.
More stories
- Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. — r/LocalLLM
- Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4 — r/LocalLLaMA
- One .char model, Consistent face, body & cloths, now works in Comfy(Custom node & workflows) MinimaxH3 & Flux2 — r/comfyui
- llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp — r/LocalLLaMA
- Is all the work that's being put into Qwen3.8 Flash Next going to set us up for a very quick uplift to Qwen4? — r/LocalLLaMA
- ComfyUI Qwen image 2.1 Enhancer (Two nodes) — r/StableDiffusion
- I built Ninfer 4080 for 16GB class GPUs — r/LocalLLaMA
- How I use a local Qwen 27B for real work on home hardware (unscripted workflow demo) — r/LocalLLM
Get the daily brief of stories like this at 6:30 every morning →