Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print.
So everyone has been yelling about how I should be using Gufo instead of halogen because it's open source and it's "just as good or better". Checking in on their GitHub (GitHub.com/gufo-org/gufo) got me immediately .. "Qwen 27B Q4: 70.56 tok/s single user, 123 tok/s with 8 users" on a strix halo de…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-01 21:18 · r/LocalLLaMA
Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print.
More stories
- Strands Labs, AWS's experimental agent-development project, unveils Strands Decider 2B, a free, open-source Jev competitor fine-tuned from an Alibaba Qwen base (Carl Franzen/VentureBeat) — Techmeme
- Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune) — r/LocalLLM
- Qwen flash next on 12+16gb vram, and 32gb ram viable? — r/LocalLLM
- Sonnet 5.5 orchestrated a local Qwen 3.8 27B! — r/ClaudeAI
- Train Edit Loras for Qwen image 2.1 in Fizgig 6.6.0 — r/StableDiffusion
- add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp — r/LocalLLaMA
- Continuity update: screen replacement, image to 3D, Qwen Image 2.1, and a lot more since 3.0 — r/StableDiffusion
- Is anyone else running insanely long unattended loops? — r/AI_Agents
Get the daily brief of stories like this at 6:30 every morning →