AINewsnow

Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print.

So everyone has been yelling about how I should be using Gufo instead of halogen because it's open source and it's "just as good or better". Checking in on their GitHub (GitHub.com/gufo-org/gufo) got me immediately .. "Qwen 27B Q4: 70.56 tok/s single user, 123 tok/s with 8 users" on a strix halo de…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-02 05:36 · r/LocalLLM
    Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print.

More stories

  1. Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context. — r/LocalLLM
  2. Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4 — r/LocalLLaMA
  3. One .char model, Consistent face, body & cloths, now works in Comfy(Custom node & workflows) MinimaxH3 & Flux2 — r/comfyui
  4. llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp — r/LocalLLaMA
  5. Is all the work that's being put into Qwen3.8 Flash Next going to set us up for a very quick uplift to Qwen4? — r/LocalLLaMA
  6. ComfyUI Qwen image 2.1 Enhancer (Two nodes) — r/StableDiffusion
  7. I built Ninfer 4080 for 16GB class GPUs — r/LocalLLaMA
  8. How I use a local Qwen 27B for real work on home hardware (unscripted workflow demo) — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →