AINewsnow

Gufo: the all-in-one strix halo inference engine

Hello, we are happy to share a project that a friend of mine and I have been working on for a while. It's an extremely optimized inference engine for the Strix Halo platform. We both have a Framework Desktop with 128 GB of RAM and were tired of juggling multiple forks of llama.cpp and other project…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-24 09:37 · r/LocalLLM
    Gufo: the all-in-one strix halo inference engine

More stories

  1. Transformers now runs llama.cpp quants — Hugging Face Blog
  2. yandex/AliceAI-Foundation-80B-A3B-Base: Russian-developed competitor to Qwen 35B and DeepSeek V4 Flash — r/LocalLLaMA
  3. I trained a 360M-param Python model from scratch on two workstation GPUs and wrote up every step, including the bugs — r/learnmachinelearning
  4. Performance tune for gemma4-26b-a4b flash attention shape. by frobnitzem · Pull Request #28450 · ggml-org/llama.cpp · GitHub — r/LocalLLaMA
  5. My foray into local ai. Two BC-250 ex mining apus running Qwen3.6-35B-A3B Q4_K_M at 60 tok/s with 64k context — r/LocalLLaMA
  6. Stanford's MAttr Tops AI Interpretability Benchmark by Nearly 3x — AlphaSignal
  7. I turned Qwen3.8-27B Q2_64 + llama.cpp into a fully TypeSafe AI-compatible Jev-like system. OpenAI API still intact! World’s first Vision-enabled Jev-like model! <10 GB VRAM, 170 ms on an RTX 3090 and ~140 tok/s in chat. 76% vs. 88% Jev-1.13 Acc. on a diverse 22,000-request typed-decision benchmark — r/LocalLLaMA
  8. GGUFs in transformers natively! — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →