AINewsnow

GLM-4.7 benchmark compared MXFP4 vs Q4_K_M vs Q4_K_XL using Radeon 6800H iGPU 680M

Using llama.cpp Ubuntu Vulkan prebuilt binary and the Acemagic miniPC S3A using an AMD Ryzen 7 6800H is a high-performance 8-core, 16-thread mobile processor launched on January 4, 2022, built on the 6nm Zen 3+ architecture loaded with 64GB of DDR5 RAM. It features a 3.2 GHz base clock, a 4.7 GHz b…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-04 21:08 · r/LocalLLaMA
    GLM-4.7 benchmark compared MXFP4 vs Q4_K_M vs Q4_K_XL using Radeon 6800H iGPU 680M

More stories

  1. Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4 — r/LocalLLaMA
  2. Prime Intellect Launches Prime Inference, Serving 600B Tokens Daily — AlphaSignal
  3. llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp — r/LocalLLaMA
  4. Two ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine — r/LocalLLaMA
  5. Running 100B+ MoE Models on a Single RTX 4090: A Practical Guide to Expert Offloading with llama.cpp — DEV Community — Machine Learning
  6. Can an Open Model Do Security Research? Cantina's apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks — MarkTechPost
  7. Fully local little parkour sim — r/LocalLLaMA
  8. LLM Inference Dashboard — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →