add GLM-5.3-Flash (GLM5-Next) support (#27773) · ggml-org/llama.cpp@649dcb1
Finally!
Read the full story at r/LocalLLaMA ↗
Timeline · 5 reports
- 2026-10-03 06:18 · r/LocalLLaMA
qwen4exp : halve the indexer score memory by ServeurpersoCom · Pull Request #29825 · ggml-org/llama.cpp - 2026-10-02 18:47 · r/LocalLLaMA
CUDA: fuse shared experts into MMVQ by am17an · Pull Request #29184 · ggml-org/llama.cpp - 2026-10-02 10:29 · r/LocalLLaMA
llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp - 2026-10-01 11:18 · r/LocalLLaMA
Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp - 2026-09-30 12:41 · r/LocalLLaMA
add GLM-5.3-Flash (GLM5-Next) support (#27773) · ggml-org/llama.cpp@649dcb1
More stories
- Qwen 3.8 Flash Next - doubled Strata throughput on 3090+5070 Ti, IQ3_S 2466 pp/167 tps, UD-Q4_K_XL 2341 pp / 126 tps (yes, really) — r/LocalLLM
- Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now — r/LocalLLaMA
- Claude and Grok built me a local monitoring setup for my AI box: two dashboards, one for the machine, one for model training — r/LocalLLM
- Cloudflare says its new Clef model means humans no longer need to be in the loop for AI agents — The Decoder
- I benchmarked Jev 1.13 against 4 local LLMs on RTX5070 12GB - amazing. — r/singularity
- Who’s the current “king” of local LLMs for you — Qwen, Gemma, Llama, something else? — r/LocalLLM
- Which LLM is best for coding/agents if I have dual R9700 GPUs? — r/LocalLLM
- Used Opus 5.5 to optimize llama.cpp inference for Swift Qwen 3.8 27B Q6_K on RTX 5090 - decode 143 tok/s prefill 2840 tok/s — r/LocalLLM
Get the daily brief of stories like this at 6:30 every morning →