GLM-5.3-Flash Benchmarks on TensorSharp and llama.cpp
This story is from 2026-08-28. It is preserved in the archive; the latest stories are on the live feed.
Model: unsloth/GLM-5.3-Flash-GGUF UD-Q2_K_XL (4 shards, 101 GiB) + mmproj-BF16. Reference: llama.cpp PR #27754 (glm5next), CUDA 12.8, SM 120. Throughput Both engines back to back in one session, flash attention on, n_ubatch 2048 on both ( llama-bench vs the parity harness --bench ). Run-to-run spre…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-28 17:08 · r/LocalLLaMA
GLM-5.3-Flash Benchmarks on TensorSharp and llama.cpp