Benchmark GLM 5.2 Unsloth GGUF model on TensorSharp
This story is from 2026-08-20. It is preserved in the archive; the latest stories are on the live feed.
I've been working on GLM-5.2 support in TensorSharp, and I finally have some back-to-back performance numbers against llama.cpp. The setup: Model: GLM-5.2-UD-IQ2_XXS (~226 GiB) GPUs: 3× RTX PRO 6000 Blackwell, 97 GiB each Distribution: layer split across all 3 GPUs Same machine, same session llama.…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-20 06:25 · r/LocalLLM
Benchmark GLM 5.2 Unsloth GGUF model on TensorSharp