Infatoshi Squeezes GLM-5.3's 753B Parameters Into 273 GiB for Multi-GPU Workstations
A 3-bit EXL3 quant squeezes a 753B uncensored GLM-5.3 Mixture of Experts into 273 GiB, making local inference possible on four RTX PRO 6000s.
Read the full story at AlphaSignal ↗
Timeline · 1 report
- 2026-10-09 09:27 · AlphaSignal
Infatoshi Squeezes GLM-5.3's 753B Parameters Into 273 GiB for Multi-GPU Workstations