ByteShape Squeezes Qwen3.8-27B Into 8.8 GB Hitting 176 Tokens per Second
ByteShape ships ShapeLearn-quantized GGUFs of Qwen3.8-27B that fit on 12 GB GPUs and hit 176 tokens per second on an RTX 5090.
Read the full story at AlphaSignal ↗
Timeline · 1 report
- 2026-10-04 16:01 · AlphaSignal
ByteShape Squeezes Qwen3.8-27B Into 8.8 GB Hitting 176 Tokens per Second