Finally!
So I have been tinkering. I think I got this V100 32gb dialed in (PG500-216 @ 185w, custom compile) Has Qwen3.8:27b - UD-Q4-KM (mtp) at 50 ish tps decode on as you can see a long string. Prefill dives fast on this, anyone have pointers on squeezing more prefill? Batch size at 4096 ubatch at 2048. (…
Read the full story at r/LocalLLM ↗
Timeline · 2 reports
- 2026-10-09 19:44 · r/GeminiAI
It’s finally here!!! 😮 - 2026-10-08 18:51 · r/LocalLLM
Finally!