2x P100 What is possible with 3.8 27b
This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.
I currently have a P100 16GB + RTX 3070 8GB running Qwen3.8 27B GGUF with llama.cpp. I am thinking about replacing the 3070 with a second P100. Right now I am getting about 8 tok/s generation at 120K context, with prompt processing around 80 to 100 tok/s. My max context is about 160K. Has anyone he…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-14 13:34 · r/LocalLLM
2x P100 What is possible with 3.8 27b