95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090
Hello everyone! A little while back I posted about LlamAmpere , a fork of Llama.cpp with Ampere-specific improvements (though it is caught up to main and will support other hardware, too). Thank you to everyone that tried it out and shared back their results across the 30xx cards. I'm happy to shar…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-28 18:35 · r/LocalLLaMA
95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090