If you have a 3090, or other 30xx for local LLMs, I have something for you
This story is from 2026-09-15. It is preserved in the archive; the latest stories are on the live feed.
I have a custom fork of llama.cpp designed around the ampere architecture specifically (though many of the upgrades also translate to faster performance of blackwell + lovelace). The recommended config supports 90+ TPS (for agentic/coding, at temp 1; greedy will of course be faster) through 100K to…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-15 01:32 · r/LocalLLaMA
If you have a 3090, or other 30xx for local LLMs, I have something for you