Finally found a model my hardware can run at full precision: me
Was getting bored trying to squeeze every last t/s out of my local model on my hardware, so I made a tps counter for my fingers instead (with real tokenizers, of course). My best is around 2 t/s. According to the page, that beats a 70B on a laptop CPU and is roughly 76x slower than an 8B on a 4090.…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-28 15:55 · r/LocalLLaMA
Finally found a model my hardware can run at full precision: me