Local LLMs: How much does good prompting + tooling actually close the gap with frontier models for mid-sized (27B–35B) weights?
I’ve been staring at LLM leaderboards and benchmark stats way too much lately, and I’ve reached a bit of a frustrating conclusion. It feels like if you want to run something locally at a usable speed, you're either forced to go really small (and deal with models that noticeably lag behind the front…
Read the full story at r/LocalLLM ↗