Single 3060 40 t/s code with 80k context building stuff
I have been working on a project for the past Month or more to make my 12gb cards much more useful in my homelab. I landed on making a custom Hybrid model using Qwen3.8-27b. I combined the body of Escha Labs W2 with an IQ3 head from the Lowgpu xxxs quant and ended up with something very smart and 8…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-03 14:51 · r/LocalLLM
Single 3060 40 t/s code with 80k context building stuff