Gem16 - custom engine for Gemma4 12B & 26B on Blackwell 16GB GPUs
It’s probably a bit niche and the models are a bit old at this point, but after reading about Ninfer a few months ago I did my own small vibe coded engine project for my 5080 Laptop GPU. Initially I thought I can only fit the 12B model with enough context into the VRAM, but with custom quantisation…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-27 22:02 · r/LocalLLaMA
Gem16 - custom engine for Gemma4 12B & 26B on Blackwell 16GB GPUs