Gemma 4 26B-A4B + MTP on RTX 5060 Ti 16GB (OCuLink) — Real-World 128k Window Logs (20W Idle / 150-200W Peak)
This story is from 2026-08-27. It is preserved in the archive; the latest stories are on the live feed.
I wanted to share real-world logs from my 24/7 homelab inference node after dialing in Multi-Token Prediction (MTP) and ngram-mod in llama.cpp (b10621). For private homelab applications—such as document RAG, agent workflows, and vision analysis—you don't necessarily need a multi-GPU workstation con…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-27 19:00 · r/LocalLLM
Gemma 4 26B-A4B + MTP on RTX 5060 Ti 16GB (OCuLink) — Real-World 128k Window Logs (20W Idle / 150-200W Peak)