Gemma4 E4B EXL3 - Alternative to Llama.cpp on Jetson Orin
My Jetson Orin–optimized engine, little-gemma V1.0, substantially outperforms llama.cpp. Even after exhausting every practical GGUF option, however, it still falls short of the ideal performance level for Gemma E4B on the Jetson Orin Nano Super 8GB. I forked exllamav3 and ported key code from littl…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-10-11 19:47 · r/LocalLLaMA
Gemma4 E4B EXL3 - Alternative to Llama.cpp on Jetson Orin