AINewsnow

Gemma4 E4B EXL3 - Alternative to Llama.cpp on Jetson Orin

My Jetson Orin–optimized engine, little-gemma V1.0, substantially outperforms llama.cpp. Even after exhausting every practical GGUF option, however, it still falls short of the ideal performance level for Gemma E4B on the Jetson Orin Nano Super 8GB. I forked exllamav3 and ported key code from littl…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-11 19:47 · r/LocalLLaMA
    Gemma4 E4B EXL3 - Alternative to Llama.cpp on Jetson Orin

More stories

  1. Qwen3.6 35B-A3B optimized for a 4 GB GTX 1050 Ti Mobile: up to 2× prefill and 3× decode vs llama.cpp — r/LocalLLM
  2. Benchmarked 4 Local Models on 3 Harnesses — r/LocalLLM
  3. What local AI and software do you use? Which is best? — r/LocalLLM
  4. Pancho the Llama talks (Part II) — r/comfyui
  5. I built an open-source GPU orchestrator that lets several local AI models share one GPU by loading/unloading them on demand (llama.cpp, Ollama, ComfyUI, TTS) — r/LocalLLM
  6. Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM — r/LocalLLaMA
  7. 7900 XTX swap vs. adding a 7600 XT — r/LocalLLM
  8. If you used an Agentic AI to help design an agentic AI for working on your personal task and projects, replying to emails, etc. What did your final prompt and build instructions look like? — r/ArtificialInteligence

Get the daily brief of stories like this at 6:30 every morning →