FreeToken vs llama.cpp on one RTX 3090: llama.cpp is 2–3× faster when the MoE fits in VRAM. On gpt-oss-120b (63 GB), FreeToken gets the first token out 7× faster at 32 concurrent users.
I've run the benchmark on a RTX 3090, 1024 tokens in / 256 out, concurrency 1–32. If the model fits on vRAM (Gemma-4-26B-A4B, byte-identical GGUF on both engines ): llama.cpp has 2.2–3.2× the throughput and 5–6× faster TTFT. FreeToken 0.1.2 can't keep 4-bit experts in VRAM at all, and it OOM'd at 8…
Read the full story at r/LocalLLaMA ↗