AINewsnow

FreeToken vs llama.cpp on one RTX 3090: llama.cpp is 2–3× faster when the MoE fits in VRAM. On gpt-oss-120b (63 GB), FreeToken gets the first token out 7× faster at 32 concurrent users.

I've run the benchmark on a RTX 3090, 1024 tokens in / 256 out, concurrency 1–32. If the model fits on vRAM (Gemma-4-26B-A4B, byte-identical GGUF on both engines ): llama.cpp has 2.2–3.2× the throughput and 5–6× faster TTFT. FreeToken 0.1.2 can't keep 4-bit experts in VRAM at all, and it OOM'd at 8…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-01 14:11 · r/LocalLLaMA
    FreeToken vs llama.cpp on one RTX 3090: llama.cpp is 2–3× faster when the MoE fits in VRAM. On gpt-oss-120b (63 GB), FreeToken gets the first token out 7× faster at 32 concurrent users.

More stories

  1. Qwen 3.8 27B on a single 3090: 114 min solo, 43 min as a worker under a GPT 6.1 SOL orchestrator — r/LocalLLM
  2. OpenAI Says It Will Not Release Newest Astra A.I. Model Over Safety Concerns — New York Times Technology
  3. OpenAI’s Dots Are Always-On AI Agents—and Its Answer to Meta’s Muse — Wired AI
  4. GPT-6 SOL AND LUNA ARE OUT!!! — Matthew Berman
  5. GPT 6.1 Artificial Analysis - Intelligence Index — r/ChatGPT
  6. Qwen flash next on 12+16gb vram, and 32gb ram viable? — r/LocalLLM
  7. Sonnet 5.5 orchestrated a local Qwen 3.8 27B! — r/ClaudeAI
  8. add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →