AINewsnow

Running 100B+ MoE Models on a Single RTX 4090: A Practical Guide to Expert Offloading with llama.cpp

Tuần này trên Hacker News có một bài hơn 600 điểm: chạy một model MoE 125B tham số trên một con RTX 4090 mà vẫn đạt tốc độ sinh token rất đáng nể. Nghe như chuyện đùa, vì 4090 chỉ có 24GB VRAM, trong khi 125B tham số ở mức quantize 4-bit đã chiếm khoảng 70-75GB. Bí quyết không nằm ở phép màu nào cả…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-10-05 01:58 · DEV Community — Machine Learning
    Running 100B+ MoE Models on a Single RTX 4090: A Practical Guide to Expert Offloading with llama.cpp

More stories

  1. Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4 — r/LocalLLaMA
  2. llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp — r/LocalLLaMA
  3. GLM-4.7 benchmark compared MXFP4 vs Q4_K_M vs Q4_K_XL using Radeon 6800H iGPU 680M — r/LocalLLaMA
  4. LLM Inference Dashboard — r/LocalLLaMA
  5. Strata looping badly with iq2_xxs — r/LocalLLM
  6. I want to hook llama.cpp up to a search provider, both locally hosted. I am hitting nothing but dead-ends. Help? — r/LocalLLM
  7. From 1x3090 to 20 DGX Sparks: my house fuses were the first bottleneck — r/LocalLLaMA
  8. Need maybe say "Use llama.cpp" — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →