AINewsnow

Cómo correr un modelo de 180B sin GPU: POCKET-Darwin-180B en GGUF con llama.cpp

This story is from 2026-10-06. It is preserved in the archive; the latest stories are on the live feed.

TL;DR POCKET-Darwin-180B-GGUF es la versión cuantizada del modelo abierto Darwin-180B-RSI, pensada para correr sin GPU. Puntos clave para quien va a desplegarlo: Pesa 111 GB en GGUF (el original en BF16 ocupa 360 GB) y se ejecuta con llama.cpp . Hasta 21 tokens por segundo solo en CPU (1 socket, 16…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-10-06 13:12 · DEV Community — AI
    Cómo correr un modelo de 180B sin GPU: POCKET-Darwin-180B en GGUF con llama.cpp

More stories

  1. Is all the work that's being put into Qwen3.8 Flash Next going to set us up for a very quick uplift to Qwen4? — r/LocalLLaMA
  2. LLM Inference Dashboard — r/LocalLLaMA
  3. Gemma 4 26B-A4B and a 37 GB Qwen3.6 MoE running in a browser tab on a 24 GB Mac — experts streamed from disk, output matches llama.cpp — r/LocalLLaMA
  4. Overclocking DDR5 For Faster MoE Prefill and Decode — r/LocalLLaMA
  5. Uniform GGUF quants silently break Qwen3.8-27B's deep thinking — reproduced on llama.cpp AND vLLM (short tasks unaffected) — r/LocalLLM
  6. llama.cpp v0.6.0 released with MTP speculative decoding for Qwen4Exp and lots more — r/LocalLLaMA
  7. Llama.cpp + WebGPU = agants.html — r/AI_Agents
  8. SkyIsNotGreen/Scion-35B-A3B · Hugging Face - Ternary MoE — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →