AINewsnow

From 1x3090 to 20 DGX Sparks: my house fuses were the first bottleneck

​ From the first LLaMA 33B I knew I wanted that magic-like intelligence locally, mine, so nobody could take it away when I needed it. I bought a 3090 for my home PC. Then LLaMA 65B appeared and I was dazzled, it looked like it had all the knowledge in the world. I made two copies, one local…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-04 14:09 · r/LocalLLaMA
    From 1x3090 to 20 DGX Sparks: my house fuses were the first bottleneck

More stories

  1. RTX 5070 TI and 32GB RAM DDR4, WHAT IS THE BEST VIDEO MODEL I COULD USE? — r/comfyui
  2. llama, server: add /v1/systemone API (models: laya, julia-1, lev, openjev, kev) by ngxson · Pull Request #29818 · ggml-org/llama.cpp — r/LocalLLaMA
  3. Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant — r/LocalLLaMA
  4. Strata looping badly with iq2_xxs — r/LocalLLM
  5. I want to hook llama.cpp up to a search provider, both locally hosted. I am hitting nothing but dead-ends. Help? — r/LocalLLM
  6. Need maybe say "Use llama.cpp" — r/LocalLLaMA
  7. I got llama.cpp inference running on the Snapdragon 8 Gen 3 Hexagon NPU from non-root Termux + Adreno OpenCL results (S24 Ultra) — r/LocalLLM
  8. Strata takes the promise of "MoE models just need a total amount of VRAM+RAM" and makes it a reality — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →