AINewsnow

Don't you feel scammed by Nvidia with them hiding P2P behind just a dozen lines of code?

This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.

So you spend your hard earn money to get some 50xx GPUs and months down the line you discover you can enable P2P with just a dozen line code change on the drivers which magically makes llama.cpp "split-mode: tensor" make the GPUs work better and less laboured (which means they'll probably last long…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-04 19:25 · r/LocalLLM
    Don't you feel scammed by Nvidia with them hiding P2P behind just a dozen lines of code?

More stories

  1. Beginner confused about Ollama vs LM Studio vs llama.cpp vs vLLM vs Unsloth — can someone explain? — r/LocalLLM
  2. Which models you run on your Nvidia v100? — r/LocalLLM
  3. Am I right in thinking llama.cpp is the only show in town for mixed (Nvidia) GPUs? — r/LocalLLM
  4. qwen4exp: add hc ops by am17an · Pull Request #28901 · ggml-org/llama.cpp — r/LocalLLaMA
  5. Huang and Zuckerberg back AI safety without a slowdown: Why Big Tech is betting on market forces to police AI — Mint AI
  6. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  7. Made a tool that tells you which GGUF quants will fit your GPU/Mac, with the llama.cpp command to run them — r/LocalLLM
  8. Connected a local model (Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_M) to GIMP via MCP tools using llama.cpp - and here's the image result from my first prompt "can you draw a picture of a flower in gimp?". Needs work. Setup follows. — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →