[Benchmark] llama.cpp batch/ubatch impacts on PP and TG
This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.
After 7 execution days (full time), I now have the values! My test is running DeepSeek v4 Flash 0731 at native size on DGX Spark machine (GB10, 128 GB unified memory). The model size is bigger than RAM, so weights will be loaded many times from the SSD when running. To improve the speed, weights ha…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-02 10:36 · r/LocalLLM
[Benchmark] llama.cpp batch/ubatch impacts on PP and TG