AINewsnow

[Benchmark] llama.cpp batch/ubatch impacts on PP and TG

This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.

After 7 execution days (full time), I now have the values! My test is running DeepSeek v4 Flash 0731 at native size on DGX Spark machine (GB10, 128 GB unified memory). The model size is bigger than RAM, so weights will be loaded many times from the SSD when running. To improve the speed, weights ha…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-02 10:36 · r/LocalLLM
    [Benchmark] llama.cpp batch/ubatch impacts on PP and TG

More stories

  1. OpenAI solves Navier-Stokes, Meta’s Muse a free AI agent that’s really good, DeepSeek V4.1 shrinks KV cache, and one doomer post causes OpenAI to consider pausing training + more AI news — ThursdAI
  2. Cactus Needle 3: A Sliceable 8-29MB Automation Foundation Model That Matches DeepSeek v4 Flash — r/LocalLLaMA
  3. Beginner confused about Ollama vs LM Studio vs llama.cpp vs vLLM vs Unsloth — can someone explain? — r/LocalLLM
  4. DeepSeek’s Insane New Architecture — Two Minute Papers
  5. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  6. Made a tool that tells you which GGUF quants will fit your GPU/Mac, with the llama.cpp command to run them — r/LocalLLM
  7. Jina AI Releases jina-ocr-v1: A 3.4B MoE Document Parser With Built-In Speculative Decoding for Low-Budget GPUs — MarkTechPost
  8. Gemini 4 is good enough - JUST RELEASE IT — r/GeminiAI

Get the daily brief of stories like this at 6:30 every morning →