AINewsnow

$5400 eBay 8x V100 server cranks on flash-next

Had opus 5.5 do the bring up.. super solid results. Using only 4 GPU results in KV of like 120k with image on. 27B TP=4 results in >200 tok/s with dflash, prefill around 2.5-3.5k. Both these are running the nvfp4 Nvidia checkpoints. Found a magic repo that unpacks into fp16 on the fly ( https://git…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-10-01 13:44 · r/LocalLLaMA
    $5400 eBay 8x V100 server cranks on flash-next

More stories

  1. Cognition Becomes First Customer for NVIDIA Vera Rubin NVL72 on CoreWeave Cloud — CoreWeave Blog
  2. Top AI and tech firms sign 'morally binding' accord to 'self-police' development after meeting at White House — Euronews Next
  3. Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples — NVIDIA Technical Blog
  4. Build agent memory with NVIDIA NeMo Agent Toolkit and Amazon S3 Vectors — AWS Machine Learning Blog
  5. China’s DeepSeek open-sources tools to help Huawei chips supplant Nvidia in AI — South China Morning Post Tech
  6. Nvidia says new tool can contain rogue AI agents in "milliseconds" — Axios AI+
  7. Productive, Durable, Fungible: How NVIDIA AI Factories Maximize Return on Investment — NVIDIA Blog
  8. Deploying an HSTU Generative Recommender with NVIDIA Dynamo-Triton — NVIDIA Technical Blog

Get the daily brief of stories like this at 6:30 every morning →