AINewsnow

What GPUs will give me GOOD speeds and on DSV4 Flash and similar models, and not have to run a mega quantized version? Budget around $15k-ish.

This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.

I wish I could spend $15k on my own homelab hardware, but no this is for work lol. Like the title says, we're looking to run DSV4 Flash (and similar tier models) locally at good speeds, both for token gen and prompt processing. By "good" I'm thinking in the range of 40-50+ t/s gen and at least 1000…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-11 18:36 · r/LocalLLaMA
    What GPUs will give me GOOD speeds and on DSV4 Flash and similar models, and not have to run a mega quantized version? Budget around $15k-ish.

More stories

  1. Google Joins OpenAI, Anthropic, Meta in Disclosing AI Hacks — Bloomberg AI
  2. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
  3. Introducing Amazon SageMaker HyperPod Inference Gateway — AWS Machine Learning Blog
  4. Introducing Astra for Law — OpenAI News
  5. Newsom signs executive order to explore new AI rules, consider ‘kill switch’ — Politico Technology
  6. Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
  7. Sources: Anthropic considers releasing a new AI model to counter OpenAI's momentum since Astra's launch, ahead of an IPO and after Amodei's call for a slowdown (Reuters) — Techmeme
  8. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology

Get the daily brief of stories like this at 6:30 every morning →