Quick Bench Report: Qwen3.8 Flash Next on 2xV100, NVLink
This story is from 2026-09-07. It is preserved in the archive; the latest stories are on the live feed.
Running a Q5_K_M with ngram table in system ram - everything else on the v100s, just barely fits with 160k context length topping out the window - 32gb on each card, 64gb total. KV cache is full fp16. starts around 30 tps decode at 0, near 160k context length drops to about 15 starts around 700 pre…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-07 00:01 · r/LocalLLM
Quick Bench Report: Qwen3.8 Flash Next on 2xV100, NVLink
More stories
- Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
- Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
- Introducing Amazon SageMaker HyperPod Inference Gateway — AWS Machine Learning Blog
- Anthropic, OpenAI, SpaceXAI, Google sued over call to ‘pace’ AI development — Politico Technology
- NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
- AI hallucination of Chinese nuclear components almost led to US military attack — Ars Technica AI
- The new AgentCore runtime: Elastic, optimized, and consistently fast starts — AWS Machine Learning Blog
- Qwen Image 2.1 PR to ComfyUI — r/StableDiffusion
Get the daily brief of stories like this at 6:30 every morning →