AINewsnow

What quant Qwen3.8 should I be using on a 5090? (Unsloth Studio)

This story is from 2026-09-06. It is preserved in the archive; the latest stories are on the live feed.

So this is confusing me a bit, and I'm not sure where the numbers are coming from. 5090 has 32gb of vram. I know that the model has to be in vram to keep speeds reasonable, and context also needs to be in vram where possible. And for things like coding and agentic tasks, I want/need as much context…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-06 04:46 · r/LocalLLM
    What quant Qwen3.8 should I be using on a 5090? (Unsloth Studio)

More stories

  1. Anthropic says Claude 'leads' 26 percent of its AI R&D work — Engadget
  2. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog
  3. Optimizing agent system prompts with Amazon Bedrock AgentCore — AWS Machine Learning Blog
  4. Introducing Amazon SageMaker HyperPod Inference Gateway — AWS Machine Learning Blog
  5. Introducing Astra for Law — OpenAI News
  6. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  7. OpenAI reveals cases of ‘concerning’ AI behaviour as it announces new disclosure system — The Guardian AI
  8. What It Takes to Bring Up a Multi-Rack NVIDIA Vera Rubin NVL72 Cluster — CoreWeave Blog

Get the daily brief of stories like this at 6:30 every morning →