AINewsnow

Given how common RTX 3090 use is for LLMs, why don't we see more INT8 W8A8 models ?

This story is from 2026-09-01. It is preserved in the archive; the latest stories are on the live feed.

Based on https://huggingface.co/hardware , the RTX 3090 is the second most used GPU by LLM enthusiasts. Because RTX 3090 has native INT8 tensors cores, it can provide better performance with INT8 W8A8. However people seems to default to FP8 or smaller quants anyway. I suppose I am missing informati…

Read the full story at r/LocalLLaMA ↗

Timeline · 1 report

  1. 2026-09-01 22:27 · r/LocalLLaMA
    Given how common RTX 3090 use is for LLMs, why don't we see more INT8 W8A8 models ?

More stories

  1. Qwen Image 2.1 PR to ComfyUI — r/StableDiffusion
  2. we made a 27b model for creative writing. performs as good as claude fable 5, at a 40x cheaper price, open weights. — r/GeminiAI
  3. Deploy Hugging Face models on Amazon SageMaker AI with coding agents — AWS Machine Learning Blog
  4. A quick Minimax H3 news round-up - 18th September 2026 — r/comfyui
  5. Change the camera movement/angle for your existing video clip - Minimax H3 V2V CrossView-Warp LoRA — r/StableDiffusion
  6. Hugging Face Hack Shows Humans Can Keep AI In Check — AI Now Institute
  7. Qwen/Qwen-Image-2.1 · Hugging Face — r/StableDiffusion
  8. this looks promising: stepfun-ai/Step-5-Preview-BF16 · Hugging Face — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →