Given how common RTX 3090 use is for LLMs, why don't we see more INT8 W8A8 models ?
This story is from 2026-09-01. It is preserved in the archive; the latest stories are on the live feed.
Based on https://huggingface.co/hardware , the RTX 3090 is the second most used GPU by LLM enthusiasts. Because RTX 3090 has native INT8 tensors cores, it can provide better performance with INT8 W8A8. However people seems to default to FP8 or smaller quants anyway. I suppose I am missing informati…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-01 22:27 · r/LocalLLaMA
Given how common RTX 3090 use is for LLMs, why don't we see more INT8 W8A8 models ?