Question from an uneducated person...don't kill me
Could a neural network use dictionary-compressed weights directly during GPU inference instead of fully decoding them first? I'm not a computer scientist. I'm a truck driver, so I'm wondering if I'm reinventing something that already exists. Suppose you quantize a model to INT4 or similar and then…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-09-29 03:00 · r/learnmachinelearning
Question from an uneducated person...don't kill me