How I Took GLiClass FP8 from 59 ms to 16 ms on an RTX 4050
Multilingual zero-shot classification on a 6 GB laptop GPU, with native FP8, Triton, CUDA Graphs, and reproducible measurements. I quantized Knowledgator’s GLiClass Multilang Ultra into a custom FP8 W8A8 checkpoint. Then I wanted to find out how much of that compression could translate into faster…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-27 14:40 · DEV Community — Machine Learning
How I Took GLiClass FP8 from 59 ms to 16 ms on an RTX 4050