How I Took GLiClass FP8 from 59 ms to 16 ms on an RTX 4050
Multilingual zero-shot classification on a 6 GB laptop GPU, with native FP8, Triton, CUDA Graphs, and reproducible measurements.
I quantized Knowledgator’s GLiClass Multilang Ultra into a custom FP8 W
badbat4560.hashnode.dev10 min read