HomeAI glossaryQuantization

Modified

Quantization

In Swedish: Kvantisering

The process of compressing an AI model by reducing the precision of the numbers inside it. A model is quantized to Q4 (4-bit), Q8 (8-bit), or other levels to fit smaller GPU memory. Lower precision = smaller file and faster execution, but marginally lower quality.

Domain:OptimeringInfrastruktur

Browse the full glossary · 243 terms in Swedish and English