AI glossary Models & Architecture
What is Quantization?
← All glossary termsReducing numeric precision of weights or activations to shrink models and speed inference.
Explained
Quantization (e.g., INT8, INT4) lowers memory footprint and accelerates inference on GPUs, TPUs, or NPUs. It is central to cost-efficient LLM deployment at scale.