Model Quantization and Compression

12 repos

Techniques and implementations for reducing neural network model size and computational requirements through quantization, knowledge distillation, and weight compression. The cluster centers on practical applications of these methods to vision-language models like Qwen and LLaVA, with a focus on INT4 quantization-aware training and AWQ (Activation-aware Weight Quantization) approaches. Repositories here demonstrate how to efficiently deploy large pre-trained models on resource-constrained hardware while maintaining model performance.

Python · 3
Jupyter Notebook · 1
knowledge-distillation ·2,923
int4 ·2,923
quantization-aware-training ·2,922
awq ·2,918
int8 ·2,710
quantization ·2,710
large-language-models ·2,709
low-precision ·2,709
fp4 ·2,709
gptq ·2,709