18 repos
Lossless compression algorithms and techniques applied to neural network models, particularly large language models (LLMs) and diffusion models optimized for GPU deployment. The cluster centers on model compression, optimization, and efficient inference, with a focus on practical implementations that reduce model size while maintaining performance. Repositories here include quantized and distilled variants of models like Qwen and DeepSeek, demonstrating how compression approaches enable deployment of sophisticated neural architectures on resource-constrained hardware.