2 repos
Libraries and tools for optimizing inference performance of large language models through techniques like quantization, pruning, and hardware acceleration. The cluster centers on frameworks like AutoGPTQ for post-training quantization, alongside educational resources and deployment utilities for running LLMs efficiently on resource-constrained hardware. Most repositories are Python-based implementations focused on the transformer inference pipeline.