12 repos
Optimized quantized and sparse versions of the LLaMA 2 language model at 7B and 13B parameter scales. The cluster centers on GGUF-formatted model files and associated predictor/inference tools for running these models efficiently on resource-constrained hardware. These repositories represent different compression and sparsification approaches applied to the same base LLaMA 2 architecture, enabling practical deployment of large language models.