Sparse LLaMA Model Variants

12 repos

Optimized quantized and sparse versions of the LLaMA 2 language model at 7B and 13B parameter scales. The cluster centers on GGUF-formatted model files and associated predictor/inference tools for running these models efficiently on resource-constrained hardware. These repositories represent different compression and sparsification approaches applied to the same base LLaMA 2 architecture, enabling practical deployment of large language models.

en ·42
transformers ·41
gguf ·41
endpoints_compatible ·37
llama ·15
relullama ·11
bamboo ·11
PowerInfer ·5
feature-extraction ·4
sparsellama ·4