10 repos
Techniques and implementations for extending the context window length of large language models, enabling them to process longer sequences of text. The cluster centers on YARN (Yet Another RoPE extensioN), a method for extending RoPE-based positional encodings, applied across multiple Llama and Mistral model variants at different scales. These repositories represent fine-tuned and quantized model checkpoints optimized for inference, with implementations compatible with text-generation frameworks like vLLM.