Large Language Model Context Extension

10 repos

Techniques and implementations for extending the context window length of large language models, enabling them to process longer sequences of text. The cluster centers on YARN (Yet Another RoPE extensioN), a method for extending RoPE-based positional encodings, applied across multiple Llama and Mistral model variants at different scales. These repositories represent fine-tuned and quantized model checkpoints optimized for inference, with implementations compatible with text-generation frameworks like vLLM.

Python · 1
custom_code ·890
text-generation ·890
text-generation-inference ·890
transformers ·890
endpoints_compatible ·890
pytorch ·890
en ·692
mistral ·627
llama ·263