Long-context Language Models in JAX

10 repos

Language models optimized for processing very long input sequences (128K to 1M tokens), implemented in JAX. These repositories explore efficient architectures and training approaches for extended-context applications, including both base and chat-tuned variants. The cluster focuses on pushing the practical limits of context length in transformer-based models while maintaining computational efficiency.

Related clusters