10 repos
Language models optimized for processing very long input sequences (128K to 1M tokens), implemented in JAX. These repositories explore efficient architectures and training approaches for extended-context applications, including both base and chat-tuned variants. The cluster focuses on pushing the practical limits of context length in transformer-based models while maintaining computational efficiency.