17 repos
PyTorch implementations of ModernBERT, a contemporary take on BERT-style transformers with fill-mask pretraining capabilities. The cluster centers on a family of encoder and decoder model variants at different scales (17M to 400M parameters), built on modern transformer architectures and designed for masked language modeling tasks. Repositories here provide reference implementations, pretrained weights, and infrastructure for working with these models in the transformers ecosystem.