26 repos
Pretrained ELECTRA models for non-English languages, including German, Turkish, and French variants. These repositories contain discriminator and generator checkpoints, along with resources for efficient language model pretraining using the ELECTRA framework—which uses replaced token detection rather than masked language modeling. The cluster focuses on making state-of-the-art transformer pretraining accessible across diverse linguistic communities and datasets like Europeana and MC4.