Cluster 647803

64 repos across 4 sub-areas

Multilingual Legal RoBERTa Models

25 repos

Pre-trained RoBERTa transformer models fine-tuned on legal text corpora across diverse languages including Polish, Maltese, Hungarian, Slovak, Slovenian, Latvian, and others. These repositories represent specialized language models built with the Hugging Face transformers framework, designed for legal domain NLP tasks like document classification and named entity recognition. The cluster demonstrates a systematic effort to extend legal AI capabilities beyond English to support multilingual legal document understanding and analysis.

RoBERTa Sentence Embeddings

20 repos

Fine-tuned RoBERTa models for generating sentence-level embeddings using various pooling strategies (CLS token, mean pooling, max pooling). These repositories provide pre-trained transformer-based endpoints compatible with standard NLP pipelines, built on PyTorch and Hugging Face Transformers. The cluster represents practical implementations of sentence representation learning, useful for semantic similarity, retrieval, and downstream NLP tasks.

RoBERTa Language Models & NLP

18 repos

Pre-trained transformer models based on RoBERTa architecture, primarily trained on Twitter data across multiple time periods (2021–2022). These repositories provide masked language modeling capabilities and PyTorch-compatible endpoints for downstream NLP tasks. The cluster centers on incremental and periodic updates to Twitter-specific RoBERTa checkpoints, useful for fine-tuning on social media text and other domain-specific natural language processing applications.

Cluster 647805

1 repos