78 repos across 3 sub-areas
RoBERTa sentence embeddings and transformers
34 repos
Transformer-based NLP models, primarily RoBERTa variants fine-tuned for sentence embedding tasks using different pooling strategies (mean, max, CLS token). These repositories focus on pre-trained language models compatible with the Hugging Face transformers library and PyTorch, with applications in semantic similarity and masked language modeling. The cluster represents modern approaches to generating dense vector representations from transformer encoders for downstream NLP tasks.
Multilingual Legal NLP Models
33 repos
Transformer-based language models fine-tuned for legal document processing across multiple European languages, including Hungarian, Slovak, Finnish, Slovenian, Polish, and Maltese. These models are built on RoBERTa architecture and distributed via Hugging Face's model hub, leveraging common tools like TensorBoard for training monitoring and safetensors for efficient model serialization. The cluster represents a focused effort to adapt pre-trained transformers for domain-specific (legal) and language-specific tasks in low-resource and mid-resource legal NLP contexts.
Cluster 462323
11 repos