42 repos across 2 sub-areas
Libraries, datasets, and benchmarking frameworks for natural language processing tasks including text classification, semantic similarity, sentence embeddings, and structured NLP pipelines. The cluster spans from low-level linguistic tools (like spaCy and AllenNLP for parsing and annotation) to high-level task frameworks (like PromptSource for prompt-based learning) and multilingual resources (Thai sentence vectors, cross-lingual benchmarks). Most repos are Python-based tools and Jupyter notebooks exploring NLP methods empirically.
Cluster 462198
40 repos
Natural Language Processing and Language Models
2 repos
Libraries, datasets, and tools for building NLP systems, language detection, and large language model applications. The cluster spans multilingual text processing (lingua-py for language identification), training datasets (Glot500, datasets), and JavaScript/Python frameworks for NLP tasks (nlp.js). Most repos are Python-based, reflecting the dominance of Python in the deep learning and PyTorch ecosystems that power modern NLP.