NLP Tasks, Datasets & Benchmarks

42 repos across 2 sub-areas

Libraries, datasets, and benchmarking frameworks for natural language processing tasks including text classification, semantic similarity, sentence embeddings, and structured NLP pipelines. The cluster spans from low-level linguistic tools (like spaCy and AllenNLP for parsing and annotation) to high-level task frameworks (like PromptSource for prompt-based learning) and multilingual resources (Thai sentence vectors, cross-lingual benchmarks). Most repos are Python-based tools and Jupyter notebooks exploring NLP methods empirically.