Natural Language Processing and NLP Tooling

57 repos across 3 sub-areas

Libraries, datasets, and frameworks for processing, analyzing, and understanding human language at scale. The cluster spans foundational NLP tasks like named entity recognition and tokenization, through to deep learning-based language models and prompt engineering. Central projects include AllenNLP (a deep learning NLP toolkit), spaCy lookup data and NLTK (classic NLP libraries), and PromptSource (a framework for working with prompting datasets), alongside benchmarks and specialized tools for languages like Thai. Most repos are Python-based, reflecting the dominance of Python in the NLP research and engineering community.