40 repos
Libraries, datasets, and tools for processing, analyzing, and understanding human language at scale. This cluster spans foundational NLP frameworks like spaCy and NLTK, transformer-based systems like AllenNLP, domain-specific implementations (Thai language processing, music information extraction), and benchmark datasets for evaluating model performance. Practitioners here work on tasks ranging from tokenization and morphological analysis to semantic similarity and structured information extraction.
bheinzerling/bpemb
Pre-trained subword embeddings in 275 languages, based on Byte-Pair Encoding (BPE)