Multilingual Abusive Language Detection

30 repos

NLP systems for detecting abusive and offensive content across multiple languages and code-mixed text, primarily using transformer-based models like BERT and T5. The cluster focuses on text classification approaches for identifying harmful language in English, Hindi, Bengali, Urdu, Kannada, and related language variants, with applications to content moderation and safer online spaces. Most repos center on dataset creation, model training, and evaluation frameworks for these multilingual and code-switched abuse detection tasks.

Python · 6
text-classification ·6,030
bert ·6,030
pytorch ·6,028
transformers ·4,118
nlp ·4,115
deep-learning ·2,789
pretrained-models ·2,789
seq2seq ·2,633
knowledge-pretraining ·2,186
fewshot-learning ·2,186