Overview of Abusive and Threatening Language Detection in Urdu at FIRE 2021
arXiv · · Notable
Summary
This paper introduces two shared tasks for abusive and threatening language detection in Urdu, a low-resource language with over 170 million speakers. The tasks involve binary classification of Urdu tweets into Abusive/Non-Abusive and Threatening/Non-Threatening categories, respectively. Datasets of 2400/6000 training tweets and 1100/3950 testing tweets were created and manually annotated, along with logistic regression and BERT-based baselines. 21 teams participated and the best systems achieved F1-scores of 0.880 and 0.545 on the abusive and threatening language tasks, respectively, with m-BERT showing the best performance.
Keywords
Urdu · abusive language detection · threatening language detection · FIRE 2021 · m-BERT
Get the weekly digest
Top AI stories from the GCC region, every week.