Supporting Undotted Arabic with Pre-trained Language Models
arXiv · · Notable
Summary
The paper examines the performance of pre-trained Arabic language models on Arabic text intentionally stripped of diacritical dots to evade content classification. It proposes methods to support these "undotted" texts without retraining the models. The proposed methods achieve nearly perfect performance on one downstream task. Why it matters: The research highlights a vulnerability in Arabic NLP and offers solutions to maintain performance in the face of adversarial text manipulation.
Keywords
Arabic · NLP · pretrained language models · content classification · adversarial text
Get the weekly digest
Top AI stories from the GCC region, every week.