Evaluation of Adversarial Robustness in Arabic Language Models
arXiv · · Significant research
Summary
This study evaluated the adversarial robustness of five state-of-the-art Arabic Language Models against various Arabic adversarial attacks at character, word, and sentence levels. It found that diacritic insertion could reduce model accuracy by up to 92%, while manipulating Arabic conjunctions led to a 58% accuracy degradation, and paraphrasing reduced performance by an average of 76%. While adversarial training improved overall resilience, particularly for MARBERT and AraBERT, challenges against character-level noise persist. Why it matters: These findings are crucial for understanding and mitigating security vulnerabilities in Arabic AI, guiding the development of more robust and safe Arabic NLP systems.
Keywords
Adversarial Robustness · Arabic Language Models · Adversarial Attacks · NLP Security · Arabic NLP
Get the weekly digest
Top AI stories from the GCC region, every week.