This paper presents a comparative study of pre-trained transformer models for Arabic question answering (QA). The study evaluates the performance of AraBERTv2-base, AraBERTv0.2-large, and AraELECTRA models on four reading comprehension datasets: Arabic-SQuAD, ARCD, AQAD, and TyDiQA-GoldP. The researchers fine-tuned these models and analyzed the results to understand the performance disparities. Why it matters: This research contributes to the advancement of Arabic NLP by evaluating and comparing state-of-the-art models on important QA tasks, addressing the scarcity of resources in this domain.
Zeerak Talat, an independent scholar, gave a talk at MBZUAI on ethical concerns in NLP. The talk covered disparities in research on biases in NLP, performance differences based on socio-economic language variations, and risks of malicious reuse of NLP tools. Talat's research considers how machine learning interacts with and impacts societies through content moderation technologies. Why it matters: As NLP technologies become more integrated into society, understanding and addressing their potential harms and ethical implications is crucial for responsible development and deployment in the region and beyond.
MBZUAI Professor Timothy Baldwin reflects on his decision to move from the University of Melbourne to help build MBZUAI. He emphasizes the importance of inclusive AI education to ensure the benefits of AI are shared globally. Baldwin notes the concentration of AI innovation in a few countries, leading to disparities in language model performance for non-dominant languages. Why it matters: The article highlights MBZUAI's role in addressing the global imbalance in AI development and promoting inclusivity in AI education and research, particularly for Arabic and other underrepresented languages.