Skip to content
GCC AI Research

Search

Results for "EMNLP"

Adapting AI to identify Arabic dialects

KAUST ·

KAUST researchers have developed a parameter-efficient learning approach to identify Arabic dialects using limited data and computing power, fine-tuning the Whisper model with a dataset of 17 dialects. The model achieves high accuracy using only 2.5% of the parameters of the larger model and 30% of the training data. Srijith Radhakrishnan presented the findings at EMNLP 2023 and Interspeech 2023. Why it matters: This research addresses the challenge of dialect identification in Arabic NLP and enables more efficient use of large language models in resource-constrained environments.

MBZUAI celebrates another year at the forefront of transformative AI

MBZUAI ·

MBZUAI has published 674 papers in 2023 and holds a global ranking of 18 in AI, CV, ML, NLP, and robotics according to CSRankings. The university presented 30 papers at ICCV, will present 53 papers at NeurIPS, and has 44 papers at EMNLP 2023. MBZUAI was also awarded its first patent by the US Patent Office for a system and method for handwriting generation. Why it matters: This demonstrates the rapid growth and increasing prominence of MBZUAI as a leading AI research institution in the region and globally.

MBZUAI researchers earn high-profile honors at EMNLP

MBZUAI ·

MBZUAI researchers received high honors at EMNLP 2025 for two research papers, placing them in the top 2% of accepted work. One paper, MAviS, is a multimodal AI system that identifies bird species by combining images, sounds, and text. The other award-winning paper focuses on uncertainty in LLM-as-a-Judge. Why it matters: The recognition highlights MBZUAI's growing influence in NLP and multimodal AI research, particularly in domain-specific applications like biodiversity conservation.

Fine-grained species recognition with MAviS: a new dataset, benchmark, and model

MBZUAI ·

MBZUAI researchers have developed MAviS, a new multimodal dataset, benchmark, and chatbot for fine-grained bird species recognition. MAviS includes images, audio, and text to help models identify subtle differences between species, especially rare and regional varieties. The related study was presented at EMNLP 2025 and selected as a "Senior Area Chair Highlight". Why it matters: This work addresses a key limitation in AI's ability to support biodiversity conservation and ecological monitoring in the region and globally.

Can we tell when AI wrote that code? This project thinks so, even when the AI tries to hide it

MBZUAI ·

MBZUAI researchers introduced Droid, a resource suite and detector family, at EMNLP 2025 designed to distinguish between AI-generated and human-written code. The project addresses the challenge of identifying AI-generated code in software development, considering the prevalence of AI-suggested code and the risks of obfuscated backdoors and feedback loops. DroidCollection includes over one million code samples across seven programming languages, three coding domains, and outputs from 43 different code models, including human-AI co-authored code and adversarially humanized machine code. Why it matters: This research is crucial for maintaining software security and integrity in the age of AI-assisted coding, providing a robust tool for detecting AI-generated code across diverse languages and domains.

Study on the paradox of ‘low-resource’ languages wins Outstanding Paper Award at EMNLP

MBZUAI ·

A study co-authored by researchers from UC Berkeley, University of the Witwatersrand, Lelapa AI, and MBZUAI received the Outstanding Paper Award at EMNLP 2024. The paper critiques the term "low-resource" languages in NLP, highlighting its limitations in capturing the diverse challenges faced by different languages. The authors propose a more detailed analysis of resourcedness to encourage targeted support for languages currently underserved by technology. Why it matters: The research challenges assumptions in NLP and promotes more nuanced approaches to supporting the world's many languages, including Arabic, in AI systems.

Dozens of studies by MBZUAI scientists presented at top natural language processing conference

MBZUAI ·

MBZUAI faculty and students will present 44 papers at the Empirical Methods in Natural Language Processing (EMNLP) conference in Singapore. Research topics include disinformation detection, social media analysis, dialogue generation, and Arabic LLMs. Preslav Nakov, Iryna Gurevych, Timothy Baldwin, Alham Fikri Aji, and Muhammad Abdul-Mageed are among the MBZUAI researchers presenting at the conference. Why it matters: MBZUAI's strong presence at a top NLP conference highlights the UAE's growing contributions to cutting-edge AI research and its increasing global prominence in the field.

Making LLM accuracy a matter of fact

MBZUAI ·

MBZUAI NLP master's graduate Hasan Iqbal developed OpenFactCheck, a framework for fact-checking and evaluating the factual accuracy of large language models. The framework consists of three modules: ResponseEvaluator, LLMEvaluator, and CheckerEvaluator. OpenFactCheck was published at EMNLP 2024 and accepted at NAACL 2025 and COLING 2025, with Iqbal playing an active role at COLING in Abu Dhabi. Why it matters: The development of automated fact-checking frameworks is crucial for ensuring the reliability and trustworthiness of information generated by increasingly prevalent LLMs, especially in the Arabic-speaking world.