MultiProSE: A Multi-label Arabic Dataset for Propaganda, Sentiment, and Emotion Detection
arXiv · · Significant research
Summary
The paper introduces MultiProSE, the first multi-label Arabic dataset for propaganda, sentiment, and emotion detection. It extends the existing ArPro dataset with sentiment and emotion annotations, resulting in 8,000 annotated news articles. Baseline models, including GPT-4o-mini and BERT-based models, were developed for each task, and the dataset, guidelines, and code are publicly available. Why it matters: This resource enables further research into Arabic language models and a better understanding of opinion dynamics within Arabic news media.
Keywords
Arabic NLP · propaganda detection · sentiment analysis · emotion detection · MultiProSE
Get the weekly digest
Top AI stories from the GCC region, every week.