Proper Noun Diacritization for Arabic Wikipedia: A Benchmark Dataset
arXiv · · Notable
Summary
A new dataset for Arabic proper noun diacritization was introduced, addressing the ambiguity caused by undiacritized proper nouns in Arabic Wikipedia. The dataset includes manually diacritized Arabic proper nouns of various origins along with their English Wikipedia glosses. GPT-4o was benchmarked on the task of recovering full diacritization from undiacritized Arabic and English forms, achieving 73% accuracy. Why it matters: The release of this dataset should facilitate further research on Arabic Wikipedia proper noun diacritization, improving the accessibility and accuracy of Arabic NLP resources.
Keywords
diacritization · Arabic · Wikipedia · proper nouns · GPT-4o
Get the weekly digest
Top AI stories from the GCC region, every week.