An Empirical Study of Pre-trained Transformers for Arabic Information Extraction
arXiv · · Significant research
Summary
This paper introduces GigaBERT, a customized bilingual BERT model pre-trained for Arabic NLP and English-to-Arabic zero-shot transfer learning. The study evaluates GigaBERT's performance on four information extraction tasks: named entity recognition, part-of-speech tagging, argument role labeling, and relation extraction. Results show that GigaBERT outperforms mBERT, XLM-RoBERTa, and AraBERT in both supervised and zero-shot transfer settings. Why it matters: GigaBERT advances Arabic NLP by providing a high-performing, publicly available model tailored for the complexities of the Arabic language and cross-lingual applications.
Keywords
GigaBERT · Arabic NLP · zero-shot learning · information extraction · transfer learning
Get the weekly digest
Top AI stories from the GCC region, every week.