Skip to content
GCC AI Research

Search

Results for "Masader"

Masader Plus: A New Interface for Exploring +500 Arabic NLP Datasets

arXiv ·

Researchers have developed Masader Plus, a web interface for browsing the Masader catalog of Arabic NLP datasets. The interface allows for data exploration, filtration, and API access to examine datasets. User interactions with the website are intended to provide a way to improve the dataset catalog itself. Why it matters: This interface lowers the barrier to entry for researchers seeking Arabic NLP datasets, facilitating more research in the field.

Masader: Metadata Sourcing for Arabic Text and Speech Data Resources

arXiv ·

Researchers created Masader, the largest public catalog for Arabic NLP datasets, containing 200 datasets annotated with 25 attributes. They developed a metadata annotation strategy applicable to other languages. The paper highlights issues within current Arabic NLP datasets and suggests recommendations. Why it matters: This curated dataset directory helps lower the barrier to entry for Arabic NLP research and development.