QASR: QCRI Aljazeera Speech Resource -- A Large Scale Annotated Arabic Speech Corpus
arXiv · · Significant research
Summary
The Qatar Computing Research Institute (QCRI) has released QASR, a 2,000-hour transcribed Arabic speech corpus collected from Aljazeera news broadcasts. The dataset features multi-dialect speech sampled at 16kHz, aligned with lightly supervised transcriptions and linguistically motivated segmentation. QCRI also released a 130M word dataset to improve language model training. Why it matters: QASR enables new research in Arabic speech recognition, dialect identification, punctuation restoration, and other NLP tasks for spoken data.
Keywords
Arabic speech recognition · speech corpus · Aljazeera · QASR · dialect identification
Get the weekly digest
Top AI stories from the GCC region, every week.