Skip to content
GCC AI Research

Search

Results for "Arabic speech recognition"

QASR: QCRI Aljazeera Speech Resource -- A Large Scale Annotated Arabic Speech Corpus

arXiv ·

The Qatar Computing Research Institute (QCRI) has released QASR, a 2,000-hour transcribed Arabic speech corpus collected from Aljazeera news broadcasts. The dataset features multi-dialect speech sampled at 16kHz, aligned with lightly supervised transcriptions and linguistically motivated segmentation. QCRI also released a 130M word dataset to improve language model training. Why it matters: QASR enables new research in Arabic speech recognition, dialect identification, punctuation restoration, and other NLP tasks for spoken data.

Egypt's Intella Raises $12.5M to Expand Arabic AI Speech Models - Dabafinance

GCC AI Startup ·

Egyptian AI startup Intella, specializing in Arabic speech recognition, has raised $12.5 million in funding. The round was led by বিনিয়োগ, with participation from other investors. Intella plans to use the capital to expand its Arabic AI speech models and related services. Why it matters: The funding will help advance Arabic language AI capabilities, which are currently underserved compared to English-centric models.

N-Shot Benchmarking of Whisper on Diverse Arabic Speech Recognition

arXiv ·

This paper benchmarks the performance of OpenAI's Whisper model on diverse Arabic speech recognition tasks, using publicly available data and novel dialect evaluation sets. The study explores zero-shot, few-shot, and full finetuning scenarios. Results indicate that while Whisper outperforms XLS-R models in zero-shot settings on standard datasets, its performance drops significantly when applied to unseen Arabic dialects.