Researchers introduce HalluTruthQA, a new fine-grained benchmark designed for hallucination detection, localization, and explanation in Arabic question answering. This benchmark comprises 2,400 expert-curated examples across Islamic knowledge, history, science, and geography, featuring character-level error spans, human explanations, and various hallucination types. The study evaluated four open-source Arabic LLMs (ALLaM-7B, Falcon-H1R-7B, Qwen3-32B, SILMA) across detection, localization, factual verification, and explanation tasks, revealing no single model outperforms others across all metrics. Why it matters: HalluTruthQA provides a critical tool for advancing the factual accuracy and reliability of Arabic LLMs by enabling more granular and comprehensive hallucination evaluation beyond response-level detection.
Researchers have introduced HalluTruthQA, a new fine-grained benchmark designed for hallucination detection, localization, and explanation in Arabic Question Answering. The benchmark comprises 2,400 expert-curated examples spanning four knowledge-intensive domains: Islamic knowledge, history, science, and geography, with detailed annotations including character-level erroneous spans and human-written explanations. Four open-source LLMs ( extsc{Allam}, extsc{Falcon-H1}, extsc{Qwen32}, and extsc{Silma}) were evaluated, demonstrating varied performance across detection, localization, factual verification, and explanation tasks. Why it matters: This benchmark offers a comprehensive tool for evaluating and enhancing the factual accuracy and trustworthiness of Arabic LLMs, promoting more sophisticated assessment beyond simple hallucination detection.
Technology Innovation Institute (TII) will make its Falcon-H1 large language model available as an NVIDIA NIM microservice. Falcon-H1 features a hybrid Transformer–Mamba architecture supporting context windows of up to 256k tokens. The model's availability on NVIDIA NIM aims to provide enterprises with a plug-and-play asset for building AI systems. Why it matters: This integration will simplify deployment and scaling of Falcon-H1 for enterprises, potentially accelerating the adoption of sovereign AI solutions in the region.
TII in Abu Dhabi has launched Falcon Arabic, the first Arabic language model in the Falcon series, which is now the best-performing Arabic AI model in the region. They also released Falcon H1, a new model designed for performance and portability, outperforming Meta’s LLaMA and Alibaba’s Qwen in the small-to-medium size category. Falcon Arabic is built on Falcon 3-7B and trained on a high-quality native Arabic dataset. Why it matters: These releases strengthen the UAE's position as a leader in Arabic language AI and democratize access to high-performance AI models.
Abu Dhabi’s Technology Innovation Institute (TII) has launched Falcon-H1 Arabic, a new large language model based on a hybrid Mamba-Transformer architecture. The Falcon-H1 family comes in 3B, 7B, and 34B parameter sizes and outperforms existing models on the Open Arabic LLM Leaderboard (OALL). The model features improvements in data quality, dialect coverage, and long-context stability. Why it matters: This release strengthens the UAE's position in Arabic AI and provides a high-performing model tailored to the linguistic and cultural needs of the region.