Technology Innovation Institute (TII) won the UAE AI Award for Emirati AI Solutions for its Falcon LLM series. AI71 also won for LAW71, an AI-powered legal solution, and RAZI71, an AI-powered healthcare solution. The award recognizes AI innovations made in the UAE that demonstrate innovation, AI ethics compliance, maturity, and scalability. Why it matters: The award highlights the UAE's commitment to developing local AI talent and solutions, particularly in open-source models, for global collaboration and positive transformation.
The Technology Innovation Institute (TII) in Abu Dhabi has launched the Falcon Foundation, a non-profit dedicated to advancing open-source generative AI models. TII is committing $300 million to fund open-source AI projects, beginning with its Falcon AI models. The foundation aims to foster collaboration among stakeholders, developers, academia, and industry to promote transparent governance and knowledge exchange in AI. Why it matters: This initiative signals the UAE's commitment to leading in AI development through open-source innovation and collaboration, potentially accelerating AI adoption and customization across various sectors.
Abu Dhabi's Advanced Technology Research Council (ATRC) has launched AI71, a new AI company building on the Falcon generative AI models developed by TII. AI71 will focus on multi-domain specializations, offering AI data control options for companies and countries looking to self-host for greater privacy. The company will be taken to market by ATRC's VentureOne subsidiary, initially targeting the medical, educational, and legal sectors. Why it matters: AI71 aims to establish Abu Dhabi and the UAE as a major AI player by providing decentralized data ownership and promoting broader access to AI technology.
Technology Innovation Institute (TII) in the UAE has launched Falcon 180B, an open access large language model with 180 billion parameters trained on 3.5 trillion tokens. Falcon 180B ranks first on the Hugging Face Leaderboard for pretrained LLMs, outperforming Meta's LLaMA 2 and nearing the performance of OpenAI's GPT-4 and Google's PaLM 2. The model is available for research and commercial use under the 'Falcon 180B TII License', based upon Apache 2.0. Why it matters: This release strengthens the UAE's position in AI development and promotes open access to advanced AI technology, fostering innovation and collaboration.
TII's Falcon 40B, a 40-billion-parameter open-source AI model, has ranked #1 on Hugging Face's Open LLM Leaderboard, surpassing models like LLaMA and StableLM. The leaderboard uses benchmarks like AI2 Reasoning Challenge, HellaSwag, MMLU, and TruthfulQA. Trained on one trillion tokens, Falcon 40B's weights are available for research and commercial use. Why it matters: This achievement positions the UAE as a leader in generative AI and promotes transparent, inclusive AI development.
Researchers introduce HalluTruthQA, a new fine-grained benchmark designed for hallucination detection, localization, and explanation in Arabic question answering. This benchmark comprises 2,400 expert-curated examples across Islamic knowledge, history, science, and geography, featuring character-level error spans, human explanations, and various hallucination types. The study evaluated four open-source Arabic LLMs (ALLaM-7B, Falcon-H1R-7B, Qwen3-32B, SILMA) across detection, localization, factual verification, and explanation tasks, revealing no single model outperforms others across all metrics. Why it matters: HalluTruthQA provides a critical tool for advancing the factual accuracy and reliability of Arabic LLMs by enabling more granular and comprehensive hallucination evaluation beyond response-level detection.
Researchers have introduced HalluTruthQA, a new fine-grained benchmark designed for hallucination detection, localization, and explanation in Arabic Question Answering. The benchmark comprises 2,400 expert-curated examples spanning four knowledge-intensive domains: Islamic knowledge, history, science, and geography, with detailed annotations including character-level erroneous spans and human-written explanations. Four open-source LLMs ( extsc{Allam}, extsc{Falcon-H1}, extsc{Qwen32}, and extsc{Silma}) were evaluated, demonstrating varied performance across detection, localization, factual verification, and explanation tasks. Why it matters: This benchmark offers a comprehensive tool for evaluating and enhancing the factual accuracy and trustworthiness of Arabic LLMs, promoting more sophisticated assessment beyond simple hallucination detection.
This study introduces a Probabilistic Graphical Model (PGM) framework utilizing Pearl's do-operator to causally audit LLM safety mechanisms, specifically isolating the effect of injecting cultural demographics into prompts. A large-scale empirical analysis was conducted across seven instruction-tuned models from diverse origins, including the UAE's Falcon3-7B, as well as models from the US, Europe, China, and India, using ToxiGen and BOLD datasets. The findings revealed a disparity between observational and interventional bias, demonstrating that standard fairness metrics can overestimate demographic bias. Western models exhibited higher causal refusal rates for specific demographic groups, while Eastern models showed low overall intervention rates with targeted sensitivities toward regional demographics. Why it matters: This research highlights the geopolitical nuances of LLM safety alignment and the potential for demographic-sensitive over-triggering to restrict benign discourse, which is particularly relevant for diverse regions like the Middle East in developing culturally-aware AI.