Skip to content
GCC AI Research

HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification

arXiv · · Significant research

Summary

HalluTruthQA-4K is an expanded corpus comprising 4,000 expert-curated Arabic question-answering instances designed for fine-grained hallucination detection and truth verification in large language models. The resource includes model-generated responses, verified reference answers, distractors, character-level erroneous spans, human explanations, and hierarchical hallucination types across four knowledge domains. It features 1,643 hallucinated and 2,357 non-hallucinated responses, with 1,843 annotated erroneous spans, and serves as the official dataset for Track 2 of the HalluScoring 2026 shared task. Why it matters: This comprehensive, expert-annotated corpus provides a crucial foundation for advancing research and development in building more reliable and factually accurate Arabic large language models.

Get the weekly digest

Top AI stories from the GCC region, every week.

Related

AraHalluEval: A Fine-grained Hallucination Evaluation Framework for Arabic LLMs

arXiv ·

The paper introduces AraHalluEval, a new framework for evaluating hallucinations in Arabic and multilingual large language models (LLMs). The framework uses 12 fine-grained hallucination indicators across generative question answering and summarization tasks, evaluating 12 LLMs including Arabic-specific, multilingual, and reasoning-based models. Results show factual hallucinations are more common than faithfulness errors, with the Arabic model Allam showing lower hallucination rates. Why it matters: This work addresses a critical gap in Arabic NLP by providing a comprehensive tool for assessing and mitigating hallucination in LLMs, which is essential for reliable AI applications in the Arabic-speaking world.

GenAI Content Detection Task 1: English and Multilingual Machine-Generated Text Detection: AI vs. Human

arXiv ·

The GenAI Content Detection Task 1 is a shared task on detecting machine-generated text, featuring monolingual (English) and multilingual subtasks. The task, part of the GenAI workshop at COLING 2025, attracted 36 teams for the English subtask and 26 for the multilingual one. The organizers provide a detailed overview of the data, results, system rankings, and analysis of the submitted systems.