Cross-Document Topic-Aligned Chunking for Retrieval-Augmented Generation

arXiv · November 8, 2025 · Significant research

NLP LLM Research Arabic AI Information Retrieval

Summary

This paper introduces Cross-Document Topic-Aligned (CDTA) chunking to address knowledge fragmentation in Retrieval-Augmented Generation (RAG) systems. CDTA identifies topics across documents, maps segments to topics, and synthesizes them into unified chunks. Experiments on HotpotQA and UAE legal texts show that CDTA improves faithfulness and citation accuracy compared to existing chunking methods, especially for complex queries requiring multi-hop reasoning.

Keywords

chunking · RAG · topic modeling · knowledge retrieval · cross-document

Read original article →

Get the weekly digest

Top AI stories from the GCC region, every week.

Retrieval Augmentation as a Shortcut to the Training Data

MBZUAI · Invalid Date

This article discusses retrieval augmentation in text generation, where information retrieved from an external source is used to condition predictions. It references recent work on retrieval-augmented image captioning, showing that model size can be greatly reduced when training data is available through retrieval. The author intends to continue this work focusing on the intersection of retrieval augmentation and in-context learning, and controllable image captioning for language learning materials. Why it matters: This research direction has the potential to improve transfer learning in vision-language models, which could be especially relevant for downstream applications in Arabic NLP and multimodal tasks.

Cross-Document Topic-Aligned Chunking for Retrieval-Augmented Generation

Summary

Keywords

Related

Retrieval Augmentation as a Shortcut to the Training Data