Researchers have proposed the Cylindrical Representation Hypothesis (CRH) to address the instability and unpredictability observed in steering large language models, an issue not fully explained by the existing Linear Representation Hypothesis (LRH). CRH suggests that overlapping concept contributions lead to a sample-specific axis-orthogonal structure, comprising a central axis for concept generation and a surrounding normal plane for steering sensitivity. This framework identifies intrinsic uncertainty at the 'sensitive sector' level within the plane, providing a principled explanation for fluctuations in steering outcomes. Experiments verify the existence of this cylindrical structure and demonstrate CRH's practical utility in interpreting real-world model steering behavior, with code available on GitHub from mbzuai-nlp. Why it matters: This research from MBZUAI offers a crucial theoretical advancement in understanding and potentially improving the control and reliability of large language models.
Researchers from MBZUAI have proposed the Cylindrical Representation Hypothesis (CRH) to explain the instability and unpredictability observed in large language model steering. CRH relaxes the orthogonality assumption of the existing Linear Representation Hypothesis, positing a cylindrical structure where a central axis captures concept differences and a surrounding normal plane controls steering sensitivity. The hypothesis suggests that the intrinsic uncertainty in identifying specific sensitive sectors within this normal plane accounts for why steering outcomes frequently fluctuate even with well-aligned directions. Why it matters: This research offers a more robust theoretical framework for understanding and potentially improving the control and reliability of large language models.
Professor Omar Knio, Dean of the Computer, Electrical and Mathematical Sciences and Engineering (CEMSE) Division at KAUST, has been named a 2026 SIAM Fellow. This prestigious recognition from the Society for Industrial and Applied Mathematics honors his leadership in uncertainty quantification and multiscale mathematics. His research areas include applications in combustion, energetic materials, geophysical fluid dynamics, high-performance computing, and data-enabled predictive science. Why it matters: This recognition highlights KAUST's global standing in applied mathematics and computational science, reinforcing its role as a hub for scientific talent and interdisciplinary research crucial for advanced technological development in Saudi Arabia.
MBZUAI researchers introduce SocialMaze, a new benchmark for evaluating social reasoning capabilities in large language models (LLMs). SocialMaze includes six diverse tasks across social reasoning games, daily-life interactions, and digital community platforms, emphasizing deep reasoning, dynamic interaction, and information uncertainty. Experiments show that LLMs vary in handling dynamic interactions, degrade under uncertainty, but can be improved via fine-tuning on curated reasoning examples.
This paper introduces Diffusion-BBO, a new online black-box optimization (BBO) framework that uses a conditional diffusion model as an inverse surrogate model. The framework employs an Uncertainty-aware Exploration (UaE) acquisition function to propose scores in the objective space for conditional sampling. The approach is shown theoretically to achieve a near-optimal solution and empirically outperforms existing online BBO baselines across 6 scientific discovery tasks.
MBZUAI researchers received high honors at EMNLP 2025 for two research papers, placing them in the top 2% of accepted work. One paper, MAviS, is a multimodal AI system that identifies bird species by combining images, sounds, and text. The other award-winning paper focuses on uncertainty in LLM-as-a-Judge. Why it matters: The recognition highlights MBZUAI's growing influence in NLP and multimodal AI research, particularly in domain-specific applications like biodiversity conservation.
Researchers from MBZUAI developed "uncertainty quantification heads" (UQ heads) to detect hallucinations in language models by probing internal states and estimating the credibility of generated text. UQ heads leverage attention maps and logits to identify potential hallucinations without altering the model's generation process or relying on external knowledge. The team found that UQ heads achieved state-of-the-art performance in claim-level hallucination detection across different domains and languages. Why it matters: This approach offers a more efficient and accurate method for identifying hallucinations, improving the reliability and trustworthiness of language models in various applications.
Omar Knio, Dean of the Computer, Electrical and Mathematical Sciences and Engineering (CEMSE) Division at KAUST, has been named a 2026 SIAM Fellow. He received this honor from the Society for Industrial and Applied Mathematics (SIAM) for his leadership in uncertainty quantification and multiscale mathematics. His research applies to fields such as combustion, energetic materials, and geophysical fluid dynamics. Why it matters: This recognition highlights KAUST's and Saudi Arabia's growing contributions to global scientific talent and interdisciplinary research in applied mathematics and computational science.