Skip to content
GCC AI Research

Search

Results for "post-training"

LLM Post-Training: A Deep Dive into Reasoning Large Language Models

arXiv ·

A new survey paper provides a deep dive into post-training methodologies for Large Language Models (LLMs), analyzing their role in refining LLMs beyond pretraining. It addresses key challenges such as catastrophic forgetting, reward hacking, and inference-time trade-offs, and highlights emerging directions in model alignment, scalable adaptation, and inference-time reasoning. The paper also provides a public repository to continually track developments in this fast-evolving field.

MBZUAI launches K2 Think V2: UAE’s fully sovereign, next-generation reasoning system

MBZUAI ·

MBZUAI, G42, and Cerebras Systems have launched K2 Think V2, a 70-billion parameter reasoning system built on the K2-V2 base model. K2 Think V2 is fully open-source, from pre-training data to post-training alignment, ensuring transparency and reproducibility. It achieves leading results on complex reasoning benchmarks like AIME2025 and GPQA-Diamond. Why it matters: This release marks a significant advancement in the UAE's AI capabilities, demonstrating leadership in building globally accessible and fully sovereign AI systems focused on reasoning.

K2 Think V2: a fully sovereign reasoning model

MBZUAI ·

MBZUAI's Institute of Foundation Models (IFM) has released K2 Think V2, a 70 billion parameter open-source general reasoning model built on K2 V2 Instruct. The model excels in complex reasoning benchmarks like AIME2025 and GPQA-Diamond, and features a low hallucination rate with long context reasoning capabilities. K2 Think V2 is fully sovereign and open, from pre-training through post-training, using IFM-curated data and a Guru dataset. Why it matters: This release contributes to closing the gap between community-owned reproducible AI and proprietary models, particularly in reasoning and long-context understanding for Arabic NLP tasks.