Fine-tuning Text-to-Image Models: Reinforcement Learning and Reward Over-Optimization
MBZUAI · Notable
Summary
The article discusses research on fine-tuning text-to-image diffusion models, including reward function training, online reinforcement learning (RL) fine-tuning, and addressing reward over-optimization. A Text-Image Alignment Assessment (TIA2) benchmark is introduced to study reward over-optimization. TextNorm, a method for confidence calibration in reward models, is presented to reduce over-optimization risks. Why it matters: Improving the alignment and fidelity of text-to-image models is crucial for generating high-quality content, and addressing over-optimization enhances the reliability of these models in creative applications.
Keywords
text-to-image · fine-tuning · reinforcement learning · reward over-optimization · diffusion models
Get the weekly digest
Top AI stories from the GCC region, every week.