From Text to image: M.Sc. graduate develops cutting-edge techniques to transform T2I generation
MBZUAI · Significant research
Summary
MBZUAI M.Sc. graduate Mohammad Hanan Ghani developed new techniques to improve text-to-image generation from long text prompts, combining large language models and diffusion models. Advised by Dr. Salman Khan, Ghani published three papers at ICLR, BMVC, and NeurIPS, with the ICLR paper focusing on generating images that accurately reflect detailed text descriptions. The new system improves upon existing techniques to generate images that closely follow the details of the input text. Why it matters: This research addresses a key limitation in current T2I models and advances the field of multimodal AI, potentially improving the capabilities of robots and autonomous devices.
Keywords
text-to-image · MBZUAI · diffusion models · large language models · multimodal AI
Get the weekly digest
Top AI stories from the GCC region, every week.