Latent Space Exploration for Safe and Trustworthy AI Models
MBZUAI · Notable
Summary
Hassan Sajjad from Dalhousie University presented research on exploring the latent space of AI models to assess their safety and trustworthiness. He discussed use cases where analyzing latent space helps understand the robustness-generalization tradeoff in adversarial training and evaluate language comprehension. Sajjad's work aims to build better AI models and increase trust in their capabilities by looking at model internals. Why it matters: Intrinsic evaluation of model internals will become important to improving AI safety and robustness.
Keywords
latent space · AI models · trustworthiness · robustness · generalization
Get the weekly digest
Top AI stories from the GCC region, every week.