Skip to content
GCC AI Research

MIRAGE: Exploring How Large Language Models Perform in Complex Social Interactive Environments

arXiv · · Notable

Summary

The paper introduces MIRAGE, a framework for evaluating LLMs' ability to simulate human behaviors in murder mystery games. MIRAGE uses four methods: TII, CIC, ICI and SCI to assess the LLMs' role-playing proficiency. Experiments show that even GPT-4 struggles with the complexities of the MIRAGE framework.

Keywords

LLM · role-playing · evaluation · MIRAGE · GPT-4

Get the weekly digest

Top AI stories from the GCC region, every week.