Overview of Theory of Mind and Social Intelligence Evals
mainThis evaluation suite tests Large Language Models (LLMs) on social intelligence and theory of mind using two primary benchmarks:
- ToMi (Theory of Mind): Based on the Sally-Anne test, this assesses a model's ability to infer false beliefs in others. It contains 5,993 question-answer pairs.
- SocialIQA: A multiple-choice benchmark (3 options per question) covering social scenarios, motivations, and reactions. It contains 2,224 question-answer pairs.
Light Versions: For rapid iteration on prompts and scaffolding, 'light' versions of both datasets are available, containing 1/10th of the original data points.