Overview of the GAIA dataset
mainGAIA (General AI Assistants Benchmark) is a benchmark designed to evaluate next-generation Large Language Models (LLMs) that possess augmented capabilities, such as access to tools, efficient prompting, and web search.
The dataset consists of over 450 non-trivial questions with unambiguous answers, categorized into three difficulty levels:
- Level 1: Achievable by high-performing LLMs.
- Level 3: Represents a significant jump in required model capabilities.
Each level includes a fully public dev set for validation and a test set containing private answers and metadata to prevent data leakage.