Overview of Spider 2.0 Benchmark Settings
mainSpider 2.0 provides three distinct task settings for evaluating Large Language Models (LLMs) on enterprise-level text-to-SQL workflows. Choose a setting based on your task type, database requirements, and budget:
| Setting | Task Type | # Examples | Databases | Cost |
|---|---|---|---|---|
| Spider 2.0-Snow | Text-to-SQL task | 547 | Snowflake (547) | NO COST! |
| Spider 2.0-Lite | Text-to-SQL task | 547 | BigQuery (214), Snowflake (198), SQLite (135) | Some cost incurred |
| Spider 2.0-DBT | Code agent task | 68 | DuckDB (DBT) (68) | NO COST! |
Note that for Spider 2.0-Lite and Spider 2.0-Snow, the project provides ground-truth tables to facilitate quick benchmarking. When using these, you must indicate that you are using oracle tables.