Overview of ArxivRollBench-Full benchmark
mainArxivRollBench is a rolling benchmark designed to evaluate the ability of large language models (LLMs) to reason over recent scientific text from arXiv papers. It uses three task formats: sequencing, cloze, and next-fragment prediction (under the SCP framework).
Key Details:
- Task Type: Multiple-choice scientific text reasoning.
- Input: Recent arXiv text fragments with four answer choices (A, B, C, or D).
- Output: A single correct answer letter.
- Domains: Computer Science, Quantitative Finance, Mathematics, Physics, Statistics, Quantitative Biology, Economics, and Electrical Engineering/System Science.
- Snapshots: Includes 2024b, 2025a, and 2026a rolling snapshots to reduce contamination-related overestimation.
- Metrics: Evaluated using accuracy (
acc). - Dataset ID:
liangzid/arxivrollbench-full(hosted on ModelScope).