LegalBench

repository·main·Indexed 20 days ago

https://github.com/hazyresearch/legalbench

An open-science benchmark for evaluating legal reasoning in English Large Language Models (LLMs). It consists of over 160 tasks curated by legal professionals, where each task is paired with a dataset of input-output pairs to measure reasoning that is either 'interesting' or 'useful' for real-world legal workflows.

Tokens
97.6K
Snippets
200
Records
592
Agent score
69%

What's inside LegalBench

  1. Overview of the learned_hands_divorce task

    main

    The learned_hands_divorce task is a binary classification task designed for issue-spotting in legal contexts. The goal is to determine if a user's post discusses legal issues related to divorce, such as filing for divorce, separation, annulment, spousal support, splitting money/property, or court processes.

    Task Details

  2. Overview of the cuad_license_grant task

    main

    The cuad_license_grant task is a binary classification task designed to determine if a specific contractual clause contains a license granted by one party to its counterparty. It is based on the CUAD dataset and focuses on the legal reasoning type of 'Interpretation'.

    Task Details

    • Task Type: Binary classification
    • Legal Reasoning Type: Interpretation
    • Sample Size: 1402
    • Source: Atticus Project
    • License: CC By 4.0
  3. Overview of the cuad_cap_on_liability task

    main

    The cuad_cap_on_liability task is a binary classification task designed to determine if a contractual clause specifies a cap on liability. This includes identifying clauses that set a maximum amount for recovery or a time limitation for a counterparty to bring claims following a breach of obligation.

    Task Details

    • Legal Reasoning Type: Interpretation
    • Task Type: Binary classification
    • Dataset Size: 1252 samples
    • Source: Atticus Project (CUAD)
    • License: CC By 4.0
  4. Overview of the learned_hands_traffic task

    main

    The learned_hands_traffic task is a binary classification task designed for issue-spotting in legal contexts. The goal is to determine if a user's post implicates legal issues related to the traffic system (e.g., traffic/parking tickets, fees, driver's licenses) or automotive matters (e.g., car accidents, injuries, vehicle quality, repairs, or purchases).

    Task Details:

    • Legal Reasoning Type: Issue-spotting
    • Task Type: Binary classification
    • Dataset Size: 562 samples
    • Source: Learned Hands
    • License: CC BY-NC-SA 4.0
  5. Overview of the legal_reasoning_causality task

    main

    The legal_reasoning_causality task is a binary classification task designed to identify whether an excerpt from a US Federal District Court opinion relies on statistical evidence (such as regression analysis) or direct evidence (such as witness testimony) to establish a causal link in labor discrimination cases.

    Task Details

    • Legal reasoning type: Rhetorical-analysis
    • Task type: Binary classification
    • Size: 59 samples
    • License: CC BY 4.0
  6. Overview of the cuad_non-transferable_license task

    main

    The cuad_non-transferable_license task is a binary classification task designed to determine if a contractual clause limits a party's ability to transfer a granted license to a third party. It is based on the CUAD dataset and focuses on legal interpretation.

    Task Details:

    • Task Type: Binary classification
    • Legal Reasoning Type: Interpretation
    • Sample Size: 548
    • Source: Atticus Project
    • License: CC By 4.0
  7. Overview of the CUAD IP Ownership Assignment task

    main

    The cuad_ip_ownership_assignment task is a binary classification task designed to determine if a contractual clause specifies that intellectual property (IP) created by one party becomes the property of the counterparty. This transfer of ownership can occur per the terms of the contract or upon the occurrence of specific events.

    Task Details

    • Task Type: Binary classification
    • Legal Reasoning Type: Interpretation
    • Dataset Size: 582 samples
    • Source: Atticus Project (CUAD)
    • License: CC By 4.0
  8. Overview of the SCALR task

    main

    SCALR is a 5-way classification benchmark designed to assess the legal reasoning and reading comprehension of large language models. The task requires identifying the correct 'holding' statement from a set of five candidates that best answers a given legal question (the 'question presented').

    Key characteristics:

    • Goal: Measure understanding of legal language rather than memorization of specific legal knowledge.
    • Source: Supreme Court of the United States cases (2001 Term and later).
    • Reasoning Type: Rhetorical-analysis.
    • Task Type: 5-way classification.
  9. Overview of the cuad_audit_rights task

    main

    The cuad_audit_rights task is a binary classification task designed to determine if a specific contractual clause grants a party the right to audit the books, records, or physical locations of a counterparty to ensure contract compliance.

    Key metadata:

    • Legal reasoning type: Interpretation
    • Task type: Binary classification
    • Dataset size: 1222 samples
    • Source: Atticus Project (CUAD)
    • License: CC By 4.0
  10. Overview of the learned_hands_education task

    main

    The learned_hands_education task is a binary classification task designed to determine if a user's post discusses legal issues related to education. This includes topics such as school accommodations for special needs, discrimination, student debt, and discipline.

    Task Details

  11. Overview of the telemarketing_sales_rule task

    main

    The telemarketing_sales_rule task evaluates a model's ability to apply specific Federal Trade Commission regulations (16 C.F.R. § 310.3(a)(1) and 16 C.F.R. § 310.3(a)(2)) to fact patterns. The goal is to determine if a described telemarketing practice constitutes a deceptive violation of the Telemarketing Sales Rule.

    Task Details:

    • Legal Reasoning Type: Application/conclusion
    • Task Type: Binary classification
    • Dataset Size: 52 samples
    • License: CC BY 4.0
  12. Overview of the definition_extraction task

    main

    The definition_extraction task involves identifying the specific term being defined within a sentence extracted from a Supreme Court opinion. This is a rhetorical-analysis task of the 'Extraction' type.

    Key Challenge: Defining sentences may not always use quotation marks to denote the term being defined (e.g., "A vacation is defined by..." defines 'vacation' without quotes).

    Task Metadata:

    • Source: Kevin Tobia
    • License: CC BY-SA 4.0
    • Size: 696 samples
    • Legal Reasoning Type: Rhetorical-analysis