Overview of BIG-Bench Hard (BBH) tasks
mainBBH is a subset of 27 challenging tasks from the BIG-Bench suite focusing on algorithmic, commonsense, and multi-step reasoning. It includes 6,511 total samples.
Task Categories:
Multiple Choice Datasets:
date_understanding,disambiguation_qa,geometric_shapes,hyperbaton,logical_deduction_five_objects,logical_deduction_seven_objects,logical_deduction_three_objects,movie_recommendation,penguins_in_a_table,reasoning_about_colored_objects,ruin_names,salient_translation_error_detection,snarks,temporal_sequences,tracking_shuffled_objects_five_objects,tracking_shuffled_objects_seven_objects,tracking_shuffled_objects_three_objects
Binary Choice Datasets:
boolean_expressions(True, False)causal_judgement(Yes, No)formal_fallacies(valid, invalid)navigate(Yes, No)sports_understanding(yes, no)web_of_lies(Yes, No)
Open Answer Datasets:
multistep_arithmetic_two(integer)object_counting(natural number)word_sorting(list of words)dyck_languages(closing brackets)
Scoring: Scoring is based on simple accuracy calculated over the samples.