mirdata Documentation

repository·master·Indexed 19 days ago

https://github.com/mir-dataset-loaders/mirdata

A Python library providing standardized loaders for Music Information Retrieval (MIR) datasets. mirdata automates dataset downloading, validation, and annotation loading into formats compatible with evaluation tools like mir_eval to facilitate reproducible research.

Tokens
19.1K
Snippets
49
Records
86
Agent score
62%

What's inside mirdata

  1. Overview of supported datasets in mirdata

    master

    The mirdata repository provides loaders for a wide variety of Music Information Retrieval (MIR) datasets. The following table summarizes the capabilities of currently supported datasets, including whether audio/annotations are downloadable, the types of annotations available (e.g., beats, tempo, pitch, chords), the number of tracks, and the licensing terms.

    Dataset Summary Table

    DatasetDownloadable?Annotation TypesTracksLicense
    AcousticBrainz Genreaudio: 🧮, annotations: ✅genre>4MCC BY-SA 4.0
    BAFaudio: 🔑, annotations: 🔑matches3425Custom
    Ballroomaudio: ✅, annotations: ✅beats, tempo, genre698CC0 1.0
    Beatlesaudio: ❌, annotations: ✅beats, chords, sections, key, vocal-activity180
    Beatport EDM keyaudio: ✅, annotations: ✅global key1486CC BY-SA 4.0
    Billboard (McGill)audio: ❌, annotations: ✅chords, sections890CC0 1.0
    BRIDaudio: ✅, annotations: ✅beats, tempo367CC BY-NC-SA 4.0
    Candombeaudio: ✅, annotations: ✅beats35CC BY-NC-SA 4.0
    cante100audio: 🔑, annotations: ✅f0, Vocal notes100cante
    CIPImusicXML: 🔑, embeddings: 🔑, annotations: 🔑difficulty levels652CC BY-NC-SA 4.0
    (CompMusic) Carnatic Rhythmaudio: 🔑, annotations: 🔑beats, meter176CC BY-NC-SA 4.0
    (CompMusic) Hindustani Rhythmaudio: 🔑, annotations: 🔑beats, meter151CC BY-NC-SA 4.0
    (CompMusic) Indian Tonicaudio: 🔑, annotations: ✅tonic2150CC BY-NC-SA 4.0
    (CompMusic) Jingju A Cappellaaudio: ✅, annotations: ✅lyrics, phonemes, syllables82CC BY-NC-SA 4.0
    (CompMusic) OTMM Makamaudio: ✅, annotations: ✅f0, tonic1000CC BY-NC-SA 4.0
    (CompMusic) Ragaaudio: 🔑, annotations: ✅f0, segments, tonic780CC BY-NC-SA 4.0
    Cuidadoaudio: ❌, annotations: ✅beats, tempo698CC0 1.0
    Dagstuhl ChoirSetmultitrack audio: ✅, annotations: ✅f0, beats, notes108CC BY 4.0
    DALIaudio: 📺, annotations: ✅lyrics, Vocal notes5358CC BY-SA 4.0
    Da-TACOSaudio: 🧮, annotations: ✅lyrics, Vocal notes15k/10kCC BY-SA 4.0
    EGFxSetaudio: ✅, annotations: ✅notes8970CC BY-SA 4.0
    Filosaxaudio: 🔑, annotations: 🔑, midi: 🔑f0, beats, chords, tempo, notes48
    FMA Keysaudio: ✅, annotations: ✅spotify_uri, key, mode, key_number, mode_number5489CC BY-SA 4.0
    Four-Way Tabla Strokeaudio: ✅, annotations: ✅tags236CC BY-SA 4.0
    Freesound One-Shotaudio: ✅, annotations: ✅tags10254CC BY-SA 4.0
    Giantsteps keyaudio: ✅, annotations: ✅global key500CC BY-SA 4.0
    Giantsteps tempoaudio: 📺, annotations: ✅global genre, global tempo664CC BY-SA 4.0
    Good Soundsaudio: ✅, annotations: ✅instruments, sound quality, instrument metadata16308CC BY-SA 4.0
    Groove MIDIaudio: ✅, midi: ✅beats, tempo, drums1150CC BY-SA 4.0
    Gtzan-Genreaudio: ❌, annotations: ✅global genre, beats, tempo1000
    Guitarsetaudio: ✅, midi: ✅beats, chords, key, tempo, notes, f0360MIT
    (CompMusic) IAMMSaudio: ✅, annotations: ✅f0, section, event32CC BY-NC-SA 4.0
    Ikalaaudio: ❌, annotations: ❌Vocal f0, lyrics252ikala
    Hainsworthaudio: ❌, annotations: ❌beats, tempo222CC0 1.0
    Haydn op20audio: N/A, midi: ✅, scores: ✅, annotations: ✅symbolic chords, symbolic key24CC BY-NC-SA 4.0
    IDMT-SMT-Audio Effectsaudio: ✅, annotations: ✅instruments, notes, fx55044CC BY-NC-ND 4.0
    IRMASaudio: ✅, annotations: ✅instruments, genre9579CC BY-NC-SA 3.0
    Jazz Trio Databaseaudio: 🔑, annotations: ✅, midi: ✅beats, Global tempo, Piano notes1294MIT
    MAESTROaudio: ✅, annotations: ✅Piano notes1282CC BY-NC-SA 4.0
    MDB-stem-synthaudio: ✅, annotations: ✅f0230CC BY-NC 4.0
    Medley-solos-DBaudio: ✅, annotations: ✅instruments21571CC BY-SA 4.0
    MedleyDB melodyaudio: 🔑, annotations: ✅Melody f0108CC BY-NC-SA 4.0
    MedleyDB pitchaudio: 🔑, annotations: ✅f0, instruments103CC BY-NC-SA 4.0
    Mridangam Strokeaudio: ✅, annotations: ✅stroke-name, tonic6977CC BY 3.0
    MTG Jamendoaudio: ✅, annotations: ✅moodtheme annotations18448CC BY-NC-SA 4.0
    MULTIVOXaudio: ✅, video: ✅, metadata: ✅spatial, 360, near-field, demographics, etc.154CC BY 4.0
    Orchsetaudio: ✅, annotations: ✅Melody f064CC BY-NC-SA 4.0
    PHENICX-Anechoicmultitrack audio: ✅, annotations: ✅Aligned/Original score notes4CC BY-NC-SA 4.0
    Queenaudio: ❌, annotations: ✅chords, sections, key51
    RWC classicalaudio: ❌, annotations: ✅beats, sections61rwc
    RWC jazzaudio: ❌, annotations: ✅beats, sections50rwc
    RWC popularaudio: ❌, annotations: ✅beats, sections, vocal-activity, chords, tempo100rwc
    Salamiaudio: ❌, annotations: ✅sections1359CC0 1.0
    Saraga Carnaticaudio: ✅, annotations: ✅f0, tempo, phrases, beats, sections, tonic249CC BY-NC-SA 4.0
    Saraga Hindustaniaudio: ✅, annotations: ✅f0, tempo, phrases, beats, sections, tonic108CC BY-NC-SA 4.0
    Saraga Carnatic Melody Synthaudio: ✅, annotations: ✅f0, events2460CC BY-NC-SA 4.0
    SIMACaudio: ❌, annotations: ❌beats, tempo595CC0 1.0
    Slakhmultitrack audio: ✅, annotations: ✅Notes, Instruments1710CC BY 4.0
    Tinysolaudio: ✅, annotations: ✅instruments, technique, notes2913CC BY 4.0
    Tonality ClassicalDBaudio: 🧮, annotations: ✅Global key881CC BY-NC-SA 4.0
    TONASaudio: 🔑, annotations: 🔑f0, notes72tonas
    vocaditoaudio: ✅, annotations: ✅f0, notes, lyrics40CC BY-NC-SA 4.0
  2. Core capabilities of mirdata

    master

    mirdata is designed to provide reproducible access to MIR datasets through the following features:

    • Downloading: Automatically fetch datasets to a common location and format.
    • Validation: Verify that all required files for a specific dataset are present and correctly placed.
    • Annotation Loading: Load annotation files into a standardized format compatible with mir_eval.
    • Metadata Parsing: Parse track-level metadata to facilitate detailed evaluations.
  3. How to define multiple versions of a dataset

    master

    When a dataset has multiple versions (e.g., different annotation sets) but uses the same loading logic, you should write a single loader and define multiple versions using the INDEXES dictionary in the dataset module.

    • Use the naming convention <datasetname>_index_<version>.json for index files.
    • The INDEXES dictionary maps version names to core.Index objects.
    • Use the default key in INDEXES to specify which version is loaded by default when calling mirdata.initialize('dataset_name').
    • To load a specific version, use mirdata.initialize('dataset_name', version='version_name').
    • To avoid downloading unnecessary data for specific versions, use the partial_download argument in core.Index to specify a subset of keys from the REMOTES dictionary.
    INDEXES = {
        "default": "1.0",
        "test": "sample",
        "1.0": core.Index(filename="example_index_1.0.json", partial_download=['audio', 'v1-annotations']),
        "2.0": core.Index(filename="example_index_2.0.json", partial_download=['audio', 'v2-annotations']),
        "sample": core.Index(filename="example_index_sample.json")
    }
  4. What are Mirdata indexes?

    master

    Indexes in Mirdata are JSON manifests that map files in a dataset to their corresponding MD5 checksums. This ensures data integrity and provides a structured way to locate files.

    An index must contain a top-level version key. It can also include one or more of the following top-level keys depending on the dataset organization:

    • metadata: Maps metadata file names to their relative paths and MD5 checksums.
    • tracks: Used for datasets organized as individual tracks (mono/multi-channel audio, spectrograms, etc.).
    • multitracks: Used for datasets comprising groups of related tracks.
    • records: Used for datasets organized as groups of tables (e.g., relational databases).
    {
        "version": "1.0.0",
        "metadata": {
            "metadata_file_1": [
                "path_to_metadata/metadata_file_1.csv",
                "bb8b0ca866fc2423edde01325d6e34f7"
            ]
        },
        "tracks": {
            "track1": {
                "audio": ["audio_files/track1.wav", "6c77777ce77a06541cdb9f0a671afb46"],
                "beats": ["annotations/track1.beats.csv", "ab8b0ca866fc2423edde01325d6e34f7"]
            }
        }
    }
  5. Use smart_open for remote filesystem compatibility

    master

    To support reading data from remote filesystems, use the smart_open library instead of the standard Python open command.

    When checking for file existence, avoid os.path.exists. Instead, attempt to open the file and catch the FileNotFoundError. This ensures compatibility with remote storage where os.path operations may fail.

    from smart_open import open
    
    # Instead of os.path.exists, use try/except with open
    try:
        with open(file_path, "r") as fhandle:
            ...
    except FileNotFoundError:
        raise FileNotFoundError(f"{file_path} not found, did you run .download?")
  6. Create a dataset index

    master

    Mirdata uses indexes (dictionaries) to store information about dataset files, their locations, and checksums. This is necessary for loading and validation.

    Steps to create an index:

    1. Create a script in scripts/ (e.g., make_dataset_index.py) that automates index generation by computing MD5 checksums for files at a given data_path.
    2. Run the script on the dataset.
    3. Save the resulting index in mirdata/datasets/indexes/ as dataset_index_<version>.json (e.g., dataset_index_1.0.json).

    Index Structure for Tracks

    Most MIR datasets use a tracks top-level key. This is a dictionary where keys are unique track IDs and values are dictionaries of files (audio, annotations, etc.) and their checksums. File paths must be relative to the dataset's top-level directory.

    Index Structure for Multitracks

    Multitrack datasets can use a multitracks top-level key to group related tracks and provide master/mix audio files.

    {
        "version": "1.0",
        "tracks": {
            "track1": {
                "audio": [
                    "audio/track1.wav",
                    "912ec803b2ce49e4a541068d495ab570"
                ],
                "annotation": [
                    "annotations/track1.csv",
                    "2cf33591c3b28b382668952e236cccd5"
                ]
            }
        },
        "metadata": {
            "metadata_file": [
                "metadata/metadata_file.csv",
                "7a41b280c7b74e2ddac5184708f9525b"
            ]
        }
    }
  7. Standardization of annotations and metadata

    master

    Mirdata aims to provide a consistent interface for different datasets while respecting their unique characteristics:

    • Annotations: Mirdata provides various Annotation classes that offer a standard interface. These are designed to be compatible with the mir_eval library.
    • Metadata: When available, metadata is exposed as attributes on the track object (e.g., track.artist).
    • Dataset Idiosyncrasies: Because different datasets have different properties, track objects may have different available attributes. For example, one dataset might have an artist attribute while another does not.
  8. Reference common music annotation types

    master

    Mirdata supports a wide variety of annotation types. Because definitions can vary by dataset, it is strongly recommended to read the specific dataset documentation to ensure the data matches your expectations.

    Common annotation categories include:

    • Rhythmic/Temporal: Beats (timestamps/positions), Meter (rhythmic meter), Tempo (BPM, can be global or time-varying), and Sections (boundary timestamps).
    • Pitch/Melody: F0 (pitch contours), Melody (F0 or Notes), Notes (pitch events), Key (musical key), Mode (musical mode), and Tonic (tonal center).
    • Harmonic: Chords (labeled events like 'A:m7').
    • Instrumental/Percussive: Drums (drum instrument events), Instruments (presence/absence), Technique (playing style, e.g., 'Pizzicato'), and Stroke Name (specific instrument stroke types).
    • Vocal/Lyrics: Lyrics (text or time-aligned), Phonemes (sung phonemes), Syllables (structured syllables), and Vocal Activity (presence of singing).
    • Other: Genre (global tags), Tags (broad, often free-form labels), Effect (audio effects), Matches (audio fingerprinting identifications), and Segments (specific event segments).
  9. Using mirdata with PyTorch or TensorFlow

    master
    While mirdata does not provide built-in loaders specifically formatted for PyTorch or TensorFlow (due to the high variety in annotation types and tasks), it serves as a foundational step. You can easily build custom PyTorch or TensorFlow data loaders on top of mirdata loaders.