datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
simplified_grooveThis is a copy of the Magenta Groove dataset
The script ´simplify_midi_pretty.py` reads the midi data and simplifies it, by removing any midi values that aren't kicks or snares, and quantizing the notes.
ycuppe-midi
YCU-PPE-III: Piano Performance MIDI Dataset
MIDI transcriptions of the YCU-PPE-III piano performance dataset (Wang et al.), used for unreferenced Performance MOS (PMOS) prediction in EVPMR.
Overview
2,627 MIDI files transcribed from WAV recordings via transkun
13 songs performed by student pianists
2,511 performances with ratings from 3 expert judges (0-100 scale each)
Labels: normalized mean score to [0, 1]
Splits: 1,757 train / 377 val / 377 test (stratified by song)… See the full description on the dataset page: https://huggingface.co/datasets/anusfoil/ycuppe-midi.MidiCaps
MidiCaps Dataset
The MidiCaps dataset [1] is a large-scale dataset of 168,385 midi music files with descriptive text captions, and a set of extracted musical features.
The captions have been produced through a captioning pipeline incorporating MIR feature extraction and LLM Claude 3 to caption the data from extracted features with an in-context learning task. The framework used to extract the captions is available open source on github.
The original MIDI files originate from the… See the full description on the dataset page: https://huggingface.co/datasets/amaai-lab/MidiCaps.midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs
midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs
Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full
piece (no time-windowing); training crops sequences from packed token bins.
The source column is the original piece metadata as JSON so a row can be
traced back to its EPR Labs source dataset.
Based on MIDI datasets gathered by EPR Labs.
Codec
name: dyadic
tokenizer vocab size: 512
max_time_step: 1.0
n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs.midi-dataset
[aria-midi-v1-pruned-ext] with midi column for midi file
midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-aug-lessmidiFilesMidiCapsmidi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512
midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512
Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full
piece (no time-windowing); training crops sequences from packed token bins.
The source column is the original piece metadata as JSON so a row can be
traced back to Maestro, GiantMIDI, ATEPP, or MusicNet.
Based on MIDI datasets gathered by EPR Labs.
Codec
name: dyadic
tokenizer vocab size: 512
max_time_step: 1.0
n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512.MIDICaps-500-labelmidi-dataset-3midi-dataset-4ABC-Lakh-MIDI-Dataset
🎵 Mader ABC V2 - Music Dataset
This dataset contains music pieces derived from the Lakh MIDI Dataset and stylistically aligned with the Million Song Dataset (for musical genres), converted to ABC notation and tokenized for NLP-style applications on music (classification, generation, clustering, ...).
It provides two configurations, each with instrument and genre metadata:
abc_texts – text in ABC format
abc_tokens – token sequence
Each configuration contains one entry per… See the full description on the dataset page: https://huggingface.co/datasets/Gapagapi1/ABC-Lakh-MIDI-Dataset.midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512-epr-labs-aug-lessmidicaps_benchmarkmidi-dataset-2
[MidiCaps] with midi column for midi files and mismatch time signature filtered
combined-30k-miditokMidiMus-1.5-BIGmidi_etlmaster_midi_testingmlops_gsodaria-midi-v1-pruned-extmidi_preprocessaria-midi-v1-unique-extMIDICaps-500midi-generation-datasetpianojudge-techniques-midimidi-classical-music-toio-json-audit
MIDI Classical Music toio JSON — aggregate audit
This metadata-only audit describes ayousanz/midi-classical-music-toio-json at
revision d07a0210bb7cff7757b9d941b131df10a752eb8c. It contains no MIDI files,
converted song JSON, filenames, or recovered source payloads.
Of 4,796 source MIDI files, 4,712 have conversions. All 84 missing conversions were
invalid under strict SMF parsing; 28 could nevertheless be recovered as RIFF/RMID or
MacBinary containers. Across the 4,712 valid… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/midi-classical-music-toio-json-audit.
