datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
maestro-v3.0.0
Maestro Dataset v3.0.0
Mirror copy as of 08/03/2025
MAESTRO (MIDI and Audio Edited for Synchronous TRacks and Organization) is a dataset composed of about 200 hours of virtuosic piano performances captured with fine alignment (~3 ms) between note labels and audio waveforms.
Source download Maestro Dataset v3.0.0
Citation
@inproceedings{
hawthorne2018enabling,
title={Enabling Factorized Piano Music Modeling and Generation with the {MAESTRO}… See the full description on the dataset page: https://huggingface.co/datasets/projectlosangeles/maestro-v3.0.0.maestro-v3.0.0maestro_synth
Dataset Card for "maestro_synth"
More Information needed
MAESTRO-v3.0-FLACmaestro-unidac4-ytsv
MAESTRO + ASAP audio and MIDI tokens (U-MusT)
Tokenized MAESTRO v3.0.0 for
U-MusT: DAC audio tokens and MT3-style MIDI event arrays,
covering roughly 199 hours of Disklavier-captured piano performance with precisely aligned MIDI.
This repository also contains ASAP-derived data. lmx/ and asap_note_events/ come from the
ASAP dataset, whose audio is itself MAESTRO. Both
carry the same license, so nothing conflicts, but the repository name mentions only one of the two
corpora it… See the full description on the dataset page: https://huggingface.co/datasets/malerlab/maestro-unidac4-ytsv.MAESTRO-E
Viewer note: default uses viewer_preview/ for responsive audio playback.
Full training/evaluation files remain available in the original folder structure.
MAESTRO-E
MAESTRO-E dataset for music practice error detection with paired performance/reference inputs and note-level error labels.
Paired Inputs for Error Detection
The model takes paired inputs:
mistake: performance audio/MIDI containing musical errors
score: paired reference score audio/MIDI (target/correct… See the full description on the dataset page: https://huggingface.co/datasets/ben2002chou/MAESTRO-E.maestro-romanticmaestro-classicalmaestro-modernMaestro20hmaestroMAESTRO-test
Mirror of the test set of MAESTRO V.3.0.0 dataset with midi and audio columns for midi file and audio in FLAC
Citation
If you use this dataset, please cite the original paper:
@inproceedings{
hawthorne2018enabling,
title={Enabling Factorized Piano Music Modeling and Generation with the {MAESTRO} Dataset},
author={Curtis Hawthorne and Andriy Stasyuk and Adam Roberts and Ian Simon and Cheng-Zhi Anna Huang and Sander Dieleman and Erich Elsen and Jesse Engel and… See the full description on the dataset page: https://huggingface.co/datasets/B-K/MAESTRO-test.MARBLEBeatTracking_ASAP_MAESTROAMoreNaturalMIDItoPianoGeneration_MaestroHEARMusicSpeechClassification_MAESTRO_LibrispeechHEARMusicTranscription_MAESTRO-5hr-Fold2HEARMusicTranscription_MAESTRO-5hr-Fold4HEARMusicTranscription_MAESTRO-5hr-Fold5MAESTRO_2004_SYNTH
MAESTRO-2004-SYNTH Dataset
This is a synthesized audio dataset using the midi of MAESTRO dataset [https://magenta.tensorflow.org/datasets/maestro].
Audio files are batch-synthesized in REAPER DAW [https://www.reaper.fm/] using superposition of triangle waves, square waves and sinusoid waves.
HEARMusicTranscription_MAESTRO-5hr-Fold3HEARMusicTranscription_MAESTRO-5hr-Fold1MAESTRO-E-preview
MAESTRO-E Preview
Small audio preview subset for ben2002chou/MAESTRO-E.
This dataset is intended for Hub viewer preview and quick sanity checks.
For full training/evaluation data, use ben2002chou/MAESTRO-E.
maestro-baroque
