wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512
midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512 Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full piece (no time-windowing); training crops sequences from packed token bins. The source column is the original piece metadata as JSON so a row can be traced back to Maestro, GiantMIDI, ATEPP, or MusicNet. Based on MIDI datasets gathered by EPR Labs. Codec name: dyadic tokenizer vocab size: 512 max_time_step: 1.0 n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512.
midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512
Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full piece (no time-windowing); training crops sequences from packed token bins. The source column is the original piece metadata as JSON so a row can be traced back to Maestro, GiantMIDI, ATEPP, or MusicNet.
Based on MIDI datasets gathered by EPR Labs.
Codec
- name:
dyadic - tokenizer vocab size:
512
max_time_step:1.0n_velocity_bins:32time_unit:0.01
The repository also stores midi_codec.json and Hugging Face tokenizer files so token ids can be rebuilt exactly.
Source datasets
maestro-sustain-v2giant-midi-sustain-v2atepp-1.1-sustain-v2music-net
Maestro validation / test are the only evaluation splits. All other datasets contribute training pieces only.
Token counts
Totals: train 412407397 tokens, validation 3250843 tokens, test 3714020 tokens.
