CoolFace
Datasetpublic

wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512

midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512 Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full piece (no time-windowing); training crops sequences from packed token bins. The source column is the original piece metadata as JSON so a row can be traced back to Maestro, GiantMIDI, ATEPP, or MusicNet. Based on MIDI datasets gathered by EPR Labs. Codec name: dyadic tokenizer vocab size: 512 max_time_step: 1.0 n_velocity_bins: 32… See the full description on the dataset page: https://huggingface.co/datasets/wmatejuk/midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512.

sourceHugging Faceupdated 1mo agoView on Hugging Face
0likes67downloads
Dataset Card

midi-tokens-dyadic-tu0.01-vb32-mts1.0-vocab512

Pre-tokenized MIDI pieces for IsoFLOP scaling-law runs. Each row is one full piece (no time-windowing); training crops sequences from packed token bins. The source column is the original piece metadata as JSON so a row can be traced back to Maestro, GiantMIDI, ATEPP, or MusicNet.

Based on MIDI datasets gathered by EPR Labs.

Codec

  • —name: dyadic
  • —tokenizer vocab size: 512
  • —max_time_step: 1.0
  • —n_velocity_bins: 32
  • —time_unit: 0.01

The repository also stores midi_codec.json and Hugging Face tokenizer files so token ids can be rebuilt exactly.

Source datasets

  • —maestro-sustain-v2
  • —giant-midi-sustain-v2
  • —atepp-1.1-sustain-v2
  • —music-net

Maestro validation / test are the only evaluation splits. All other datasets contribute training pieces only.

Token counts

datasetsource splitoutput splitpiecesnotestokensskipped
maestro-sustain-v2traintrain9625659329285698830
maestro-sustain-v2validationvalidation13763942532508430
maestro-sustain-v2testtest17774141037140200
giant-midi-sustain-v2traintrain10852386985271988089562
atepp-1.1-sustain-v2traintrain11677319560051799677760
music-nettraintrain323110266250607820

Totals: train 412407397 tokens, validation 3250843 tokens, test 3714020 tokens.