datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
midi-audio-abc_60smidi, synthesized audio, ABC code triples
(this dataset contains those with audio duration in 5-60s, sampled from the full set with max 300s duration)
(token_length_abc field represents the token count of the abc text w.r.t. Qwen3's tokenizer)
midi files are from bread-midi-dataset
synthesized audio: use Don Allen's Timbres of Heaven as soundfont and FluidSynth as synthesizer
abc notation: mid2abc by EasyABC (midi2abc.py)
Citation
@misc{jiang2025advancingfoundationmodelmusic… See the full description on the dataset page: https://huggingface.co/datasets/Yi3852/midi-audio-abc_60s.solo-leveling-60s-assetsautocorrelation-spend-time-60s-diversettm_validation_dataset_60sec3_grpo_train_data_60smovie_gen_60s_real_refine_step4_nvfp4hdtf_400_audio_60s_chunksDitto_videos_hdtf_400_audio_60s_chunks_all_preprocess
