datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
midi-audio-abc_300smidi, synthesized audio, ABC code triples
(this dataset contains those with audio duration in 5-300s, several subsets with smaller duration
60s
30s
10s)
(token_length_abc field represents the token count of the abc text w.r.t. Qwen3's tokenizer)
midi files are from bread-midi-dataset
synthesized audio: use Don Allen's Timbres of Heaven as soundfont and FluidSynth as synthesizer
abc notation: mid2abc by EasyABC (midi2abc.py)
Citation
@misc{jiang2025advancingfoundationmodelmusic… See the full description on the dataset page: https://huggingface.co/datasets/Yi3852/midi-audio-abc_300s.midi-audio-abc_60smidi, synthesized audio, ABC code triples
(this dataset contains those with audio duration in 5-60s, sampled from the full set with max 300s duration)
(token_length_abc field represents the token count of the abc text w.r.t. Qwen3's tokenizer)
midi files are from bread-midi-dataset
synthesized audio: use Don Allen's Timbres of Heaven as soundfont and FluidSynth as synthesizer
abc notation: mid2abc by EasyABC (midi2abc.py)
Citation
@misc{jiang2025advancingfoundationmodelmusic… See the full description on the dataset page: https://huggingface.co/datasets/Yi3852/midi-audio-abc_60s.midi-audio-abc_longmidi, synthesized audio, ABC code triples
(this dataset contains those with audio duration in 5 min - 2 hours, less than 5 min data are in 300s
and there are several subsets with smaller duration
60s
30s
10s)
(token_length_abc field represents the token count of the abc text w.r.t. Qwen3's tokenizer)
midi files are from bread-midi-dataset
synthesized audio: use Don Allen's Timbres of Heaven as soundfont and FluidSynth as synthesizer
abc notation: mid2abc by EasyABC (midi2abc.py)… See the full description on the dataset page: https://huggingface.co/datasets/Yi3852/midi-audio-abc_long.midi-audio-abc_30smidi, synthesized audio, ABC code triples
(this dataset contains those with audio duration in 5-30s, sampled from the full set with max 300s duration)
(token_length_abc field represents the token count of the abc text w.r.t. Qwen3's tokenizer)
midi files are from bread-midi-dataset
synthesized audio: use Don Allen's Timbres of Heaven as soundfont and FluidSynth as synthesizer
abc notation: mid2abc by EasyABC (midi2abc.py)
Citation
@misc{jiang2025advancingfoundationmodelmusic… See the full description on the dataset page: https://huggingface.co/datasets/Yi3852/midi-audio-abc_30s.MODELOSDETESTEminds14
MInDS-14
MINDS-14 is training and evaluation resource for intent detection task with spoken data. It covers 14
intents extracted from a commercial system in the e-banking domain, associated with spoken examples in 14 diverse language varieties.
Example
MInDS-14 can be downloaded and used as follows:
from datasets import load_dataset
minds_14 = load_dataset("PolyAI/minds14", "fr-FR") # for French
# to download all data for multi-lingual fine-tuning uncomment… See the full description on the dataset page: https://huggingface.co/datasets/abc-123-456/minds14.midi-audio-abc_10smidi, synthesized audio, ABC code triples
(this dataset contains those with audio duration in 5-10s, sampled from the full set with max 300s duration)
(token_length_abc field represents the token count of the abc text w.r.t. Qwen3's tokenizer)
midi files are from bread-midi-dataset
synthesized audio: use Don Allen's Timbres of Heaven as soundfont and FluidSynth as synthesizer
abc notation: mid2abc by EasyABC (midi2abc.py)
Citation
@misc{jiang2025advancingfoundationmodelmusic… See the full description on the dataset page: https://huggingface.co/datasets/Yi3852/midi-audio-abc_10s.abc.hassaniya_ASRindian-speech-audio-extendedswamiji-artifact-abc-grid
Where does the end-of-clip artifact come from?
Reported symptom: short polite replies "mess up at the end", and a "huge high"
is audible after the words finish — sometimes even when the sentence ends in a
full stop. Three candidate causes: the model, the streaming, or the Opus codec.
Each phrase here is generated twice — ending in ! and ending in . — and
each generation is rendered three ways, so exactly one variable moves at a
time. Six players per row.
rendering
what it… See the full description on the dataset page: https://huggingface.co/datasets/sw-voice/swamiji-artifact-abc-grid.BABYMONSTERTESTEvocalsound-throat-sneezeabcMODELSTESTE2abcs_16k_headset_templevibravox_abcs_merge_headset_temple_16kabc_data
