datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DFADD_MLAAD_DiffSSD_VoxCeleb2audio-diffusion-1024Over 20,000 256x256 mel spectrograms of 5 second samples of music from my Spotify liked playlist. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models.
x_res = 1024
y_res = 1024
sample_rate = 44100
n_fft = 2048
hop_length = 512
GLOBE_V3_age_N_allsplitsaudio-diffusion-512Over 20,000 512x512 mel spectrograms of 5 second samples of music from my Spotify liked playlist. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models.
x_res = 512
y_res = 512
sample_rate = 22050
n_fft = 2048
hop_length = 512
odia-english-ASRaudio-diffusion-instrumental-hiphop-256256x256 mel spectrograms of 5 second samples of instrumental Hip Hop. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models.
x_res = 256
y_res = 256
sample_rate = 22050
n_fft = 2048
hop_length = 512
audio-diffusion-breaks-25630,000 256x256 mel spectrograms of 5 second samples that have been used in music, sourced from WhoSampled and YouTube. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models.
x_res = 256
y_res = 256
sample_rate = 22050
n_fft = 2048
hop_length = 512
tau-2024-mobile-dev-miniGLOBE_V3_gender_N_allsplitsMSPI_WAV_Diff_Curriculum
IEMOCAP with Curriculum Learning Metrics
This dataset enhances the original IEMO_WAV_Diff_2 dataset with inter-evaluator agreement metrics
for curriculum learning following Lotfian & Busso (2019).
Additional Columns
curriculum_order: Training order (1=highest agreement, train first)
overall_agreement: Combined agreement score (0-1, higher is better)
fleiss_kappa: Categorical agreement (-1 to 1, higher is better)
krippendorff_alpha: Krippendorff's alpha for categorical… See the full description on the dataset page: https://huggingface.co/datasets/cairocode/MSPI_WAV_Diff_Curriculum.audio-diffusion-256Over 20,000 256x256 mel spectrograms of 5 second samples of music from my Spotify liked playlist. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models.
x_res = 256
y_res = 256
sample_rate = 22050
n_fft = 2048
hop_length = 512
Different-Voicescstr-vctk-age-miniworst100-testclean-clips
Worst-100 test-clean clips — audio, transcripts, and the vocabulary finding
The 100 LibriSpeech test-clean clips where the block-4 production model (4.60% WER) made
the most word errors — with audio embedded so the failures can be listened to, plus the
model's transcript next to the reference for each clip.
The finding this dataset produced
47% of the word errors in these clips are on words that never appeared in the 30-hour
training vocabulary at all (20,066… See the full description on the dataset page: https://huggingface.co/datasets/Diffusion-ASR/worst100-testclean-clips.feji-78-different-moodsnew_data_set_same_model_diff_dataGLOBE_V2_age_5klass-synthIEMO_WAV_Diff_2lass-synth-retrieval-miniexpresso_conv_miniaudio-diffusion-breaks-25630,000 256x256 mel spectrograms of 5 second samples that have been used in music, sourced from WhoSampled and YouTube. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models.
x_res = 256
y_res = 256
sample_rate = 22050
n_fft = 2048
hop_length = 512
lj_speech_DifferentStructure
Dataset Card for "lj_speech_DifferentStructure"
More Information needed
expresso_read_minicstr-vctk-accents-minilj_speech_DifferentStructure_removedVocabs
Dataset Card for "lj_speech_DifferentStructure_removedVocabs"
More Information needed
GLOBE_v2_test_splitGLOBE_V2_testGLOBE_V2_gender_5kcstr-vctk-gender-mini
