CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cmu-mlsp /DFADD_MLAAD_DiffSSD_VoxCeleb2audio1M<n<10M1 likes606 downloads1y agoHugging Face02teticio /audio-diffusion-1024Over 20,000 256x256 mel spectrograms of 5 second samples of music from my Spotify liked playlist. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models. x_res = 1024 y_res = 1024 sample_rate = 44100 n_fft = 2048 hop_length = 512 imageimage-to-image10K<n<100K0 likes348 downloads4y agoHugging Face03diffunity /GLOBE_V3_age_N_allsplitsaudio10K<n<100K0 likes199 downloads8mo agoHugging Face04teticio /audio-diffusion-512Over 20,000 512x512 mel spectrograms of 5 second samples of music from my Spotify liked playlist. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models. x_res = 512 y_res = 512 sample_rate = 22050 n_fft = 2048 hop_length = 512 imageimage-to-image10K<n<100K2 likes126 downloads3y agoHugging Face05Mohan-diffuser /odia-english-ASRaudioautomatic-speech-recognition1K<n<10K0 likes98 downloads1y agoHugging Face06teticio /audio-diffusion-instrumental-hiphop-256256x256 mel spectrograms of 5 second samples of instrumental Hip Hop. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models. x_res = 256 y_res = 256 sample_rate = 22050 n_fft = 2048 hop_length = 512 imageimage-to-image10K<n<100K7 likes76 downloads4y agoHugging Face07teticio /audio-diffusion-breaks-25630,000 256x256 mel spectrograms of 5 second samples that have been used in music, sourced from WhoSampled and YouTube. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models. x_res = 256 y_res = 256 sample_rate = 22050 n_fft = 2048 hop_length = 512 imageimage-to-image10K<n<100K2 likes74 downloads4y agoHugging Face08diffunity /tau-2024-mobile-dev-miniaudio10K<n<100K0 likes69 downloads8mo agoHugging Face09diffunity /GLOBE_V3_gender_N_allsplitsaudio10K<n<100K0 likes52 downloads8mo agoHugging Face10cairocode /MSPI_WAV_Diff_Curriculum IEMOCAP with Curriculum Learning Metrics This dataset enhances the original IEMO_WAV_Diff_2 dataset with inter-evaluator agreement metrics for curriculum learning following Lotfian & Busso (2019). Additional Columns curriculum_order: Training order (1=highest agreement, train first) overall_agreement: Combined agreement score (0-1, higher is better) fleiss_kappa: Categorical agreement (-1 to 1, higher is better) krippendorff_alpha: Krippendorff's alpha for categorical… See the full description on the dataset page: https://huggingface.co/datasets/cairocode/MSPI_WAV_Diff_Curriculum.audio1K<n<10K0 likes51 downloads1y agoHugging Face11teticio /audio-diffusion-256Over 20,000 256x256 mel spectrograms of 5 second samples of music from my Spotify liked playlist. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models. x_res = 256 y_res = 256 sample_rate = 22050 n_fft = 2048 hop_length = 512 imageimage-to-image10K<n<100K6 likes34 downloads4y agoHugging Face12Greenbean /Different-Voicesaudio1K<n<10K0 likes34 downloads1mo agoHugging Face13diffunity /cstr-vctk-age-miniaudio10K<n<100K0 likes21 downloads8mo agoHugging Face14Diffusion-ASR /worst100-testclean-clips Worst-100 test-clean clips — audio, transcripts, and the vocabulary finding The 100 LibriSpeech test-clean clips where the block-4 production model (4.60% WER) made the most word errors — with audio embedded so the failures can be listened to, plus the model's transcript next to the reference for each clip. The finding this dataset produced 47% of the word errors in these clips are on words that never appeared in the 30-hour training vocabulary at all (20,066… See the full description on the dataset page: https://huggingface.co/datasets/Diffusion-ASR/worst100-testclean-clips.audion<1K0 likes21 downloads2mo agoHugging Face15AhmetSemih /feji-78-different-moodsgatedaudion<1K0 likes16 downloads4d agoHugging Face16Almoooo /new_data_set_same_model_diff_dataaudion<1K0 likes15 downloads1y agoHugging Face17diffunity /GLOBE_V2_age_5kaudio10K<n<100K0 likes14 downloads9mo agoHugging Face18diffunity /lass-synthaudio1K<n<10K0 likes14 downloads8mo agoHugging Face19cairocode /IEMO_WAV_Diff_2audio1K<n<10K0 likes13 downloads1y agoHugging Face20diffunity /lass-synth-retrieval-miniaudio1K<n<10K0 likes11 downloads9mo agoHugging Face21diffunity /expresso_conv_miniaudio1K<n<10K0 likes9 downloads9mo agoHugging Face22ChrisOpenSource /audio-diffusion-breaks-25630,000 256x256 mel spectrograms of 5 second samples that have been used in music, sourced from WhoSampled and YouTube. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models. x_res = 256 y_res = 256 sample_rate = 22050 n_fft = 2048 hop_length = 512 imageimage-to-image10K<n<100K0 likes9 downloads2mo agoHugging Face23HamdanXI /lj_speech_DifferentStructure Dataset Card for "lj_speech_DifferentStructure" More Information needed audio1K<n<10K0 likes8 downloads3y agoHugging Face24diffunity /expresso_read_miniaudio1K<n<10K0 likes8 downloads9mo agoHugging Face25diffunity /cstr-vctk-accents-miniaudio10K<n<100K0 likes8 downloads8mo agoHugging Face26HamdanXI /lj_speech_DifferentStructure_removedVocabs Dataset Card for "lj_speech_DifferentStructure_removedVocabs" More Information needed audio1K<n<10K0 likes7 downloads3y agoHugging Face27diffunity /GLOBE_v2_test_splitaudio1K<n<10K0 likes7 downloads9mo agoHugging Face28diffunity /GLOBE_V2_testaudio1K<n<10K0 likes7 downloads9mo agoHugging Face29diffunity /GLOBE_V2_gender_5kaudio10K<n<100K0 likes7 downloads9mo agoHugging Face30diffunity /cstr-vctk-gender-miniaudio10K<n<100K0 likes7 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.