CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01huggingface-course /audio-course-imagesimagen<1K0 likes11k downloads3y agoHugging Face02nader39 /audio-filesaudion<1K0 likes5.9k downloads57m agoHugging Face03Lakshya1234 /my-audio-appaudion<1K0 likes1.1k downloads6h agoHugging Face04GrunCrow /BIRDeep_AudioAnnotations BIRDeep Audio Annotations The BIRDeep Audio Annotations dataset is a collection of bird vocalizations from Doñana National Park, Spain. It was created as part of the BIRDeep project, which aims to optimize the detection and classification of bird species in audio recordings using deep learning techniques. The dataset is intended for use in training and evaluating models for bird vocalization detection and identification. The research code and further information is available at… See the full description on the dataset page: https://huggingface.co/datasets/GrunCrow/BIRDeep_AudioAnnotations.audioaudio-classificationn<1K2 likes884 downloads11mo agoHugging Face05raymondt /cuu_tieu_audioimagen<1K0 likes786 downloads19h agoHugging Face06raymondt /van_hai_audioimagen<1K0 likes758 downloads19h agoHugging Face07raymondt /dao_tam_audioimagen<1K0 likes679 downloads19h agoHugging Face08raymondt /tien_lo_audio426imagen<1K0 likes672 downloads20h agoHugging Face09raymondt /huyen_khong_audioimagen<1K0 likes672 downloads14h agoHugging Face10raymondt /huyen_vu_audioimagen<1K0 likes605 downloads19h agoHugging Face11hammoualiyoucef20 /quran-audioaudion<1K0 likes601 downloads10d agoHugging Face12TenzinL /BIRDeep_AudioAnnotations BIRDeep Audio Annotations The BIRDeep Audio Annotations dataset is a collection of bird vocalizations from Doñana National Park, Spain. It was created as part of the BIRDeep project, which aims to optimize the detection and classification of bird species in audio recordings using deep learning techniques. The dataset is intended for use in training and evaluating models for bird vocalization detection and identification. The research code and further information is available at… See the full description on the dataset page: https://huggingface.co/datasets/TenzinL/BIRDeep_AudioAnnotations.audioaudio-classificationn<1K0 likes511 downloads7mo agoHugging Face13jiaheillu /sovits_audio_preview 预览. 简体中文| English| 日本語 本仓库用于预览so-vits-svc-4.0训练出的各种语音模型的效果,点击角色名自动跳转对应训练参数。 推荐用谷歌浏览器,其他浏览器可能无法正确加载预览的音频。 正常说话的音色转换较为准确,歌曲包含较广的音域且bgm和声等难以去除干净,效果有所折扣。 有推荐的歌想要转换听听效果,或者其他内容建议,点我发起讨论 下面是预览音频,上下左右滑动可以看到全部 角色名 角色原声A 被转换人声BA音色替换B A音色翻唱(点击直接下载) 散兵 夢で会えたら 胡桃 ......... ......... moonlight shadow, 云烟成雨… See the full description on the dataset page: https://huggingface.co/datasets/jiaheillu/sovits_audio_preview.audion<1K8 likes503 downloads3y agoHugging Face14lmms-lab-audio /Omni_Bench_fixaudio1K<n<10K1 likes490 downloads1y agoHugging Face15hf-audio /gradient_accumulation_exampleimagen<1K0 likes419 downloads2y agoHugging Face16RyanWW /audiobench_rendertextaudio10K<n<100K0 likes382 downloads1y agoHugging Face17vnahata /vaani-audio-image-retrieval Vaani audio–image retrieval (MTEB) Multilingual audio↔image retrieval over 62 Indian languages, derived from Project Vaani (IISc Bangalore / ARTPARK). Vaani records image-prompted speech: a speaker is shown a photograph and describes it aloud in their own language. Each recording is therefore grounded in a specific image, which is what makes audio↔image retrieval well defined without any extra annotation. Prepared for MTEB as VaaniA2IRetrieval and VaaniI2ARetrieval.… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/vaani-audio-image-retrieval.audioaudio-classification1K<n<10K0 likes336 downloads23d agoHugging Face18teticio /audio-diffusion-1024Over 20,000 256x256 mel spectrograms of 5 second samples of music from my Spotify liked playlist. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models. x_res = 1024 y_res = 1024 sample_rate = 44100 n_fft = 2048 hop_length = 512 imageimage-to-image10K<n<100K0 likes299 downloads4y agoHugging Face19cssen /audio-dataset-flickr-soundnetaudion<1K2 likes265 downloads3y agoHugging Face20AhunInteligence /Amharic_Audio_and_Spectrograms Amharic Audio Spectrogram Dataset Dataset Info Total samples in full dataset: 662,611 Samples in this preview: 1,000 Audio duration: 2.49 ± 1.60 seconds Sample rate: 16kHz Spectrogram dimensions: 80 mel bins × variable time steps Sample Data Audio Sample Spectrogram License Apache 2.0 audio1K<n<10K0 likes247 downloads1y agoHugging Face21tardellirs /enem-audiodescricao Audiodescrição profissional de imagens do ENEM: corpus em português para acessibilidade e avaliação de modelos de visão Professional audio descriptions of ENEM exam figures: a Brazilian Portuguese corpus for accessibility research and vision-language model evaluation. Corpus de audiodescrições escritas por profissionais para as figuras do ENEM, extraídas dos cadernos "ledor" que o INEP publica para participantes com deficiência visual, alinhadas à figura, ao enunciado, às… See the full description on the dataset page: https://huggingface.co/datasets/tardellirs/enem-audiodescricao.imageimage-to-textn<1K1 likes199 downloads22d agoHugging Face22Darknsu /mead_hdtf_400_merge_video_audio_frames_onlyimage1M<n<10M0 likes165 downloads4mo agoHugging Face23Rapidata /kiki-bouba-audio-20k 🔊 Kiki–Bouba, Spoken Aloud (20k Global Responses) Dataset Summary This dataset is the audio companion to Rapidata/psychology-association-kiki-bouba-etc. In the original dataset, respondents read the question "Which one is called 'Kiki'?" as written text. Here, respondents instead hear the word spoken aloud — the task shows the same two shapes (a rounded blob and a spiky star) while a short audio clip of "kiki" or "bouba" plays as context. The annotator UI… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/kiki-bouba-audio-20k.audion<1K1 likes157 downloads1mo agoHugging Face24teticio /audio-diffusion-512Over 20,000 512x512 mel spectrograms of 5 second samples of music from my Spotify liked playlist. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models. x_res = 512 y_res = 512 sample_rate = 22050 n_fft = 2048 hop_length = 512 imageimage-to-image10K<n<100K2 likes141 downloads3y agoHugging Face25CrashOverrideX /quillan-audio-media Quillan-Ronin: Multimodal Audio & Media Dataset Full multimodal audio and visual collection produced and designed by Quillan-Ronin: Lossless FLAC Master Recordings (The Sound of Alchemy, singles) High-Bitrate MP3 Releases (Draming of the Sky, Rock Album, ai beats, stems) Visual Concepts & Model Diagrams audioaudio-to-audion<1K0 likes140 downloads24d agoHugging Face26deep9539 /flickr-audio-imageaudio10K<n<100K0 likes129 downloads2mo agoHugging Face27rachit8562 /mel_spectogram_bird_audio Dataset Card for "mel_spectogram_bird_audio" More Information needed image10K<n<100K3 likes115 downloads4y agoHugging Face28Quazitron420 /video-dataset-audio_dataset Video Dataset - audio_dataset Dataset Description This dataset contains video frames extracted from annotated video segments, along with annotations, transcriptions, and corresponding video clips. Combined from tasks: task06, task07, task08 Dataset Structure frames/ — extracted frames (first frame from each segment) segments/ — video clips for each annotation interval annotations/ — original JSON annotation transcriptions/ — transcription files… See the full description on the dataset page: https://huggingface.co/datasets/Quazitron420/video-dataset-audio_dataset.imageimage-classificationn<1K0 likes114 downloads10mo agoHugging Face29raymondt /kiem_vu_audioimagen<1K0 likes110 downloads18h agoHugging Face30raymondt /kiem_y_audioimagen<1K0 likes110 downloads18h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.