datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audio-course-imagesaudio-filesmy-audio-appBIRDeep_AudioAnnotations
BIRDeep Audio Annotations
The BIRDeep Audio Annotations dataset is a collection of bird vocalizations from Doñana National Park, Spain. It was created as part of the BIRDeep project, which aims to optimize the detection and classification of bird species in audio recordings using deep learning techniques. The dataset is intended for use in training and evaluating models for bird vocalization detection and identification.
The research code and further information is available at… See the full description on the dataset page: https://huggingface.co/datasets/GrunCrow/BIRDeep_AudioAnnotations.cuu_tieu_audiovan_hai_audiodao_tam_audiotien_lo_audio426huyen_khong_audiohuyen_vu_audioquran-audioBIRDeep_AudioAnnotations
BIRDeep Audio Annotations
The BIRDeep Audio Annotations dataset is a collection of bird vocalizations from Doñana National Park, Spain. It was created as part of the BIRDeep project, which aims to optimize the detection and classification of bird species in audio recordings using deep learning techniques. The dataset is intended for use in training and evaluating models for bird vocalization detection and identification.
The research code and further information is available at… See the full description on the dataset page: https://huggingface.co/datasets/TenzinL/BIRDeep_AudioAnnotations.sovits_audio_preview
预览.
简体中文|
English|
日本語
本仓库用于预览so-vits-svc-4.0训练出的各种语音模型的效果,点击角色名自动跳转对应训练参数。
推荐用谷歌浏览器,其他浏览器可能无法正确加载预览的音频。
正常说话的音色转换较为准确,歌曲包含较广的音域且bgm和声等难以去除干净,效果有所折扣。
有推荐的歌想要转换听听效果,或者其他内容建议,点我发起讨论
下面是预览音频,上下左右滑动可以看到全部
角色名
角色原声A
被转换人声BA音色替换B
A音色翻唱(点击直接下载)
散兵
夢で会えたら
胡桃
.........
.........
moonlight shadow,
云烟成雨… See the full description on the dataset page: https://huggingface.co/datasets/jiaheillu/sovits_audio_preview.Omni_Bench_fixgradient_accumulation_exampleaudiobench_rendertextvaani-audio-image-retrieval
Vaani audio–image retrieval (MTEB)
Multilingual audio↔image retrieval over 62 Indian languages, derived from
Project Vaani (IISc Bangalore /
ARTPARK).
Vaani records image-prompted speech: a speaker is shown a photograph and describes it
aloud in their own language. Each recording is therefore grounded in a specific image,
which is what makes audio↔image retrieval well defined without any extra annotation.
Prepared for MTEB as
VaaniA2IRetrieval and VaaniI2ARetrieval.… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/vaani-audio-image-retrieval.audio-diffusion-1024Over 20,000 256x256 mel spectrograms of 5 second samples of music from my Spotify liked playlist. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models.
x_res = 1024
y_res = 1024
sample_rate = 44100
n_fft = 2048
hop_length = 512
audio-dataset-flickr-soundnetAmharic_Audio_and_Spectrograms
Amharic Audio Spectrogram Dataset
Dataset Info
Total samples in full dataset: 662,611
Samples in this preview: 1,000
Audio duration: 2.49 ± 1.60 seconds
Sample rate: 16kHz
Spectrogram dimensions: 80 mel bins × variable time steps
Sample Data
Audio Sample
Spectrogram
License
Apache 2.0
enem-audiodescricao
Audiodescrição profissional de imagens do ENEM: corpus em português para acessibilidade e avaliação de modelos de visão
Professional audio descriptions of ENEM exam figures: a Brazilian Portuguese corpus for accessibility research and vision-language model evaluation.
Corpus de audiodescrições escritas por profissionais para as figuras do ENEM, extraídas
dos cadernos "ledor" que o INEP publica para participantes com deficiência visual, alinhadas
à figura, ao enunciado, às… See the full description on the dataset page: https://huggingface.co/datasets/tardellirs/enem-audiodescricao.mead_hdtf_400_merge_video_audio_frames_onlykiki-bouba-audio-20k
🔊 Kiki–Bouba, Spoken Aloud (20k Global Responses)
Dataset Summary
This dataset is the audio companion to
Rapidata/psychology-association-kiki-bouba-etc.
In the original dataset, respondents read the question "Which one is called 'Kiki'?" as written text.
Here, respondents instead hear the word spoken aloud — the task shows the same two shapes
(a rounded blob and a spiky star) while a short audio clip of "kiki" or "bouba" plays as context.
The annotator UI… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/kiki-bouba-audio-20k.audio-diffusion-512Over 20,000 512x512 mel spectrograms of 5 second samples of music from my Spotify liked playlist. The code to convert from audio to spectrogram and vice versa can be found in https://github.com/teticio/audio-diffusion along with scripts to train and run inference using De-noising Diffusion Probabilistic Models.
x_res = 512
y_res = 512
sample_rate = 22050
n_fft = 2048
hop_length = 512
quillan-audio-media
Quillan-Ronin: Multimodal Audio & Media Dataset
Full multimodal audio and visual collection produced and designed by Quillan-Ronin:
Lossless FLAC Master Recordings (The Sound of Alchemy, singles)
High-Bitrate MP3 Releases (Draming of the Sky, Rock Album, ai beats, stems)
Visual Concepts & Model Diagrams
flickr-audio-imagemel_spectogram_bird_audio
Dataset Card for "mel_spectogram_bird_audio"
More Information needed
video-dataset-audio_dataset
Video Dataset - audio_dataset
Dataset Description
This dataset contains video frames extracted from annotated video segments, along with annotations, transcriptions, and corresponding video clips. Combined from tasks: task06, task07, task08
Dataset Structure
frames/ — extracted frames (first frame from each segment)
segments/ — video clips for each annotation interval
annotations/ — original JSON annotation
transcriptions/ — transcription files… See the full description on the dataset page: https://huggingface.co/datasets/Quazitron420/video-dataset-audio_dataset.kiem_vu_audiokiem_y_audio
