datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Allo-AVAvad-human-ava-speech
Positive Transfer Of The Whisper Speech Transformer To Human And Animal Voice Activity Detection
We proposed WhisperSeg, utilizing the Whisper Transformer pre-trained for Automatic Speech Recognition (ASR) for both human and animal Voice Activity Detection (VAD). For more details, please refer to our paper
Positive Transfer of the Whisper Speech Transformer to Human and Animal Voice Activity Detection
Nianlong Gu, Kanghwi Lee, Maris Basha, Sumit Kumar Ram, Guanghao You, Richard H.… See the full description on the dataset page: https://huggingface.co/datasets/nccratliri/vad-human-ava-speech.nasle-mana-clean-chunked-30s-avasanj
Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks
Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here.
Splits
Split
Rows
Audio
Columns
labeled
4,981
41.41 hours
audio, label
to_transcribe
11,127
92.72 hours
audio
The labeled split contains the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s-avasanj.librivox_small_aud_onlyThis dataset is sampled from LibriVox audio (135+ hours). The purpose of this dataset is for training neural audio codecs or autoencoder to reconstruct high-fidelity audio data.
AvaMERGAVAGDdia-AvaAvd-test
AVA-AVD — test split (audio-visual speaker diarization in the wild)
Copie du split test d'AVA-AVD (Xu et al., ACM MM 2022), un corpus de
diarisation dans des conditions "in the wild" construit sur AVA Active Speaker.
Chaque clip (~5 min) est extrait des vidéos AVA à des offsets précis définis
par le repo officiel, puis l'audio est re-synchronisé : les timestamps des
RTTMs ont été recalculés en temps clip-local (soustrait min_start du
RTTM d'origine) pour être directement utilisables… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-AvaAvd-test.ava-avdKalliope-FalconH1-Avadec-LJSpeechAVATAR
AVATAR: What’s Making That Sound Right Now? Video-centric Audio-Visual Localization
AVATAR stands for Audio-Visual localizAtion benchmark for a spatio-TemporAl peRspective in video.
AVATAR is a benchmark dataset designed to evaluate video-centric audio-visual localization (AVL) in complex and dynamic real-world scenarios.Unlike previous benchmarks that rely on static image-level annotations and assume simplified conditions, AVATAR offers high-resolution temporal annotations over… See the full description on the dataset page: https://huggingface.co/datasets/mipal/AVATAR.Allo-AVAAvationATCliudmila-rubleuskaia-avantury-studyezusa-vyrvicha
Авантуры студыёзуса Вырвіча
Metadata
Author: Людміла Рублеўская
Title: Авантуры студыёзуса Вырвіча
Narrator:
Source Group: Дзіцячыя
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/liudmila-rubleuskaia-avantury-studyezusa-vyrvicha.avak-englishkhmer-voice-avatarvad-human-ava-speech
Positive Transfer Of The Whisper Speech Transformer To Human And Animal Voice Activity Detection
We proposed WhisperSeg, utilizing the Whisper Transformer pre-trained for Automatic Speech Recognition (ASR) for both human and animal Voice Activity Detection (VAD). For more details, please refer to our paper
Positive Transfer of the Whisper Speech Transformer to Human and Animal Voice Activity Detection
Nianlong Gu, Kanghwi Lee, Maris Basha, Sumit Kumar Ram, Guanghao You, Richard H.… See the full description on the dataset page: https://huggingface.co/datasets/0x3/vad-human-ava-speech.mydatasetliudmila-rubleuskaia-avantury-prantsisha-vyrvicha-shkaliara-i-shpega
Авантуры Пранціша Вырвіча, шкаляра і шпега
Metadata
Author: Людміла Рублеўская
Title: Авантуры Пранціша Вырвіча, шкаляра і шпега
Narrator:
Source Group: Дзіцячыя
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/liudmila-rubleuskaia-avantury-prantsisha-vyrvicha-shkaliara-i-shpega.chunked-cherrypicked-booksAllo-AVA
Allo-AVA: A Large-Scale Multimodal Dataset for Allocentric Avatar Animation
Overview
Allo-AVA (Allocentric Audio-Visual Avatar) is a large-scale multimodal dataset designed for research and development in avatar animation. It focuses on generating natural and contextually appropriate gestures from text and audio inputs in an allocentric (third-person) perspective. The dataset addresses the scarcity of high-quality, synchronized multimodal data capturing the intricate… See the full description on the dataset page: https://huggingface.co/datasets/SaifPunjwani/Allo-AVA.
