datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vietnamese-acoustic-boundary-verifier-data
Vietnamese Acoustic Boundary & Speaker Purity Dataset (Gemini 3.8 Flash Distilled)
This dataset contains 202 curated Vietnamese audio samples with fine-grained acoustic boundary annotations distilled from Google Gemini 3.8 Flash (thinkingLevel="LOW").
It is specifically designed to train and evaluate multimodal models (e.g., Gemma 4 E4B Audio) on acoustic quality control for speech synthesis and speaker diarization pipelines.
Dataset Structure
Each sample is… See the full description on the dataset page: https://huggingface.co/datasets/tungnguyenlam/vietnamese-acoustic-boundary-verifier-data.vox1-veri-full
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
1211
1211
40
# of samples
133777
14865
4874
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
kinyarwanda_cleaned_testset_verified_20HRSkinyarwanda_cleaned_testset_verified_200HRSSpeaker_Verificationvox2-veri-full
VoxCeleb 2
VoxCeleb2 contains over 1 million utterances for 6,112 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
5,994
5,994
118
# of samples
982,808
109,201
36,237
Data Fields
ID (string): The ID of the sample with format <spk_id--utt_id_start_stop>.
duration (float64): The duration of the segment in seconds.
wav (string): The filepath of the waveform.
start (int64): The… See the full description on the dataset page: https://huggingface.co/datasets/yangwang825/vox2-veri-full.vox2-veri-3s
VoxCeleb 2
VoxCeleb2 contains over 1 million utterances for 6,112 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
5,994
5,994
118
# of samples
982,808
109,201
36,237
Data Fields
ID (string): The ID of the sample with format <spk_id--utt_id_start_stop>.
duration (float64): The duration of the segment in seconds.
wav (string): The filepath of the waveform.
start (int64): The… See the full description on the dataset page: https://huggingface.co/datasets/yangwang825/vox2-veri-3s.amharic_cleaned_testset_verifiedthuyg20-tts-clean-verifiedkinyarwanda_cleaned_testset_verifiedhavacilik-veriseti
ATC Veri Kümesi - Whisper Modeli ile İnce Ayar
Bu veri kümesi, OpenAI'nin Whisper modelini, Hava Trafik Kontrolü (ATC) iletişimlerinde transkripsiyon doğruluğunu artırmak amacıyla ince ayar yapmak için oluşturulmuştur. Veri kümesi, iki ana kaynaktan alınan transkripsiyonlar ve karşılık gelen ses dosyalarını içermektedir: ATCO2 ve UWB-ATCC korpusu, özellikle havacılıkla ilgili iletişimler için seçilmiştir. Veri kümesi, Otomatik Konuşma Tanıma (ASR) projelerinde kullanılmak üzere… See the full description on the dataset page: https://huggingface.co/datasets/mehmedadymn/havacilik-veriseti.gaia-verified
GAIA-Verified
A corrected subset of GAIA's validation split: 147 of the original 165 tasks.
27 of GAIA's 165 validation tasks (16.4%) are defective. 18 could not be
salvaged and were removed; 9 had a gold answer that is simply wrong and were corrected.
Questions were never rewritten and the scorer was never patched.
Defective
27 / 165 (16.4%)
Removed
18
Golds corrected
9
GAIA-Verified
147 tasks
Every verdict, with its evidence, is in the audit:… See the full description on the dataset page: https://huggingface.co/datasets/anivenu/gaia-verified.kinyarwanda_cleaned_testset_verified_100HRSThai-Voice-Test-Verification
Thanarit/Thai-Voice
Combined Thai audio dataset from multiple sources
Dataset Details
Total samples: 10
Total duration: 0.01 hours
Language: Thai (th)
Audio format: 16kHz mono WAV
Volume normalization: -20dB
Sources
Processed 1 datasets in streaming mode
Source Datasets
GigaSpeech2: Large-scale multilingual speech corpus
Usage
from datasets import load_dataset
# Load with streaming to avoid downloading everything
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Thanarit/Thai-Voice-Test-Verification.vox1-veri-3s
VoxCeleb 1
VoxCeleb1 contains over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.
Verification Split
train
validation
test
# of speakers
1211
1211
40
# of samples
299246
33672
4874
References
https://www.robots.ox.ac.uk/~vgg/data/voxceleb/vox1.html
kyrgyz-asr-verified-v1🇰🇬 Kyrgyz ASR-Verified Speech Corpus
metric
value
Clips
411,834
Audio
874 hours
Speakers
280
Verification CER
mean 0.79%, median 0.00%, p90 2.35%
License & access
KSTU members only. This dataset is released under a custom KSTU Internal Dataset License (other, see the LICENSE file):
🎓 Access and use are restricted to KSTU (Kyrgyz State Technical University) members and KSTU-affiliated persons.
🚫 Requests from outside KSTU will not be approved. Request… See the full description on the dataset page: https://huggingface.co/datasets/kstunlp/kyrgyz-asr-verified-v1.livestreamkinyarwanda_cleaned_testset_verified_10HRSses_verisiwolof-kallaama-external-verificationVeriSpeak
VeriSpeak
VeriSpeak is a spoken-statement factual-verification benchmark. Each example is a short
synthesized speech clip of a single declarative sentence about a public figure, labeled
correct or incorrect depending on whether the spoken statement is factually true.
The task: given the audio (and optionally its transcript), decide whether the claim it
makes is accurate. It targets speech-native fact-checking / hallucination detection.
Dataset at a glance… See the full description on the dataset page: https://huggingface.co/datasets/abhiram4572/VeriSpeak.speaker-overlap-verifiedsonyc_ust_verifiedwolof-google-fleurs-external-verificationspeaker_verification_dataset_collectioniaaa-speaker-verification
