datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Quran-Ayah-Corpus
Quran-Ayah-Corpus: A Multi-Reciter Arabic Quranic Speech Dataset
Dataset Description:
Ayah-Corpus is a large-scale, multi-reciter Arabic speech dataset meticulously curated for Automatic Speech Recognition (ASR) tasks. It consists of high-quality audio recordings of Quranic verses (Ayahs) paired with their corresponding exact transcriptions. The audio is sourced from two primary repositories: Al-Quran.cloud and EveryAyah.com.
This dataset is specifically designed to… See the full description on the dataset page: https://huggingface.co/datasets/rabah2026/Quran-Ayah-Corpus.multilingual-accent-speech
🎙️ Silencio Network: Voice AI Sample Dataset
📊 This is a sample. The full Silencio corpus contains 100,000+ hours across 170+ countries and 100+ languages.
📧 Contact: sofia@silencioai.com for custom datasets, bulk licensing, or specific language requests.
🌍 Why Silencio Data?
Silencio data is collected in the wild from a massive, opt-in community (2M+ contributors across 180+ countries), giving you:
✅ Real-world accents, dialects, devices, and… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/multilingual-accent-speech.french_tv_media_dataset_2026
Dataset Card : A Multi-Domain Pseudo-Labeled ASR Corpus
Résumé (Abstract)
Ce corpus présente un jeu de données de reconnaissance automatique de la parole (ASR) en langue française, totalisant 97 heures d'audio annoté. Il est dérivé de flux de diffusion (broadcast) issus de France Télévisions, couvrant une diversité de domaines acoustiques et linguistiques (Information, Société, Divertissement, Documentaire, Sport). L'annotation a été réalisée via une méthodologie… See the full description on the dataset page: https://huggingface.co/datasets/Archime/french_tv_media_dataset_2026.ToneWebinars
ToneWebinars
darija_yt_2026
darija_yt_2026
Partition upload generated automatically.
Namespace: ohsn
Repo: ohsn/darija_yt_2026
Video count: 3511
Duration hours: 1565.31
This dataset contains raw audio files and per-video metadata generated from the Darija YouTube extraction pipeline.
french_tv_media_dataset_2026
Dataset Card : A Multi-Domain Pseudo-Labeled ASR Corpus
Résumé (Abstract)
Ce corpus présente un jeu de données de reconnaissance automatique de la parole (ASR) en langue française, totalisant 97 heures d'audio annoté. Il est dérivé de flux de diffusion (broadcast) issus de France Télévisions, couvrant une diversité de domaines acoustiques et linguistiques (Information, Société, Divertissement, Documentaire, Sport). L'annotation a été réalisée via une méthodologie… See the full description on the dataset page: https://huggingface.co/datasets/madoss/french_tv_media_dataset_2026.sinhala-tts-dataset-archive-20260429-082457
Sinhala TTS Dataset
Clean, segmented Sinhala speech from the "Unlimited History" YouTube series by @sunchare.
Stats
Metric
Value
Utterances
218
Train
208
Val
10
Hours
0.51
Mean duration
8.5s
Sample rate
22050 Hz
Pipeline
Raw YouTube audio -> HTDemucs -> VoiceFixer + DeepFilterNet3 ->
Diarization -> Silero-VAD -> ASR (faster-whisper: C:\Users\kosal\sinhala-tts\whisper-small-si-ct2) -> Quality filtering (SNR>=20.0dB)
Format… See the full description on the dataset page: https://huggingface.co/datasets/outlawmold/sinhala-tts-dataset-archive-20260429-082457.complete-voiceai-speech-dataset-anonymized
🎙️ Silencio Network: Voice AI Sample Dataset
📊 This is a sample. The full Silencio corpus contains 100,000+ hours across 170+ countries and 100+ languages.
📧 Contact: sofia@silencioai.com for custom datasets, bulk licensing, or specific language requests.
🌍 Why Silencio Data?
Silencio data is collected in the wild from a massive, opt-in community (2M+ contributors across 180+ countries), giving you:
✅ Real-world accents, dialects, devices, and… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/complete-voiceai-speech-dataset-anonymized.audio-mia-batch-20260312
Audio MIA Batch 20260312
This dataset contains 6,998 audio files (64 GB) downloaded from YouTube videos.
Dataset Structure
Each row contains:
audio: Audio bytes (playable in the dataset viewer)
video_id: YouTube video ID
category: Content category
source_term: Search term used
query: Full search query
title: Video title
url: YouTube URL
uploader: Channel name
channel_id: YouTube channel ID
upload_date: Upload date (YYYY-MM-DD)
duration: Video duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/potsawee/audio-mia-batch-20260312.bharatvani-hindi-showcase
BharatVani Hindi Speech Corpus • Public Interactive Showcase
150-Hour Enterprise Devanagari Hindi Speech Corpus & Precomputed Latents
Curated & Mastered by BharatVani AI • TheCreatorOS
1. Interactive Dataset Preview
This repository is the official public evaluation showcase for the 150-Hour BharatVani Hindi Speech Corpus (103,784 Studio Clips).
Use the Dataset Viewer above to play real audio clips, inspect the word-level timestamp alignments, and… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-showcase.reviewed_sample
Silencio Network: Multilingual Accent Speech Dataset (Sample)
Overview
Silencio data is valuable because it’s collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don’t capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed), which… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/reviewed_sample.ToneSlavic
ToneSlavic
sound-wave-leyu-ethiopian-multilingual-speech-corpus-2026
🇪🇹 Leyu Ethiopian Multilingual Speech Corpus 2026
Official research-grade multilingual speech dataset submitted for the Leyu Platform Data Collection Competition 2026, organized by gheero (Leyu Platform Team).
🏢 Platform & Team Verification Details
Team Name: SoundWaveET
Hugging Face Organization: SoundWaveET
Dataset Repository: SoundWaveET/leyu-ethiopian-multilingual-speech-corpus-2026
Platform Live Frontend: https://leyusound.netlify.app
Platform Live API:… See the full description on the dataset page: https://huggingface.co/datasets/LeyuCompetition/sound-wave-leyu-ethiopian-multilingual-speech-corpus-2026.oc_voice_2026_word_timings
oc_voice_2026_word_timings
Turn-level audio/text dataset for target-speaker phone-call style utterances, with ElevenLabs Scribe V2 word-level timestamps.
Format
One row is one target-speaker/agent turn.
Columns:
conversation_id: string
turn_index: int32, zero-based within the filtered target-speaker conversation
target_speaker: string, always agent
text: training text
asr_text: fresh ElevenLabs transcript for this final audio
text_source: matched_substring or… See the full description on the dataset page: https://huggingface.co/datasets/Cyleux/oc_voice_2026_word_timings.multilingual-accent-speech-v2
Silencio Network: Multilingual Accent Speech Dataset (Sample)
Overview
Silencio data is valuable because it's collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don't capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed)… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/multilingual-accent-speech-v2.betrac-2026-with-meta
betrac-2026-with-meta
Annotated speech dataset created with Whissle Annotator — a multimodal annotation pipeline for speech, NLP, and visual analysis.
Source Dataset
This dataset is derived from the following HuggingFace dataset(s):
BeTraC/betrac-2026:0:50
BeTraC/betrac-2026:50:50
BeTraC/betrac-2026:100:50
BeTraC/betrac-2026:150:50
BeTraC/betrac-2026:200:50
BeTraC/betrac-2026:250:50
BeTraC/betrac-2026:300:50
BeTraC/betrac-2026:350:50
BeTraC/betrac-2026:400:50… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/betrac-2026-with-meta.ToneSpeak
ToneSpeak
yoruba-second-sbpn-demucs-20260826
yoruba-second-sbpn-demucs-20260826
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the existing same-speaker… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/yoruba-second-sbpn-demucs-20260826.mac-m4pro-fresh-diarization-demucs-20260902
gdrive-sbpn-fresh-diarization-demucs-mac-m4pro-20260902
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/mac-m4pro-fresh-diarization-demucs-20260902.2026-dwesui-g04-admedvoice
DWESUI 2026 - Grupa 4 - ADMEDVOICE (medyczna)
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny.
Zespol (atrybucja): Grupa 4 (DWESUI 2026)
Zrodlo oryginalne: https://huggingface.co/datasets/TCA/zwesui-grupa-4-medyczn
Domena: medyczna
Licencja zrodla: nagrania YouTube CC-BY + Kaggle ADMEDVOICE + TTS
Status: kopia publiczna w organizacji kursowej (zespół… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g04-admedvoice.leyu-oromo-speech-corpus-2026
Leyu Afaan Oromo Speech Corpus 2026
Official speech dataset submission for the Leyu Data Collection Competition 2026.
Organization & Team
Hugging Face Org: SoundWaveET
Dataset Repo: SoundWaveET/leyu-oromo-speech-corpus-2026
corpus5-t2inserts-sample-200-20260916IALP-2026-data
IALP-2026: Whisper Open-Set Data-Selection — Query / Dev / Test Sets
Supporting data for the study "Whisper-Based Open-Set Data Selection for NSC
Adaptation." This repository holds the fixed target-query, validation, and
evaluation sets used across all experiments. Each part is a self-contained
.tar.gz.
All audio is 16 kHz mono. Each split ships with:
audio/ — audio files (FLAC, except GigaSpeech which is WAV PCM_16)
wav.scp — <utt_id> audio/<file> (Kaldi-style, relative paths)… See the full description on the dataset page: https://huggingface.co/datasets/pengyizhou/IALP-2026-data.leyu-amharic-speech-corpus-2026
🇪🇹 Leyu Ethiopian Multilingual Speech Corpus 2026
Official research-grade multilingual speech dataset submitted for the Leyu Platform Data Collection Competition 2026, organized by gheero (Leyu Platform Team).
🏢 Platform & Team Verification Details
Team Name: SoundWaveET
Hugging Face Organization: SoundWaveET
Dataset Repository: SoundWaveET/leyu-ethiopian-multilingual-speech-corpus-2026
Platform Live Frontend: https://leyusound.netlify.app
Platform Live API:… See the full description on the dataset page: https://huggingface.co/datasets/SoundWaveET/leyu-amharic-speech-corpus-2026.asr-merged-fr-20260814-144411
Test_Merged_datasets_3
Dataset ASR fusionné : audio + texte matérialisés en Parquet sur le Hub
(colonnes audio, text, source_dataset).
Nombre d'exemples poussés : 346.
Sources
BrunoHays/Accueil_UBS (config=default, split=test, audio=audio, text=sentence, max_samples=500)
eustlb/french-long-form-test (config=default, split=test, audio=audio, text=sentence, max_samples=500)
Normalisation texte
Mode : training_v3 (training_v3 | legacy | none)… See the full description on the dataset page: https://huggingface.co/datasets/Zeldeo/asr-merged-fr-20260814-144411.mac-m4pro-confirmation-3drive-20260903
sbpn-mac-m4pro-confirmation-three-drive-20260902
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the existing… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/mac-m4pro-confirmation-3drive-20260903.ToneBooks
ToneBooks
pi_bench
pi-bench
pi-bench is a multi-task audio benchmark prepared for public hosting and evaluation reproducibility. The repository is organized as a Hugging Face dataset with one dataset config per task file under data/, so each benchmark subset is visible and loadable independently.
Overview
The current release contains 11 task-specific configs spanning three broad categories:
Counterfactual/contextual QA (CTC_*)
Clarification-seeking QA (Trivia_Clarification_*… See the full description on the dataset page: https://huggingface.co/datasets/neurips2026-pi-bench/pi_bench.2026-dwesui-g01-neurologia
DWESUI 2026 - Grupa 1 - neurologia
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny.
Zespol (atrybucja): Grupa 1 (DWESUI 2026)
Zrodlo oryginalne: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia
Domena: neurologia
Licencja zrodla: nagrania YouTube CC-BY + synteza TTS
Status: kopia publiczna w organizacji kursowej (zespół opublikował zbiór… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g01-neurologia.yoruba-datasold-sbpn-demucs-20260825
yoruba-datasold-sbpn-demucs-20260825
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved using the existing same-speaker… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/yoruba-datasold-sbpn-demucs-20260825.
