datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mTEDx-ptbr
Multilingual TEDx (Portuguese speech and transcripts)
NOTE: This dataset contains only the Portuguese portion of the mTEDx dataset, already processed and segmented into parts.
Multilingual TEDx (mTEDx) is a multilingual speech recognition and translation corpus to facilitate the training of ASR and SLT models in additional languages.
The corpus comprises audio recordings and transcripts from TEDx Talks in 8 languages (Spanish, French, Portuguese, Italian, Russian, Greek, Arabic… See the full description on the dataset page: https://huggingface.co/datasets/dominguesm/mTEDx-ptbr.kinyarwanda_afrivoice_all_domains_v0.1
Afrivoice Kinyarwanda — All Domains
Combined dataset across 5 domains from the original source.
Note: the source dataset also includes a scripted_education domain, excluded
here due to a cluster of corrupted audio files in one of its shards.
Attribution
Original dataset: DigitalUmuganda/Afrivoice_Kinyarwanda
License: CC-BY-4.0
Attribution: Digital Umuganda
This dataset is derived from the above source and released under the same CC-BY-4.0 license.… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.Public-Domain-Musickinyarwanda_afrivoice_all_domains_v0.2
Kinyarwanda AfriVoice — All Domains (v0.2)
Cleaned version of ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.1.
Changes from v0.1
Removed rows with empty/null transcription values across all splits (train/validation/test)
Audio and domain labels unchanged; only null-transcription rows were dropped
Source
Original data from DigitalUmuganda/Afrivoice_Kinyarwanda (CC-BY-4.0),
extracted and concatenated across 5 domains (agriculture, education… See the full description on the dataset page: https://huggingface.co/datasets/ElizabethMwangi/kinyarwanda_afrivoice_all_domains_v0.2.partial-spoof-cross-domain-audit-data
Partial-Spoof Cross-Domain Audit: detector score outputs
Per-utterance and per-frame score outputs of three detectors (MRM, BAM, CFPRF) on
PartialSpoof, LlamaPartialSpoof, PartialEdit, and HQ-MPSD. Together with the analysis code they reproduce the reported
results of the cross-domain operational audit.
An earlier version of this study was submitted to IJCB 2026; that submission was
withdrawn and was never published. These arrays support one manuscript, currently in
preparation… See the full description on the dataset page: https://huggingface.co/datasets/sukhdeveyash/partial-spoof-cross-domain-audit-data.public_domain_sounds_3secs
Public Domain Sounds
This is a backup of the 635 copyright-free sound recordings submitted to pdsounds.org before April 2009.
all files split into 3 seconds chunks
LICENSE NOTICE
pdsounds.org - all sounds archive
March 2, 2009
ALL 635 SOUNDS in this archive are entirely public domain and copyright free.
No rights are reserved.
The sounds were recorded by volunteers and uploaded to pdsounds.org with the
condition that henceforth they are license-free and public domain.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/public_domain_sounds_3secs.domains_asr_datasethealth_domain
Healthcare
Collected with the Collector platform for Ghanaian language data.
Tasks: Translation, Speech recognition
Languages: en → tw
Rows: 1
Audio: 0.1 minutes
Voices: female: 1
Privacy
Rows identify their speaker by an opaque code only — female_01, male_02, or
speaker_01 where no voice was stated. No names, ages, regions, phone numbers
or contact details are included. Codes are numbered per dataset, so codes in
two datasets from this platform cannot be… See the full description on the dataset page: https://huggingface.co/datasets/nasare34/health_domain.DomainSpeech
Multi-domain academic audio data for evaluating ASR model
Dataset Summary
This dataset, named "DomainSpeech," is meticulously curated to serve as a robust evaluation tool for Automatic Speech Recognition (ASR) models. Encompassing a broad spectrum of academic domains including Agriculture, Sciences, Engineering, and Business. A distinctive feature of this dataset is its deliberate design to present a more challenging benchmark by maintaining a technical terminology… See the full description on the dataset page: https://huggingface.co/datasets/AcaSp/DomainSpeech.youtube-data-various-domain
Dataset Card for "youtube-data-various-domain"
More Information needed
picovoice-wake-word-benchmark
Picovoice Wake Word Benchmark Dataset
This dataset contains a collection of wake word recordings used for benchmarking wake word detection systems. The dataset has been reformatted from the original Picovoice Wake Word Benchmark repository for easier use with Hugging Face's ecosystem.
Dataset Description
The dataset contains over 300 recordings of six different wake words from more than 50 distinct speakers. These recordings were originally used to benchmark different… See the full description on the dataset page: https://huggingface.co/datasets/domdomegg/picovoice-wake-word-benchmark.DomesticEnvironmentSoundEventDetection_DESED-PublicEvalllm-lingodomains_asr_datasetdomingosDominguesdomains_asr_datasetdomingoCadudomingospunjabi-dataset-testDominguinhosAll
Dataset Card for "All"
More Information needed
joao-frangoVOCAL-APENAS2domiroZWACK
Dataset Card for "ZWACK"
More Information needed
All-2
Dataset Card for "All-2"
More Information needed
vertical-domainsuladzimir-arlou-dom-z-damavikami
Дом з дамавікамі
Metadata
Author: Уладзімір Арлоў
Title: Дом з дамавікамі
Narrator:
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size: about 250… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/uladzimir-arlou-dom-z-damavikami.
