datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SEA-Spoof
SEA-Spoof
Access And License
Access requires author approval. Please email the authors before requesting or using the dataset:
wu_jinyang@a-star.edu.sg
imcc.sg@gmail.com
This dataset is released for non-commercial academic research only.
Use is restricted to academic institutions and approved research users. Commercial use is not permitted, and this dataset may not be used by commercial companies or for commercial products, services, model training, evaluation… See the full description on the dataset page: https://huggingface.co/datasets/Jack-ppkdczgx/SEA-Spoof.Neapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/michaelcacioli/Neapolitan-Spoken-Corpus.LIEPA-3_read-spon_16kHz
LIEPA-3 Bendrinis Garsynas
Didysis lietuvių kalbos garsynas (LIEPA-3) – atviras kalbos duomenų rinkinys, skirtas šnekos atpažinimo tikslams ir moksliniams tyrimams.
Aprašymas
Šis duomenų rinkinys apima LIEPA-3 garsyno read ir spon dalis (fonetinė phon ir dialektų dial dalys neįtrauktos). Duomenys paruošti Whisper mokymui.
Pakeitimų istorija
1 versija — 2026-08-25 (20260825)
Pradinis pilno garsyno įkėlimas: 6,729,976 įrašai (be P)
Kalbėtojų… See the full description on the dataset page: https://huggingface.co/datasets/liepa-project/LIEPA-3_read-spon_16kHz.The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden
Dataset Card for "The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden"
Homepage: Spoken Wikipedia CorporaRepository: Bitbucket RepositoryPaper: Publication at LREC 2016Leaderboard: Interspeech 2018 PaperPoint of Contact: nats@nats.gitlab.io
Dataset Summary
The Spoken Wikipedia Corpora (SWC) is a collection of aligned spoken Wikipedia articles including articles in Dutch. It includes approximately 210 hours of audio, transcriptions, and metadata. The corpus is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Jaspernl/The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden.ghomala-spoken-bible
Ghomálá' Spoken New Testament — aligned audio + trilingual text
Part of the Lingo / NativeAI language-preservation project. This is
~20 hours of spoken Ghomálá' (Ghomala, ISO bbj; a Grassfields Bantu language of
West Cameroon) — recorded readings of the New Testament — aligned chapter-by-chapter
with parallel text in Ghomálá', French, and English.
Spoken-language data is exactly what oral-first Cameroonian languages lack, which makes
this a rare resource for building ASR, TTS… See the full description on the dataset page: https://huggingface.co/datasets/flagship-ai/ghomala-spoken-bible.SpokenWikipedia-retrieval
Spoken Wikipedia speech-text retrieval (MTEB)
Volunteer readings of Wikipedia articles paired with the article lead, in Dutch,
English, German, Spanish and French.
Recordings come from Wikimedia Commons, which is free by site policy, and the lead
text from each Wikipedia, which is CC-BY-SA. The set is published as cc-by-sa-4.0.
Only the first 60 seconds of each reading is kept, since readers start
at the lead. One recording per article.
Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/SpokenWikipedia-retrieval.SpokenWords-GA-EN-MTed
Dataset Card for Dataset Name
This is the Irish portion of the Spoken Words dataset (available at MLCommons/ml_spoken_words),
with merged splits “train”, “validation”, and “test”, augmented with machine translation.
The Irish sentences are automatically translated into English using Google Translation API.
The dataset includes approximately 3 hours and 2 minutes of audio (03:02:02), spoken by multiple narrators.
Dataset Structure
Dataset({
features: ['keyword'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/SpokenWords-GA-EN-MTed.SpokenPortugueseGeographicalSocialVarieties
Spoken Portuguese - Geographical and Social Varieties
dataset source: https://www.clul.ulisboa.pt
(1995-1997 - European Commission DGXXII, Programme LINGUA/SOCRATES)
The project is concluded and the materials are published in CD-ROM, with the exclusive publishing support of Instituto Camões, under the title Português Falado - Documentos Autênticos: Gravações áudio com transcrição alinhada. Its distribution outside of Portugal is ensured by Instituto Camões and in Portugal by CLUL.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/SpokenPortugueseGeographicalSocialVarieties.sundanese-spontaneous-conversation-sample
Sundanese Spontaneous Conversation — Free Sample
Unscripted two-speaker Sundanese (su-ID) conversation recorded in
West Java, Indonesia. This is a free 30-minute sample of a larger
commercially licensable corpus.
Why this exists
Sundanese has approximately 40 million native speakers, yet
spontaneous conversational data is almost absent:
Resource
Type
Volume
Common Voice Spontaneous 4.0
Spontaneous, 2 speakers
0.62 h
OpenSLR SLR36 / SLR44
Read speech
—… See the full description on the dataset page: https://huggingface.co/datasets/Indospeech/sundanese-spontaneous-conversation-sample.SpokenPortugueseGeographicalSocialVarieties_splitssentence splits from SpokenPortugueseGeographicalSocialVarieties generated via forced alignment
common_voice_spontaneous_speech_4_0_de
Mozilla Common Voice Spontaneous Speech 4.0 - German (Parquet & IPA)
Clean Parquet format of Mozilla Spontaneous Speech 4.0 German (264 samples).
eng-sports-radio-psst-iu-emotion-splits
English Sports Radio Non-Neutral Emotion IU Splits
Public non-neutral subset of NathanRoll/eng-sports-radio-psst-iu.
Each row is one intonation unit with exactly three columns:
audio: embedded 16 kHz mono audio for the IU
text: a leading emotion special token followed by the Parakeet transcript
accent: broadcast-location proxy accent label
Neutral examples were removed. The remaining rows are split by emotion:
joy: 247 rows, 0.287 audio hours
surprise: 153 rows, 0.188 audio… See the full description on the dataset page: https://huggingface.co/datasets/NathanRoll/eng-sports-radio-psst-iu-emotion-splits.parakeet-tdt-blind-spots
Blind Spots of nvidia/parakeet-tdt-0.6b-v2
This dataset documents 14 systematically identified blind spots in NVIDIA's parakeet-tdt-0.6b-v2 automatic speech recognition model. The errors span 8 distinct categories and reveal a consistent pattern: the model struggles with inputs outside the distribution of its Western English-centric training data.
Model Under Test
Property
Value
Model
nvidia/parakeet-tdt-0.6b-v2
Parameters
600M
Architecture… See the full description on the dataset page: https://huggingface.co/datasets/TieIncred/parakeet-tdt-blind-spots.LIEPA-3_read-spon-dial_16kHz
LIEPA-3 Bendrinis Garsynas
Didysis lietuvių kalbos garsynas (LIEPA-3) – atviras kalbos duomenų rinkinys, skirtas šnekos atpažinimo tikslams ir moksliniams tyrimams.
Aprašymas
Šis duomenų rinkinys apima LIEPA-3 garsyno read, spon ir dial dalis (fonetinė phon dalis neįtraukta). Duomenys paruošti Whisper mokymui.
Dėmesio: ši garsyno versija yra sumažinto diskretizavimo dažnio (16 kHz). Originalus garsynas:… See the full description on the dataset page: https://huggingface.co/datasets/liepa-project/LIEPA-3_read-spon-dial_16kHz.2026-zwesui-g02-sport-probka
ZWESUI 2026 - Grupa 2 - sport - probka
Publiczna próbka ilustrująca domenę zbioru testowego ASR zbudowanego przez Grupę 2
(ZWESUI 2026) w ramach przedmiotu Warsztaty z ewaluacji systemów rozpoznawania mowy
(UAM WMI). Pełny zbiór jest utrzymywany jako held-out w organizacji
uam-wmi-asr-eval-labs — dostęp na prośbę
(HF @michaljunczyk, micjun@amu.edu.pl).
Pełny zbiór (held-out, prywatny): uam-wmi-asr-eval-labs/2026-zwesui-g02-sport
Oryginał zespołu: niepubliczny (konto zespołu)… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-zwesui-g02-sport-probka.
