datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Neapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/michaelcacioli/Neapolitan-Spoken-Corpus.The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden
Dataset Card for "The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden"
Homepage: Spoken Wikipedia CorporaRepository: Bitbucket RepositoryPaper: Publication at LREC 2016Leaderboard: Interspeech 2018 PaperPoint of Contact: nats@nats.gitlab.io
Dataset Summary
The Spoken Wikipedia Corpora (SWC) is a collection of aligned spoken Wikipedia articles including articles in Dutch. It includes approximately 210 hours of audio, transcriptions, and metadata. The corpus is licensed… See the full description on the dataset page: https://huggingface.co/datasets/Jaspernl/The_Spoken_Wikipedia_Corpora_Dutch_ASR_Hiidden.ghomala-spoken-bible
Ghomálá' Spoken New Testament — aligned audio + trilingual text
Part of the Lingo / NativeAI language-preservation project. This is
~20 hours of spoken Ghomálá' (Ghomala, ISO bbj; a Grassfields Bantu language of
West Cameroon) — recorded readings of the New Testament — aligned chapter-by-chapter
with parallel text in Ghomálá', French, and English.
Spoken-language data is exactly what oral-first Cameroonian languages lack, which makes
this a rare resource for building ASR, TTS… See the full description on the dataset page: https://huggingface.co/datasets/flagship-ai/ghomala-spoken-bible.SpokenWikipedia-retrieval
Spoken Wikipedia speech-text retrieval (MTEB)
Volunteer readings of Wikipedia articles paired with the article lead, in Dutch,
English, German, Spanish and French.
Recordings come from Wikimedia Commons, which is free by site policy, and the lead
text from each Wikipedia, which is CC-BY-SA. The set is published as cc-by-sa-4.0.
Only the first 60 seconds of each reading is kept, since readers start
at the lead. One recording per article.
Built by… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/SpokenWikipedia-retrieval.SpokenWords-GA-EN-MTed
Dataset Card for Dataset Name
This is the Irish portion of the Spoken Words dataset (available at MLCommons/ml_spoken_words),
with merged splits “train”, “validation”, and “test”, augmented with machine translation.
The Irish sentences are automatically translated into English using Google Translation API.
The dataset includes approximately 3 hours and 2 minutes of audio (03:02:02), spoken by multiple narrators.
Dataset Structure
Dataset({
features: ['keyword'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/SpokenWords-GA-EN-MTed.SpokenPortugueseGeographicalSocialVarieties
Spoken Portuguese - Geographical and Social Varieties
dataset source: https://www.clul.ulisboa.pt
(1995-1997 - European Commission DGXXII, Programme LINGUA/SOCRATES)
The project is concluded and the materials are published in CD-ROM, with the exclusive publishing support of Instituto Camões, under the title Português Falado - Documentos Autênticos: Gravações áudio com transcrição alinhada. Its distribution outside of Portugal is ensured by Instituto Camões and in Portugal by CLUL.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/SpokenPortugueseGeographicalSocialVarieties.SpokenPortugueseGeographicalSocialVarieties_splitssentence splits from SpokenPortugueseGeographicalSocialVarieties generated via forced alignment
tajik-spoken-instructions
Tajik Spoken Instructions (Q&A)
Synthetic Tajik speech of dictionary and language-exercise questions, each paired with its
written answer — spoken instruction in, text answer out.
648,842 clips · ~540 hours · 2 voices (male + female)
Questions cover word meanings, antonyms, etymology, usage and grammar
Numbers expanded to spoken Tajik
Deduplicated; Latin-script and mis-encoded rows removed
Contents
file
what
audio_*.tar
the wav files
manifest.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Tohirju/tajik-spoken-instructions.tajik-spoken-questions-omni
Tajik Spoken Questions — Omni-a3
35,068 clips · 32.62 hours · Tajik (tg) · 24 kHz mono
⚠️ This audio is SYNTHETIC (TTS-generated), not real speech
Every question clip here was generated by a text-to-speech model, not recorded from a human.
It is deliberately published as a separate repo from Tohirju/tajik-audio and
Tohirju/tajik-asr-full-data, which contain real recorded speech — do not mix them without
knowing which is which. Do not use this set to train or… See the full description on the dataset page: https://huggingface.co/datasets/Tohirju/tajik-spoken-questions-omni.
