datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FalaBracarense_splitsdataset website: projectofalabracarense
Licence
CC - BY - NC - ND
Restrictions: Academic - Non Commercial Use, Attribution, No Derivatives
tw_parliament_split
Parliament
Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the
zh_tw split of disco-eth/WorldSpeech.
It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD
parliamentary proceedings.
This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style
text filtering pipeline. Audio is preserved from the upstream dataset and cast
as a Hugging Face Audio(sampling_rate=24000) feature.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/NickWeng/tw_parliament_split.japanese-anime-speech-v2-split-150k
japanese-anime-speech-v2-split-150k
joujiboi/japanese-anime-speech-v2 的 150,000 筆隨機子集,已切好 train / test。
split
rows
train
135,000
test
15,000
total
150,000
Columns
audio — 16 kHz mp3,與原始資料完全相同(未重新編碼)
sentence — 轉錄文字(原始欄位名為 transcription)
How it was built
來源的 sfw(271,788 筆)與 nsfw(20,849 筆)兩個 split 都有使用,並依原始比例分配名額
(sfw 139,313 / nsfw 10,687)。
在每個 split 的全部列上做無放回均勻抽樣,因此每筆資料被選中的機率相同。
抽出後整體打亂,前 15,000 筆為 test,其餘為 train,train / test… See the full description on the dataset page: https://huggingface.co/datasets/hhim8826/japanese-anime-speech-v2-split-150k.ghana-female-twi-speech-asr-8word-splits
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 8-Word Speech Segments
51139 speech-text pairs split from 30-min recordings.
Processing pipeline
Source audio from ghananlpcommunity/ghana-female-twi-tts-full-length
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghana-female-twi-speech-asr-8word-splits.ghana-female-twi-8sec-splits
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 8-Word Speech Segments
25951 speech-text pairs split from 30-min recordings.
Processing pipeline
Source audio from ghananlpcommunity/ghana-female-twi-tts-full-length
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-female-twi-8sec-splits.ghana-female-twi-asr-16word-splits
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi 16-Word Speech Segments
25951 speech-text pairs split from 30-min recordings.
Processing pipeline
Source audio from ghananlpcommunity/ghana-female-twi-tts-full-length
Full-file CTC forced alignment (MMS-300M) for… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/ghana-female-twi-asr-16word-splits.Clean_One_Speaker_STT_Split_EN_AR
Clean_One_Speaker_STT_Split_EN_AR
Mixed STT dataset — Saudi dialectal Arabic + Modern Standard Arabic + English —
with train/validation/test splits, built for fine-tuning multilingual ASR (e.g. Whisper)
without catastrophic forgetting of English.
Composition
Dialectal Arabic (~72%) — Sebssihakim/Clean_One_Speaker_SADA_Split
(cleaned single-speaker SADA; SDAIA / Saudi Broadcasting Authority, CC BY-NC-SA 4.0)
MSA (~8%) — FLEURS ar_eg, full official splits (CC-BY;… See the full description on the dataset page: https://huggingface.co/datasets/Sebssihakim/Clean_One_Speaker_STT_Split_EN_AR.japanese-anime-speech-v2-splitdataset split from joujiboi/japanese-anime-speech-v2
SpokenPortugueseGeographicalSocialVarieties_splitssentence splits from SpokenPortugueseGeographicalSocialVarieties generated via forced alignment
ArquivoDialetalCLUP_splitsdataset info: https://cl.up.pt/arquivo/
splits from ArquivoDialetalCLUP generated via forced alignment
license CC BY-NC-ND
eng-sports-radio-psst-iu-emotion-splits
English Sports Radio Non-Neutral Emotion IU Splits
Public non-neutral subset of NathanRoll/eng-sports-radio-psst-iu.
Each row is one intonation unit with exactly three columns:
audio: embedded 16 kHz mono audio for the IU
text: a leading emotion special token followed by the Parakeet transcript
accent: broadcast-location proxy accent label
Neutral examples were removed. The remaining rows are split by emotion:
joy: 247 rows, 0.287 audio hours
surprise: 153 rows, 0.188 audio… See the full description on the dataset page: https://huggingface.co/datasets/NathanRoll/eng-sports-radio-psst-iu-emotion-splits.dhivehi-audio-casts-1000-splitted_talks_en_mn_split
TED & TEDx Parallel Corpus (English-Mongolian)
The dataset is composed of two distinct subsets:
TED Talks (split en): English-language talks sourced from the official TED platform, paired with high-quality, human-generated Mongolian subtitles.
TEDxUlaanbaatar (split mn): Mongolian-language talks from local TEDx events in Ulaanbaatar, paired with the original Mongolian subtitles and machine-translated English subtitles.
This version of the dataset features segmented audio and text… See the full description on the dataset page: https://huggingface.co/datasets/bilguun/ted_talks_en_mn_split.split_sample
Silencio Network: Multilingual Accent Speech Dataset (Sample)
Overview
Silencio data is valuable because it’s collected in the wild from a massive, opt-in community (1.2M users across 180+ countries), giving buyers real-world accents, dialects, devices, and environments that lab or scraped datasets don’t capture. Every recording is tied to explicit, traceable consent and processed with privacy-first pipelines (GDPR/CCPA compliant, anonymized, PII hashed), which… See the full description on the dataset page: https://huggingface.co/datasets/jml2026/split_sample.Clean_One_Speaker_SADA_Split
Clean_One_Speaker_SADA_Split
Train/validation/test version of
Sebssihakim/Clean_One_Speaker_SADA,
a cleaned single-speaker subset of the SADA corpus (SDAIA / Saudi Broadcasting Authority).
Splits
80/10/10, stratified by speaker_dialect, seed 42:
train ~30.8k, validation ~3.9k, test ~3.9k rows.
The Maghrebi dialect (7 rows) was removed — too few samples to stratify.
Splits are segment-level: the same source show/speaker may appear in more than one split.
Rare… See the full description on the dataset page: https://huggingface.co/datasets/Sebssihakim/Clean_One_Speaker_SADA_Split.
