datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mosel
Dataset Description, Collection, and Source
The MOSEL corpus is a multilingual dataset collection including up to 950K hours of open-source speech recordings covering the 24 official languages of the European Union. We collect data by surveying labeled and unlabeled speech corpora under open-source compliant licenses.
In particular, MOSEL includes the automatic transcripts of 441k hours of unlabeled speech from VoxPopuli and LibriLight. The data is transcribed using Whisper large… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/mosel.arabic-audio-collection-mostafa-mahmoud
Mostafa Mahmoud Arabic Speech Dataset
Dataset Summary
The Mostafa Mahmoud Arabic Speech Dataset is a large-scale Arabic speech corpus containing approximately 187 hours of speech recordings and corresponding transcripts derived from publicly available lectures, interviews, television appearances, and talks by Dr. Mostafa Mahmoud.
The dataset was created to support Arabic speech technology research and development, including:
Automatic Speech Recognition (ASR)… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-mostafa-mahmoud.sada-validation-preprocessed
Details
This is the SADA 2022 dataset with the input_features whish are log mels and the cleaned_labels which is the tokenized version of the cleaned_text. You can directly use this as the validation dataset when training Whisper Tiny, Small, Base & Medium models, as they all use the same tokenizer. Please double check this as well from the original model repo.
In addtition, the following filters were applied to this data:
All audios are less than 30 seconds and greater than 0… See the full description on the dataset page: https://huggingface.co/datasets/mosama/sada-validation-preprocessed.bulgarian-parliament-moss-1to1
Bulgarian Parliament MOSS 1-to-1
Private machine-labelled training candidates under construction. Not gold data
or a held-out evaluation set. Public redistribution terms remain unverified.
Source: DimitarV/eurospeech-bg-single-speaker at
6aa43432891867bcd249c0fe944ac6286888231b, derived from
disco-eth/EuroSpeech.
Reference transcripts are parliamentary stenographic text. Speaker identities
are inferred clusters, not named or human-verified speakers.
Active batch list… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-parliament-moss-1to1.30_hours_krio_voicemoss-emolia-elise-hq-captioned
MOSS · Emolia + Elise + Inline-Bursts — HQ, captioned
A high-quality, richly captioned slice of the MOSS-local voice-acting corpus: expressive speech
clips scored by a panel of acoustic detectors, filtered to the top by a composite reward, and captioned
in the voice-acting format (a "how the voice sounds / how to perform it" description plus the script
with inline vocal-burst tags). Audio is shipped both as flac (WebDataset tars) and as pre-computed
MOSS-Audio-Tokenizer codes… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-emolia-elise-hq-captioned.switchboard-tierb-codeswitch
SwitchBoard Tier B — African code-switched speech
87 consented utterances of intra-sentential code-switching — Nigerian Pidgin,
Yorùbá, Hausa and Kiswahili each mixed with English inside a single sentence —
recorded from 8 bilingual volunteers at the Deep Learning Indaba 2026, Lagos.
Collected for the MLC (Africa) × Intron Agentic Voice AI Challenge as an
evaluation set for telco/fintech voice agents. 8.75 minutes total.
What this is for
Measuring whether a speech… See the full description on the dataset page: https://huggingface.co/datasets/mosesdaudu/switchboard-tierb-codeswitch.moshi-tool-audio
Moshi Tool-Calling — Audio-Grounded Dataset
Audio-grounded data teaching Moshi / PersonaPlex to emit tool-call special
tokens in its inner monologue when it hears a request — and to stay quiet
otherwise (listening/idle frames are trained to PAD).
Each row is a code tensor codes[17, T] at 12.5 Hz:
rows
stream
content
0
text monologue
PAD while listening/idle, `<
1:9
Moshi audio
silence
9:17
user audio
the spoken question (edge-tts), Mimi-encoded
mask=1 marks… See the full description on the dataset page: https://huggingface.co/datasets/abrarfahim/moshi-tool-audio.mostafa-mahmoud
Mostafa Mahmoud — Cleaned Egyptian Arabic (Qwen3-TTS ready)
A cleaned, single-speaker, speech-only, 24 kHz version of
oddadmix/arabic-audio-collection-mostafa-mahmoud
(Dr. Mostafa Mahmoud — Egyptian Arabic), prepared for fine-tuning
Qwen3-TTS.
Speaker
mostafa_mahmoud (single)
Dialect
Egyptian Arabic
Clips
24,444
Duration
~106.7 h
Sample rate
24 kHz, mono
Clip length
1–30 s
Schema
column
type
value
audio
Audio(24000)
cleaned… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/mostafa-mahmoud.arabic-audio-collection-mostafa-mahmoud
Mostafa Mahmoud Arabic Speech Dataset
Dataset Summary
The Mostafa Mahmoud Arabic Speech Dataset is a large-scale Arabic speech corpus containing approximately 187 hours of speech recordings and corresponding transcripts derived from publicly available lectures, interviews, television appearances, and talks by Dr. Mostafa Mahmoud.
The dataset was created to support Arabic speech technology research and development, including:
Automatic Speech Recognition (ASR)… See the full description on the dataset page: https://huggingface.co/datasets/EYang27/arabic-audio-collection-mostafa-mahmoud.mosla
Overview
The MOSLA dataset ("MOSLA") is a longitudinal, multimodal, multilingual, and controlled dataset created by inviting participants to learn one
of three target languages (Arabic, Spanish, and Chinese) from scratch over a span of two years, exclusively through online instruction,
and recording every lesson using Zoom. The dataset is semi-automatically annotated with speaker/language IDs and transcripts by both human
annotators and fine-tuned state-of-the-art speech models.… See the full description on the dataset page: https://huggingface.co/datasets/octanove/mosla.
