datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seamless-interaction-jefferson-annotations
Seamless Interaction Jefferson-Style Annotations
An automatic, turn-oriented annotation layer for the
Meta Seamless Interaction Dataset.
It compares the dataset's traditional transcript with an ASR-derived
Jefferson-style condition and supplies speech-act, communicative-purpose,
interactional-signal, alignment, and quality fields.
This is a derived noncommercial research dataset. It does not redistribute
the source audio. Every record retains the original interaction ID, split… See the full description on the dataset page: https://huggingface.co/datasets/kennethli319/seamless-interaction-jefferson-annotations.enhanced-audiosnippets-long-2-8M
Enhanced Audiosnippets Long 2.8M
Enhanced version of mitermix/audiosnippets_long_2_8M with speech enhancement, emotion annotations, speaker embeddings, and comprehensive metadata analysis.
Dataset Summary
Metric
Value
Total samples
2,633,037
Total audio hours
4,932 h
Duration range
3.0s - 1124.3s
Mean duration
6.7s
Audio format
WAV, 48kHz mono
Tar files
1,410
Processing Pipeline
Each audio sample was processed through:
Speech… See the full description on the dataset page: https://huggingface.co/datasets/ai-music4you3/enhanced-audiosnippets-long-2-8M.ark-asr-3b-open-asr-leaderboard-results
ARK-ASR-3B Open ASR Leaderboard Results
Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English
short-form hf-audio/open-asr-leaderboard splits.
These manifests were generated on a local 8x RTX 4090 machine and scored with
the shared Open ASR Leaderboard scorer:
PYTHONPATH=. python - <<'PY'
from normalizer.eval_utils import score_results
score_results(
'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official',
'AutoArk-AI/ARK-ASR-3B',
)
PY
Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.egyptian-arabic-tts-diacritized
Egyptian Arabic TTS Corpus (Diacritized)
97,163 utterances / ~334 hours of Egyptian Arabic speech at 24 kHz, with
diacritized transcripts — the short vowels that Arabic script does not
write.
Why diacritics
Arabic is an abjad: short vowels are unwritten, so كتب may be kataba,
kutiba, or kutub. A TTS model with no Arabic pretraining cannot infer which,
and guesses — which native listeners hear as a foreign accent with constant
mispronunciation.
This is invisible to… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/egyptian-arabic-tts-diacritized.pseudolabel-malaya-speech-stt-train-whisper-large-v3KazMix-3
KazMix-3
Kazakh three-speaker overlapping-speech dataset for target-speaker ASR (TS-ASR), released with the Persona-ASR project. Given a short enrollment utterance of a target speaker and a 3-speaker mixture, the task is to transcribe only the target speaker, or reject the utterance when the target is absent.
This repository ships the mixture manifests and generation scripts, not the audio. Mixtures are derived from the Kazakh Speech Dataset (KSD, OpenSLR 140); download KSD and… See the full description on the dataset page: https://huggingface.co/datasets/issai/KazMix-3.youtube-300h-movies-soniox
Persian Speech Corpus — full three-pool release
309.72 hours · 157,279 clips · 736 source videos · Soniox transcripts on every clip.
This is the complete quality-gated output of the persian-expressive-corpus
pipeline. It is organised into three mutually exclusive pools. Read the pool
column before using a clip — they are not interchangeable.
pool
clips
hours
transcript
emotion label
QC status
recommended use
A
66,108
108.34
yes
yes, 7-class
passed all gates
expressive… See the full description on the dataset page: https://huggingface.co/datasets/ErfanRou/youtube-300h-movies-soniox.coral-v3-conversation-pnc-da
CoRal v3 Conversation PnC DA
RyeAI/coral-v3-conversation-pnc-da is a text-only Danish punctuation and
capitalization companion for the conversation training split of
CoRal-project/coral-v3.
It contains 102,226 restored transcript rows and no audio bytes.
Pinned companion revision: 6e4fbafde87fbffadd58bbe39a3a2e09e884351a.
SHA-256 of data/train-00000-of-00001.parquet:
79d30671f238e884a5b71b682bc811043e7df075566ce5565a236e66ffe32000.
Each row can be joined back to the gated source… See the full description on the dataset page: https://huggingface.co/datasets/RyeAI/coral-v3-conversation-pnc-da.az-asr-voa-305h
Labelling
field
value
label_origin
script
speech_register
broadcast
channel
wideband-16k
provenance
inferred
Editorial broadcast text aligned to VOA audio. Terminal punctuation at 77.4% (against 98.1% for LocalDoc) is consistent with segments cut from continuous broadcast rather than at sentence ends.
Adding this to a call model made it worse. Chinar-F8 v4 mixed in 147.8 h of it and strict WER on human-transcribed calls moved 44.23% -> 56.69%. Register, not… See the full description on the dataset page: https://huggingface.co/datasets/cillegio/az-asr-voa-305h.gigaspeech-part-3
Gigaspeech Part 3
This is Part 3 of 8 of a large-scale speech dataset, split to accommodate HuggingFace's repository size limits.
Multi-Part Dataset
This dataset is split across multiple repositories:
Part 1: shahdsaf/gigaspeech-part-1
Part 2: shahdsaf/gigaspeech-part-2
Part 3 (current): shahdsaf/gigaspeech-part-3
Part 4: shahdsaf/gigaspeech-part-4
Part 5: shahdsaf/gigaspeech-part-5
Part 6: shahdsaf/gigaspeech-part-6
Part 7: shahdsaf/gigaspeech-part-7
Part 8:… See the full description on the dataset page: https://huggingface.co/datasets/shahdsaf/gigaspeech-part-3.yt3_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/yt3_chunked_speech_restorised
Aligned dataset: instinct-org/yt3_chunked_speech_restorised_nfa_aligned
Rows: 506444 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised_nfa_aligned.
