datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ark-asr-3b-open-asr-leaderboard-results
ARK-ASR-3B Open ASR Leaderboard Results
Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English
short-form hf-audio/open-asr-leaderboard splits.
These manifests were generated on a local 8x RTX 4090 machine and scored with
the shared Open ASR Leaderboard scorer:
PYTHONPATH=. python - <<'PY'
from normalizer.eval_utils import score_results
score_results(
'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official',
'AutoArk-AI/ARK-ASR-3B',
)
PY
Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.EGYSpeak
EGYSpeak
A curated dataset of 147,979 single-speaker Egyptian Arabic (pure dialect) audio clips with transcriptions, sourced from the fadisarwat/egyptian-arabic-lines Kaggle dataset and processed through a rigorous ASR pipeline.
Quick Start
1. Download the dataset:
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="MohamedGomaa30/EGYSpeak",
repo_type="dataset",
local_dir="EGYSpeak",
)
2. Extract the dataset:
from… See the full description on the dataset page: https://huggingface.co/datasets/MohamedGomaa30/EGYSpeak.pseudolabel-malaya-speech-stt-train-whisper-large-v3KazMix-3
KazMix-3
Kazakh three-speaker overlapping-speech dataset for target-speaker ASR (TS-ASR), released with the Persona-ASR project. Given a short enrollment utterance of a target speaker and a 3-speaker mixture, the task is to transcribe only the target speaker, or reject the utterance when the target is absent.
This repository ships the mixture manifests and generation scripts, not the audio. Mixtures are derived from the Kazakh Speech Dataset (KSD, OpenSLR 140); download KSD and… See the full description on the dataset page: https://huggingface.co/datasets/issai/KazMix-3.youtube-300h-movies-soniox
Persian Speech Corpus — full three-pool release
309.72 hours · 157,279 clips · 736 source videos · Soniox transcripts on every clip.
This is the complete quality-gated output of the persian-expressive-corpus
pipeline. It is organised into three mutually exclusive pools. Read the pool
column before using a clip — they are not interchangeable.
pool
clips
hours
transcript
emotion label
QC status
recommended use
A
66,108
108.34
yes
yes, 7-class
passed all gates
expressive… See the full description on the dataset page: https://huggingface.co/datasets/ErfanRou/youtube-300h-movies-soniox.test321
test321
This is a merged speech dataset containing 118 audio segments from 2 source datasets.
Dataset Information
Total Segments: 118
Speakers: 4
Languages: tr
Emotions: happy, angry, sad, neutral
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test321.ted-talks-in-chinese-zhongwen
TED中文 Podcast
聚焦华语地区的创意,本节目从上万个TED和TEDx演讲中,为您精选中文演讲,以及少量中文配音的经典英语演讲。演讲人包括科技和人文专家、关心当下与未来的思考者、关注挑战与探索的实践者。英雄不论出处,谁有创意谁讲。让这些演讲成为一把把钥匙,开启你的好奇心,升级你的行动力。
Focusing on creativity within the Chinese-speaking world, this program curates Chinese-language talks from tens of thousands of TED and TEDx presentations, along with a select few English classics dubbed into Chinese. Our speakers span technology and humanities experts, thinkers engaged with the present and future, and practitioners… See the full description on the dataset page: https://huggingface.co/datasets/bdx33/ted-talks-in-chinese-zhongwen.quantized-librispeech-train-360
