datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
youtube-300h-movies-soniox
Persian Speech Corpus — full three-pool release
309.72 hours · 157,279 clips · 736 source videos · Soniox transcripts on every clip.
This is the complete quality-gated output of the persian-expressive-corpus
pipeline. It is organised into three mutually exclusive pools. Read the pool
column before using a clip — they are not interchangeable.
pool
clips
hours
transcript
emotion label
QC status
recommended use
A
66,108
108.34
yes
yes, 7-class
passed all gates
expressive… See the full description on the dataset page: https://huggingface.co/datasets/ErfanRou/youtube-300h-movies-soniox.persian-youtube-330h-soniox
persian-youtube-330h-soniox
Persian conversational / voice-search / meeting ASR fine-tuning set built from
YouTube audio, labelled by Soniox stt-async-v5. Machine-labelled. Not ground
truth. Built 2026-09-22 by ytcrawl/build_dataset.py.
What this is
74,538 training segments (331.5 h) and 7,442 validation
segments (32.9 h), 16 kHz mono FLAC, 5–28 s each, cut from
2,708 videos on 106 channels across 9 domains.
Validation is held out by channel (a creator is never on… See the full description on the dataset page: https://huggingface.co/datasets/ErfanRou/persian-youtube-330h-soniox.
