datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
youtube-300h-movies-soniox
Persian Speech Corpus — full three-pool release
309.72 hours · 157,279 clips · 736 source videos · Soniox transcripts on every clip.
This is the complete quality-gated output of the persian-expressive-corpus
pipeline. It is organised into three mutually exclusive pools. Read the pool
column before using a clip — they are not interchangeable.
pool
clips
hours
transcript
emotion label
QC status
recommended use
A
66,108
108.34
yes
yes, 7-class
passed all gates
expressive… See the full description on the dataset page: https://huggingface.co/datasets/ErfanRou/youtube-300h-movies-soniox.yap-movie-clips
Yap Movie Clips
164,882 short video clips of film dialogue from 366 films, one
sentence per clip, in 12 languages. Every clip comes with the sentence as
written in the film's official subtitles, word-level timings, the surrounding
subtitle cues, an independent speech-to-text transcript of the same audio, and
the phoneme sequence the sentence was expected to contain versus what a
phoneme recogniser actually heard.
This is the corpus behind the listening and pronunciation cards at… See the full description on the dataset page: https://huggingface.co/datasets/anchpop/yap-movie-clips.
