datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
clipquill-asr-benchmark
Measuring whisper-tiny vs whisper-base in a browser tab
Word error rate, wall-clock timing, transfer size and peak memory for two
quantised Whisper tiers running entirely client-side in a real Chrome window,
with the scripts that produced every number.
If you are building an in-browser transcription page, the two results worth
knowing before you pick a model tier:
On clean synthetic audio the two tiers tie. If that is all you test, you
will conclude the tier does not matter… See the full description on the dataset page: https://huggingface.co/datasets/sophia8888/clipquill-asr-benchmark.yap-movie-clips
Yap Movie Clips
164,882 short video clips of film dialogue from 366 films, one
sentence per clip, in 12 languages. Every clip comes with the sentence as
written in the film's official subtitles, word-level timings, the surrounding
subtitle cues, an independent speech-to-text transcript of the same audio, and
the phoneme sequence the sentence was expected to contain versus what a
phoneme recogniser actually heard.
This is the corpus behind the listening and pronunciation cards at… See the full description on the dataset page: https://huggingface.co/datasets/anchpop/yap-movie-clips.
