datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lipreadingconfusable-100-lipreading
Confusable-100: an English vocabulary built to break lip reading
A small, deliberately adversarial corpus for probing one specific failure mode of
visual speech recognition: consonants articulated by the tongue leave no distinctive
trace on the lips, so words differing only in those consonants are not separable from
video in principle, not merely in practice.
This is an evaluation probe, not a training set. It is one speaker and 3.4 minutes
of speech. Its purpose is to expose… See the full description on the dataset page: https://huggingface.co/datasets/diddmstjr/confusable-100-lipreading.lipreading-wordslip-reading_data3Lip_reading_Speech_Video_Corpus
SPECIFICATION:
This dataset covers 250 individuals, with each person recording no less than 600 short sentences, and the effective video duration for each individual is half an hour, which can be used for tasks such as face recognition and object detection.
For more details:https://dataoceanai.com/datasets/cv/lip-speech-video-was-collected-for-250-people/
ID:
King-VD-018
