datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chinese-LiPS
Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides
⭐ Introduction
The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios.
🚀 Dataset Details
Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.Lipi-Ghor-bn-882-SSTT
🗣️ Lipi-Ghor | লিপিঘর — Bengali Speech Dataset (bn-882-SSTT)
Lipi-Ghor (লিপিঘর, meaning "House of Scripts") is a large-scale Bengali speech dataset designed for automatic speech recognition (ASR), speaker diarization, and spoken language research. It is one of the largest open Bengali speech corpora with aligned speaker, transcription, and timestamp annotations.
Built by Team_Villagers as part of DL Sprint 4.0.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/Sanjidh090/Lipi-Ghor-bn-882-SSTT.chinese-lips-speech-slide-probe
Chinese-LiPS Speech + Slide Probe
A self-contained probe set for testing whether visual slide context helps
simultaneous speech translation — with the input as audio, not transcripts.
Why audio matters: feeding a transcript to a text LLM deletes the acoustic
ambiguity (homophones, polysemy) that slide context is meant to resolve; the
transcript already commits to one reading. Any honest test of "does vision help
streaming ST" must consume speech.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-speech-slide-probe.Multimodal_Greek_Sign_Language_and_Lip_Reading
Multimodal_Greek_Sign_Language_and_Lip_Reading (v1)
The Multimodal_Greek_Sign_Language_and_Lip_Reading (v1) is a comprehensive dataset designed for research and development in multimodal machine learning, speech recognition, vision recognition, sign language recognition, sign language translation and accessibility technologies.
Description
The dkourem/Multimodal_Greek_Sign_Language_and_Lip_Reading-v1 is a comprehensive dataset designed for research and development… See the full description on the dataset page: https://huggingface.co/datasets/dkourem/Multimodal_Greek_Sign_Language_and_Lip_Reading.chinese-lips-longform-debug
Chinese-LiPS Long-Form (zh long streaming speech)
Reconstructed continuous long-speech streams from
BAAI/Chinese-LiPS, for
slide-aware / streaming speech-translation development and evaluation. Each
source video (one speaker, one scripted lecture with slides) was released as
pre-segmented clips; here they are re-joined into the full talk.
Two variants of the same 3 talks (~97 min speech total):
config
how segments are placed
use
orig_timeline
at their original session… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-longform-debug.Lipi-Ghor-bn-882-SSTT
🗣️ Lipi-Ghor | লিপিঘর — Bengali Speech Dataset (bn-882-SSTT)
Lipi-Ghor (লিপিঘর, meaning "House of Scripts") is a large-scale Bengali speech dataset designed for automatic speech recognition (ASR), speaker diarization, and spoken language research. It is one of the largest open Bengali speech corpora with aligned speaker, transcription, and timestamp annotations.Built by Team_Villagers as part of DL Sprint 4.0.
Dataset Details
Dataset Description
Lipi-Ghor… See the full description on the dataset page: https://huggingface.co/datasets/teamvillagers/Lipi-Ghor-bn-882-SSTT.WildVid-LIP
WildVid-LIP: In-The-Wild Temporal Anchors for Visual Speech Recognition
WildVid-LIP is a large-scale, open-source dataset mapping over 100,000 curated temporal segments from unconstrained, real-world YouTube videos. It provides precise timestamp anchors optimized for training Visual Speech Recognition (VSR / Lip-Reading), audio-visual synchronization, and multimodal self-supervised models.
Instead of distributing heavy, monolithic video files—which introduces platform friction… See the full description on the dataset page: https://huggingface.co/datasets/Rizul2159/WildVid-LIP.Lipi-Ghor-bn-882-SSTT
🗣️ Lipi-Ghor | লিপিঘর — Bengali Speech Dataset (bn-882-SSTT)
Lipi-Ghor (লিপিঘর, meaning "House of Scripts") is a large-scale Bengali speech dataset designed for automatic speech recognition (ASR), speaker diarization, and spoken language research. It is one of the largest open Bengali speech corpora with aligned speaker, transcription, and timestamp annotations.Built by Team_Villagers as part of DL Sprint 4.0.
Dataset Details
Dataset Description
Lipi-Ghor… See the full description on the dataset page: https://huggingface.co/datasets/frovolts/Lipi-Ghor-bn-882-SSTT.
