datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MLD-VC
🎥 MLD-VC: Multimodal Dataset for Video Conferencing
When AVSR Meets Video Conferencing: Dataset, Degradation, and the Hidden Mechanism Behind Performance Collapse (CVPR 2026)
📄 [Paper] | 🤗 [Hugging Face Dataset]
📌 Overview
MLD-VC is the first multimodal dataset specifically designed for Audio-Visual Speech Recognition (AVSR) in real-world video conferencing (VC) scenarios.
Unlike traditional AVSR datasets collected in controlled offline environments, MLD-VC… See the full description on the dataset page: https://huggingface.co/datasets/nccm2p2/MLD-VC.MLDSUM_NEW
