playwithmino/mandarin_avsr
Mandarin AVSR Synthetic Mandarin audio-visual speech separation mixes used by byd-avss. Each example is a mixture, a clean target waveform, and a silent grayscale mouth video of the target speaker. Built from Chinese-LiPS, AISHELL-6 Whisper, and WHAM noise. Use of this set must also follow the licenses of those source corpora. Download pip install -U "huggingface_hub[cli]" hf download playwithmino/mandarin_avsr --repo-type dataset --local-dir dataset/mandarin_avsr… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/mandarin_avsr.
043
