CoolFace
Datasetpublic

playwithmino/mandarin_avsr

Mandarin AVSR Synthetic Mandarin audio-visual speech separation mixes used by byd-avss. Each example is a mixture, a clean target waveform, and a silent grayscale mouth video of the target speaker. Built from Chinese-LiPS, AISHELL-6 Whisper, and WHAM noise. Use of this set must also follow the licenses of those source corpora. Download pip install -U "huggingface_hub[cli]" hf download playwithmino/mandarin_avsr --repo-type dataset --local-dir dataset/mandarin_avsr… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/mandarin_avsr.

sourceHugging Faceotherupdated 19d agoView on Hugging Face
0likes43downloads
filecv.tar.gz3.79 GBdownload
filepreview.tar.gz5.8 MBdownload
filetr_1mix.tar.gz7.86 GBdownload
filetr_2mix.tar.gz7.81 GBdownload
filetr_3mix.tar.gz7.73 GBdownload
filetr_4mix.tar.gz7.70 GBdownload
filetr_5mix.tar.gz7.64 GBdownload
filetr_all.tar.gz536 KBdownload
filett.tar.gz4.01 GBdownload

playwithmino/mandarin_avsr · main · files are served by the source, never re-hosted here