avsr
Datasets
All datasets matching “avsr”AVSR-CT-AIavsr-tr-datasetavsr-recordingsmandarin_avsr
Mandarin AVSR
Synthetic Mandarin audio-visual speech separation mixes used by byd-avss. Each example is a mixture, a clean target waveform, and a silent grayscale mouth video of the target speaker.
Built from Chinese-LiPS, AISHELL-6 Whisper, and WHAM noise. Use of this set must also follow the licenses of those source corpora.
Download
pip install -U "huggingface_hub[cli]"
hf download playwithmino/mandarin_avsr --repo-type dataset --local-dir dataset/mandarin_avsr… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/mandarin_avsr.avsr-leaderboard
Japanese AVSR Leaderboard
This is an AVSR leaderboard that evaluates AVSR/ASR models using an internally collected out-of-domain evaluation dataset for AVSR benchmarking.
Evaluation Dataset
We randomly sampled sentences from the JSUT corpus, had about five speakers read them aloud while simultaneously recording their faces, and collected 660 audio-visual samples that passed manual quality checks.
Metrics
The model performance is evaluated using CER (Character… See the full description on the dataset page: https://huggingface.co/datasets/enactic/avsr-leaderboard.AVSR-Vietnamese-Dataset
