datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yubao_videos
YuBao: A New Chinese Dialect Speech Benchmark
Paper | Code
This repository contains the video metadata for the YuBao (語保) dataset, as presented in the paper "Towards Comprehensive Semantic Speech Embeddings for Chinese Dialects".
YuBao is a comprehensive collection from the Chinese Language Resources Protection Project, featuring speech, dialect transcripts, phonetic (IPA) transcriptions, and Mandarin translations for parallel items (1,000 characters, 1,200 words, and 50 sentences)… See the full description on the dataset page: https://huggingface.co/datasets/kalbin/yubao_videos.videos
