botp/WordVoice-5A
WordVoice-5A Dataset 🚀 A Large-Scale Bilingual Word-level Five-Annotation Dataset for WordVoice 📖 Dataset Description / 数据集简介 WordVoice-5A is a large-scale bilingual (Mandarin and English) dataset containing approximately 4.7k hours of speech with fine-grained word-level acoustic annotations, designed for high-precision controllable Text-to-Speech (TTS). It addresses the scarcity of large-scale, high-quality word-aligned datasets with explicit acoustic… See the full description on the dataset page: https://huggingface.co/datasets/botp/WordVoice-5A.
0140
