CoolFace
Datasetpublicgated

panlr/teochew_wild

Teochew-Wild:首个正字标注的野外潮州话数据集 本数据集(Teochew-Wild)是从网络上发音清晰、噪声较少的音视频内容中获取的,原始音视频的数据来源为:民生新闻、潮汕讲古、地方电视节目、故事书、抖音自媒体口播等,我借鉴了Emilla提出的数据集自动处理流水线,对原始数据进行归一化、降噪和剪切(部分自动剪切效果差的使用手工修正); Teochew-Wild总共包括20个发音标准、念错率低的潮汕母语说话人、共12500条音频片段,包含潮州市区、汕头市区、澄海、榕江音、潮安南部等多个区域的口音,语料内容覆盖书面用语与口头用语,并同时提供正字和拼音标注,是首个公开可用、标注准确率高的潮州话数据集,主要面向语音识别和语音合成任务。 文件说明 (File Structure Explanation) ├── label_for_qwen_asr/ # 预处理标签文件夹,完全适配Qwen-ASR模型读取格式 ├── README.md # 项目说明文档(本文档)… See the full description on the dataset page: https://huggingface.co/datasets/panlr/teochew_wild.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
44likes160downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
panlr/teochew_wild · CoolFace