speech-dialogue
chinese-dialogue-speech-dataset
中文多轮对话语音合成数据集
数据集概述
这是一个大规模的中文多轮对话语音合成数据集,包含 46,080 个多轮对话,涵盖文学问答、自然对话和诗词文化等多个领域。
数据统计
对话数量: 46,080 个
音频文件: 约 275,000 个 WAV 文件
音频总时长: 约 1,000-1,200 小时
音频格式: WAV, 16kHz 采样率
分批数量: 10 个压缩包
使用方法
1. 下载数据
from huggingface_hub import hf_hub_download
import tarfile
# 下载单个批次
batch_file = hf_hub_download(
repo_id="MYJOKERML/chinese-dialogue-speech-dataset",
filename="batch_001.tar.gz",
repo_type="dataset"
)
# 解压
with… See the full description on the dataset page: https://huggingface.co/datasets/MYJOKERML/chinese-dialogue-speech-dataset.500-Hours-Japanese-48khz-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset
Description
일본어(Japan) 48kHz 풀 듀플렉스 자연 대화 스마트폰 음성 데이터셋으로, 주어진 주제를 바탕으로 자유롭게 대화하는 방식으로 수집되었습니다. 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 다양한 지역의 폭넓은 화자로부터 데이터를 수집하여 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다. 또한 다양한 AI 기업을 통해 데이터 품질을 검증했습니다.
데이터 수집, 저장 및 활용 전 과정에서 개인정보 보호 및 관련 법규를 엄격하게 준수하며, 사용자의 개인정보와 법적 권리를 보호합니다. 본 데이터셋은 GDPR, CCPA, PIPL을 준수합니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1971?source=hf.kr
Specifications… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/500-Hours-Japanese-48khz-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset.Towards-Joint-Modeling-of-Dialogue-Response-and-Speech-Synthesis-based-on-Large-Language-Modeldaily_dialogue_mixed_chinese_english_speech_tts211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset
Description
211시간 규모의 태국어(Thai-Thailand) 풀 듀플렉스 자연 대화 스마트폰 음성 데이터셋으로, 일반적인 주제를 기반으로 대화를 시뮬레이션하여 녹음했습니다. 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 총 654명의 태국어 원어민 화자로부터 데이터를 수집했으며, 다양한 화자와 지역적 특성을 반영하여 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1594?source=hf.kr
Format
16kHz, 16 bit, WAV, 모노 채널
Content category
일반적인 주제를 기반으로 대화를 시뮬레이션하여 녹음
Recording condition
낮은… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset.211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample
Description
211 Hours - Thai(Thailand) Full-Duplex Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers(654 native speakers), geographicly speaking, enhancing model performance in real and complex tasks.
For more details, please refer to the link:… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample.
