CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MYJOKERML /chinese-dialogue-speech-dataset 中文多轮对话语音合成数据集 数据集概述 这是一个大规模的中文多轮对话语音合成数据集,包含 46,080 个多轮对话,涵盖文学问答、自然对话和诗词文化等多个领域。 数据统计 对话数量: 46,080 个 音频文件: 约 275,000 个 WAV 文件 音频总时长: 约 1,000-1,200 小时 音频格式: WAV, 16kHz 采样率 分批数量: 10 个压缩包 使用方法 1. 下载数据 from huggingface_hub import hf_hub_download import tarfile # 下载单个批次 batch_file = hf_hub_download( repo_id="MYJOKERML/chinese-dialogue-speech-dataset", filename="batch_001.tar.gz", repo_type="dataset" ) # 解压 with… See the full description on the dataset page: https://huggingface.co/datasets/MYJOKERML/chinese-dialogue-speech-dataset.text-to-speech100K<n<1M3 likes229 downloads11mo agoHugging Face02Nexdata-kr /500-Hours-Japanese-48khz-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset Description 일본어(Japan) 48kHz 풀 듀플렉스 자연 대화 스마트폰 음성 데이터셋으로, 주어진 주제를 바탕으로 자유롭게 대화하는 방식으로 수집되었습니다. 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 다양한 지역의 폭넓은 화자로부터 데이터를 수집하여 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다. 또한 다양한 AI 기업을 통해 데이터 품질을 검증했습니다. 데이터 수집, 저장 및 활용 전 과정에서 개인정보 보호 및 관련 법규를 엄격하게 준수하며, 사용자의 개인정보와 법적 권리를 보호합니다. 본 데이터셋은 GDPR, CCPA, PIPL을 준수합니다. 자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1971?source=hf.kr Specifications… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/500-Hours-Japanese-48khz-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset.audion<1K1 likes87 downloads14d agoHugging Face03xinyuzhou2000 /Towards-Joint-Modeling-of-Dialogue-Response-and-Speech-Synthesis-based-on-Large-Language-Modeltext10K<n<100K6 likes72 downloads3y agoHugging Face04yiwu2 /daily_dialogue_mixed_chinese_english_speech_ttsaudio10K<n<100K6 likes69 downloads2y agoHugging Face05Nexdata-kr /211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset Description 211시간 규모의 태국어(Thai-Thailand) 풀 듀플렉스 자연 대화 스마트폰 음성 데이터셋으로, 일반적인 주제를 기반으로 대화를 시뮬레이션하여 녹음했습니다. 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 총 654명의 태국어 원어민 화자로부터 데이터를 수집했으며, 다양한 화자와 지역적 특성을 반영하여 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다. 자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1594?source=hf.kr Format 16kHz, 16 bit, WAV, 모노 채널 Content category 일반적인 주제를 기반으로 대화를 시뮬레이션하여 녹음 Recording condition 낮은… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-dataset.audion<1K0 likes64 downloads14d agoHugging Face06Nexdata-AI /211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample Description 211 Hours - Thai(Thailand) Full-Duplex Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers(654 native speakers), geographicly speaking, enhancing model performance in real and complex tasks. For more details, please refer to the link:… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/211-Hours-Thai-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample.audion<1K0 likes43 downloads1mo agoHugging Face07Nexdata-AI /352-Hours-Urdu-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample Description Urdu Full-Duplex Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/352-Hours-Urdu-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample.audion<1K0 likes28 downloads1mo agoHugging Face08Nexdata-AI /500-Hours-Japanese-48khz-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample Description Japanese(Japan) 48khz Full-Duplex Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/500-Hours-Japanese-48khz-Full-Duplex-Spontaneous-Dialogue-Smartphone-speech-Data-Sample.textn<1K0 likes26 downloads1mo agoHugging Face09jlking /speechdialogue_dataaudio10K<n<100K0 likes22 downloads1y agoHugging Face10DataoceanAI /Ethiopian_Amharic_Free_Dialogue_Speech_Corpus SPECIFICATION: Product Type: Ethiopian Amharic language, free dialogue, mobile 16K 【Corpus Type】 Family, health, travel, education, work, cuisine, marriage, movies, music, socializing, celebrities, weather, sports, and other common topics of daily life. Natural context, applicable to all industries. Pronunciation Person Information: Gender: Male 50%, Female 50% Age: The pronunciation people mainly cover the age range of 16-45. Accent: The pronunciation people mainly come from… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Ethiopian_Amharic_Free_Dialogue_Speech_Corpus.0 likes11 downloads2y agoHugging Face11DataoceanAI /Albanian_Free_Dialogue_Speech_Corpus SPECIFICATION: Corpus Type: Family, health, travel, education, work, gourmet food, marriage, movies, music, socializing, celebrities, weather, sports, and other common topics of daily life. Natural context, applicable to all industries. Pronunciation Person Information: Gender: Male 45%, Female 55% Age: The pronunciation people mainly cover the age range of 16-45. Accent: Speakers are from Tirana. For more details:… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Albanian_Free_Dialogue_Speech_Corpus.0 likes4 downloads2y agoHugging Face12DataoceanAI /Free_dialogue_in_Odia_Speech_Corpus Product Type: Odia language from India, free conversation, mobile 16K For more details, please refer to the link: https://dataoceanai.com/datasets/asr/free-dialogue-in-odia-speech-corpus/ Specification: ID: King-ASR-946 Size: 52 hours Language: Odia Corpus Type: Home, health, travel, education, work, gourmet food, marriage, movies, music, socializing, celebrities, weather, sports, and other common topics in daily life Natural context, applicable to the entire industry… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Free_dialogue_in_Odia_Speech_Corpus.0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.