CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nccm2p2 /AiEdit 🎧 AiEdit Dataset 📖 Introduction AiEdit is a large-scale, cross-lingual speech editing dataset designed to advance research and evaluation in Speech Editing tasks. We have constructed an automated data generation pipeline comprising the following core modules: Text Engine: Powered by Large Language Models (LLMs), this engine intelligently processes raw text to execute three types of editing operations: Addition, Deletion, and Modification. Speech Synthesis & Editing:… See the full description on the dataset page: https://huggingface.co/datasets/nccm2p2/AiEdit.audio10K<n<100K0 likes197 downloads7mo agoHugging Face02JunXueTech /AiEdit 🎧 AiEdit Dataset 📖 Introduction AiEdit is a large-scale, cross-lingual speech editing dataset designed to advance research and evaluation in Speech Editing tasks. We have constructed an automated data generation pipeline comprising the following core modules: Text Engine: Powered by Large Language Models (LLMs), this engine intelligently processes raw text to execute three types of editing operations: Addition, Deletion, and Modification. Speech Synthesis & Editing:… See the full description on the dataset page: https://huggingface.co/datasets/JunXueTech/AiEdit.audio10K<n<100K3 likes157 downloads8mo agoHugging Face03K-University-AIED /Pathological-child-voice Speech Dataset for AI-Based Language Assessment in Children The "Speech Database of Typically Developing and Speech-Impaired Children" is an open speech dataset designed to support the development of AI-based language assessment systems. It contains speech samples from children aged 2 to 9 who are either typically developing or have reduced consonant articulation accuracy. This dataset is based on standardized Korean articulation tools: APAC (Articulation and Phonology… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/Pathological-child-voice.audion<1K0 likes46 downloads1mo agoHugging Face04K-University-AIED /hallym_AI_OpenDataset Hallym Adult and Child Speech Dataset This dataset contains speech recordings and transcriptions collected from adult and child speakers for AI-based speech and language research. Dataset Overview Total Records: 2,714 Speakers: 49 (adult: 25, child: 24) Groups: adult, child File Format: WAV (audio) + TXT (transcription) Speaker Statistics Group Count Gender Age Range Adult 25명 남/여 50~78세 Child 24명 남/여 3~8세 Dataset Fields… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/hallym_AI_OpenDataset.audio1K<n<10K0 likes14 downloads7mo agoHugging Face05K-University-AIED /korean_monosyllabic_speech Korean Monosyllabic Speech Perception Test Dataset 「Korean Monosyllabic Speech Perception Test Database」set is an open speech dataset created for evaluating monosyllables (meaningless, meaningful) and researching error patterns in elderly individuals with mild to moderate hearing loss. This dataset selected only monosyllables with a correct response rate of 80~100% out of 3192 possible Korean consonant-vowel combination sounds. The dataset is distributed under the CC BY NC ND 4.0… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/korean_monosyllabic_speech.audion<1K0 likes13 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.