CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ud-nlp /human-robot-conversation-korean Human-Robot Conversation Dataset (Korean) - 660+ Hours Dataset (Korean) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data Dataset characteristics: Characteristic Data Description Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-korean.audioautomatic-speech-recognitionn<1K1 likes51 downloads6mo agoHugging Face02ud-nlp /human-robot-conversation-english Human-Robot Conversation Dataset (English) - 660+ Hours Dataset (English) contains 660+ hours of audio featuring dialogues between AI and a human in English across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data Dataset characteristics: Characteristic Data Description Audio of dialogues between… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-english.audioautomatic-speech-recognitionn<1K1 likes30 downloads6mo agoHugging Face03ud-nlp /korean-speech-recognition Korean Speech Recognition Dataset - 10+ hours Dataset comprises 10 hours of high-quality telephone audio recordings in Korean, featuring 20 native speakers. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data Dataset characteristics: Characteristic Data Description Audio of… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/korean-speech-recognition.audioautomatic-speech-recognitionn<1K0 likes21 downloads10mo agoHugging Face04ud-nlp /vietnamese-speech-recognition Vietnamese Speech Dataset - 10+ hours Dataset comprises 10+ hours of telephone dialogues in Vietnamese, collected from 20 native speakers across various topics and domains. It is designed for research in speech recognition, focusing on various recognition models, primarily aimed at meeting the requirements for automatic speech recognition (ASR) systems. - Get the data Dataset characteristics: Characteristic Data Description Audio of telephone dialogues in… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/vietnamese-speech-recognition.audioautomatic-speech-recognitionn<1K0 likes16 downloads10mo agoHugging Face05ud-nlp /japanese-speech-recognition-dataset Japanese Telephone Dialogues Dataset - 10 Hours Dataset comprises 10 hours of high-quality telephone audio recordings in Japanese, featuring 20+ native speakers and achieving a 95% sentence accuracy rate. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data Dataset characteristics:… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/japanese-speech-recognition-dataset.audioautomatic-speech-recognitionn<1K0 likes14 downloads10mo agoHugging Face06ud-nlp /slovenian-speech-recognition Slovenian Speech Dataset - 10+ hours Dataset comprises 10+ hours of audio recordings featuring 20+ speakers engaged in telephone dialogues in the Slovenian language. It contains speech data designed for training robust language models and automatic speech recognition systems in real-world conversational scenarios. - Get the data Dataset characteristics: Characteristic Data Description Audio of telephone dialogues in Slovenian for training NLP models in… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/slovenian-speech-recognition.audioautomatic-speech-recognitionn<1K0 likes13 downloads10mo agoHugging Face07ud-nlp /human-robot-conversation-german Human-Robot Conversation Dataset (German) - 660+ Hours Dataset (German) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data Dataset characteristics: Characteristic Data Description Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-german.audioautomatic-speech-recognitionn<1K1 likes10 downloads6mo agoHugging Face08archivartaunik /ivan-bunin-sonechny-udar-uladzimir-ragautsou Сонечны ўдар Metadata Author: Іван Бунін Title: Сонечны ўдар Narrator: Уладзімір Рагаўцоў Source Group: Аўдыёкнігі Source: БЛР#аўдыякніга Notes The original audio files are preserved as-is: no conversion; no re-encoding; no filename changes inside each split folder, except removing one common top-level archive folder when present. To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders. Target maximum… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/ivan-bunin-sonechny-udar-uladzimir-ragautsou.audion<1K0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.