datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zeroth-korean
Zeroth-Korean
Zeroth-Korean
The data set contains transcriebed audio data for Korean. There are 51.6 hours transcribed Korean audio for training data (22,263 utterances, 105 people, 3000 sentences) and 1.2 hours transcribed Korean audio for testing data (457 utterances, 10 people). This corpus also contains pre-trained/designed language model, lexicon and morpheme-based segmenter(morfessor).
Zeroth project introduces free Korean speech corpus and aims to make Korean… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/zeroth-korean.korean_datasetZeroth-STT-Korean
Zeroth-STT-Korean Dataset
Description
This is a shuffled version of the Zeroth-STT-Ko dataset.
Citation
Zeroth-Korean Dataset, created by [Lucas Jo(@Atlas Guide Inc.) and Wonkyum Lee(@Gridspace Inc.)], 2023.
Available at https://github.com/goodatlas/zeroth under CC-BY-4.0 license.
Junhoee/STT_Korean_Dataset_80000 Dataset, created by [Junhoee], 2024.
Available at https://huggingface.co/datasets/Junhoee/STT_Korean_Dataset_80000
korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.Korean-Japanese-Code-Switching-Speech
Korean-Japanese-Code-Switching-Speech
This dataset contains Korean-Japanese code-switching speech recordings with sentence-level transcriptions. It was introduced in the paper Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs.
Since there is an extremely small amount of Korean-Japanese code-switching data available, this dataset was designed to be used as a small-scale evaluation dataset.
The dataset consists of code-switching recordings… See the full description on the dataset page: https://huggingface.co/datasets/thetaone-ai/Korean-Japanese-Code-Switching-Speech.korean-asr
korean-asr — Korean ASR pseudo-labels for YODAS2
This repository contains transcripts and segment metadata only. It does not contain audio.
Every row points into espnet/yodas2 by
(shard, video_id, start, end), so you fetch the audio from the upstream dataset and cut it
yourself. See Reconstructing the audio.
split
utterances
hours
train
1,034,181
6,974.3
heldout
46,542
314.6
dev (subset of heldout)
3,000
20.4
Total labelled: 1,080,723 utterances / 7,288.9… See the full description on the dataset page: https://huggingface.co/datasets/nhatminh/korean-asr.human-robot-conversation-korean
Human-Robot Dataset
The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the Korean language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in robotic systems and conversational AI technologies.… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-korean.human-robot-conversation-korean
Human-Robot Conversation Dataset (Korean) - 660+ Hours
Dataset (Korean) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-korean.Korean-Speech-Dataset
🎧 Korean Speech Dataset
The Korean Speech Dataset is a large-scale speech audio dataset designed to provide high-quality and structured audio data for advanced AI and machine learning systems. It includes 192 hours of audio data across 628 files, delivered in MP3 and WAV formats, with a total size of 447 MB. This well-balanced audio dataset ensures diverse and representative voice data, with 52% female and 48% male speakers, and an age distribution ranging from 18 to 50+ years. The… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Korean-Speech-Dataset.korean-speech-recognition
Korean Speech Dataset
Dataset comprises 10+ hours of audio recordings from 20+ speakers, featuring telephone-quality speech data from native korean speakers. It provides a diverse collection of spoken language for automatic speech recognition tasks and serves as essential training data for model training in NLP and speech detection research.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in automatic speech recognition… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/korean-speech-recognition.korean-speech-recognition
Korean Speech Recognition Dataset - 10+ hours
Dataset comprises 10 hours of high-quality telephone audio recordings in Korean, featuring 20 native speakers. Designed for advancing speech recognition models and language processing, this extensive speech data corpus covers diverse topics and domains, making it ideal for training robust automatic speech recognition (ASR) systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/korean-speech-recognition.YodaLingua-Korean
YodaLingua-Korean
YodaLingua is a high-quality speech dataset designed for training text-to-speech (TTS) systems, ASR models, and any application requiring clean, well-aligned audio–text pairs.This release contains the Korean portion of the multilingual YodaLingua collection.
🧾 Dataset Overview
Property
Value
Total clips
53,903 audio–transcription pairs
Total duration
143 hours
Speakers
2,219 distinct speakers
Audio format
MP3 • mono • 24 kHz • 16-bit… See the full description on the dataset page: https://huggingface.co/datasets/Thomcles/YodaLingua-Korean.korean-speech-samples
Korean Speech Samples
This sample shows Korean contributor speech in a consistent audio format. It is meant to help buyers review recording quality, language coverage, and metadata structure before scoping a larger delivery.
What This Shows
Korean speech recordings from contributor collection workflows
Clip-level metadata for format and review context
Ground-truth transcripts for understanding sample content
Dataset Specifications
Field… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/korean-speech-samples.
