datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
korean-full-duplex-synthetic-dataset-preview
Korean Full-Duplex Synthetic Dataset Preview
Overview
Public preview of a Korean full-duplex synthetic speech dataset. This
repository contains 100 conversations sampled from a corpus of 89,273
conversations (2,000.5 hours); it does not publish the full corpus audio.
Preview contents
100 conversation WAV files
data/representative.jsonl
24 kHz, mono, 16-bit PCM
Events: normal, barge_in, backchannel, cutoff_by_user
Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.korean-asr
korean-asr — Korean ASR pseudo-labels for YODAS2
This repository contains transcripts and segment metadata only. It does not contain audio.
Every row points into espnet/yodas2 by
(shard, video_id, start, end), so you fetch the audio from the upstream dataset and cut it
yourself. See Reconstructing the audio.
split
utterances
hours
train
1,034,181
6,974.3
heldout
46,542
314.6
dev (subset of heldout)
3,000
20.4
Total labelled: 1,080,723 utterances / 7,288.9… See the full description on the dataset page: https://huggingface.co/datasets/nhatminh/korean-asr.
