CoolFace
Datasetpublic

Wi-Fi/korean-full-duplex-synthetic-dataset-preview

Korean Full-Duplex Synthetic Dataset Preview Overview Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio. Preview contents 100 conversation WAV files data/representative.jsonl 24 kHz, mono, 16-bit PCM Events: normal, barge_in, backchannel, cutoff_by_user Annotation format… See the full description on the dataset page: https://huggingface.co/datasets/Wi-Fi/korean-full-duplex-synthetic-dataset-preview.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
1likes144downloads
Dataset Card

Korean Full-Duplex Synthetic Dataset Preview

Overview

Public preview of a Korean full-duplex synthetic speech dataset. This repository contains 100 conversations sampled from a corpus of 89,273 conversations (2,000.5 hours); it does not publish the full corpus audio.

Preview contents

  • —100 conversation WAV files
  • —data/representative.jsonl
  • —24 kHz, mono, 16-bit PCM
  • —Events: normal, barge_in, backchannel, cutoff_by_user

Annotation format

Each conversation includes:

  • —conversation_id: conversation identifier
  • —events[].speaker: user or assistant
  • —events[].text: utterance text
  • —events[].start_sec, events[].end_sec: dialogue-timeline timestamps
  • —events[].event_type: normal, barge_in, backchannel, or cutoff_by_user
  • —label_sequence: round-level interaction labels

Full-corpus statistics

  • —Text dialogues: 729,635
  • —Text utterances: 10,694,922
  • —Materialized conversations: 89,273
  • —Materialized duration: 2,000.5 hours

The 2,000.5-hour duration is the measured sum of the materialized conversation WAV durations.

Generation, license, and attribution

Contact

  • —mukss0909056@snu.ac.kr