spoken
Datasets
All datasets matching “spoken”duplexgen-spoken
DuplexGen Spoken
Rendered spoken audio for the DuplexGen turn-taking dialogues — the exact
set of clips used to fine-tune the full-duplex model (PP-DG) in
DuplexGen: Adaptive Synthesis of Human–AI Turn-Taking Dialogues.
Each clip is a full render of one generated dialogue variation: the mixed
two-speaker dialogue audio, the isolated per-turn utterances, and the inserted
backchannel clips, plus per-clip metadata. Audio is synthesized with
Chatterbox TTS; the dialogue
text it… See the full description on the dataset page: https://huggingface.co/datasets/DuplexGen/duplexgen-spoken.ml_spoken_wordsMultilingual Spoken Words Corpus is a large and growing audio dataset of spoken
words in 50 languages collectively spoken by over 5 billion people, for academic
research and commercial applications in keyword spotting and spoken term search,
licensed under CC-BY 4.0. The dataset contains more than 340,000 keywords,
totaling 23.4 million 1-second spoken examples (over 6,000 hours). The dataset
has many use cases, ranging from voice-enabled consumer devices to call center
automation. This dataset is generated by applying forced alignment on crowd-sourced sentence-level
audio to produce per-word timing estimates for extraction.
All alignments are included in the dataset.SpokenWOZ-Test-Audiospoken-squad-t2aspoken-multiturn-sft
Spoken Multi-turn SFT Japanese
Japanese spoken multi-turn SFT dataset generated from kanhatakeyama/AutoMultiTurnByCalm3-22B using CosyVoice2 TTS.
Dataset Description
This dataset contains Japanese multi-turn SFT (Supervised Fine-Tuning) data with spoken questions.
q1: First question (text + audio)
a1: First answer (text only)
q2: Follow-up question (text + audio)
a2: Second answer (text only)
Samples
ID
Q1
Q1 Audio
A1
Q2
Q2 Audio
A2
0
鉄は強磁性体ですか?… See the full description on the dataset page: https://huggingface.co/datasets/Atotti/spoken-multiturn-sft.spoken-alpaca-gpt4
