datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spoken-multiturn-sft
Spoken Multi-turn SFT Japanese
Japanese spoken multi-turn SFT dataset generated from kanhatakeyama/AutoMultiTurnByCalm3-22B using CosyVoice2 TTS.
Dataset Description
This dataset contains Japanese multi-turn SFT (Supervised Fine-Tuning) data with spoken questions.
q1: First question (text + audio)
a1: First answer (text only)
q2: Follow-up question (text + audio)
a2: Second answer (text only)
Samples
ID
Q1
Q1 Audio
A1
Q2
Q2 Audio
A2
0
鉄は強磁性体ですか?… See the full description on the dataset page: https://huggingface.co/datasets/Atotti/spoken-multiturn-sft.multiturn_ks
khursanirevo/multiturn_ks
Dataset Description
Multiturn dialogue dataset with speaker-separated stereo audio and multi-language transcripts from 139 YouTube videos.
Features
Audio: Stereo audio with speaker separation (speaker 0 = left channel, speaker 1 = right channel)
Segments: Speaker turn-level annotations with timestamps for English and Malay
Multi-language: Transcripts in 9 languages (en, ms, zh-Hans, zh-Hant, ru, id, ar, ja, ko)
Video ID: YouTube video… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/multiturn_ks.Malaysian-Multiturn-Chat-Assistant
Malaysian-Multiturn-Chat-Assistant
Generate synthetic multi-turn chat assistant with complex system prompt using mesolitica/Malaysian-Qwen2.5-72B-Instruct.
After that generate synthetic voice using mesolitica/Malaysian-Dia-1.6B also verified with Force Alignment to make sure the pronunciations almost correct.
A conversation must at least have 2 audio. We follow chat template from Qwen/Qwen2-Audio-7B-Instruct.
how to prepare the dataset
huggingface-cli download \… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Multiturn-Chat-Assistant.Malaysian-UltraChat-Speech-Multiturn-Instructions
Malaysian-UltraChat-Speech-Multiturn-Instructions
We filter Malaysian short user questions from mesolitica/malaysian-ultrachat that suitable to convert to voice prompt and generate synthetic voice using mesolitica/Malaysian-Dia-1.6B also verified with Force Alignment to make sure the pronunciations almost correct.
A conversation must at least have 2 audio. We follow chat template from Qwen/Qwen2-Audio-7B-Instruct.
how to prepare the dataset
huggingface-cli download \… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-UltraChat-Speech-Multiturn-Instructions.malay-multiturn-dialogueswb-fisher-casper-multiturn-sft-v4multiturn-speech-eval
