z050209/va-spoken-qa-agentvoice-stage2
Stage-2 spoken-QA training corpus (gemma4_talker) 97,697 train / 935 val spoken QA pairs. Questions: REAL user audio from VoiceAssistant-400K (audio_q.*.tar, filenames match input_audio basenames in the manifests). Answers: synthesized in ONE fixed agent voice (LibriSpeech train-clean-100 narrator ref via ResembleAI/Chatterbox; audio_a.*.tar matching assistant_audio). Manifests carry transcripts, answer text, and GLM-4-Voice speech tokens for both sides (glm_in_tokens question /… See the full description on the dataset page: https://huggingface.co/datasets/z050209/va-spoken-qa-agentvoice-stage2.
Stage-2 spoken-QA training corpus (gemma4_talker)
97,697 train / 935 val spoken QA pairs. Questions: REAL user audio from VoiceAssistant-400K (audio_q.*.tar, filenames match input_audio basenames in the manifests). Answers: synthesized in ONE fixed agent voice (LibriSpeech train-clean-100 narrator ref via ResembleAI/Chatterbox; audio_a.*.tar matching assistant_audio). Manifests carry transcripts, answer text, and GLM-4-Voice speech tokens for both sides (glm_in_tokens question / glm_tokens answer, 12.5Hz, THUDM/glm-4-voice-tokenizer).
Sources & licenses: VoiceAssistant-400K (research use), LibriSpeech CC-BY-4.0 (speaker reference), synthesized audio via Chatterbox. Companion model: z050209/gemma2-9b-glm-speechlm-s1.
