CoolFace
Datasetpublic

z050209/va-spoken-qa-agentvoice-stage2

Stage-2 spoken-QA training corpus (gemma4_talker) 97,697 train / 935 val spoken QA pairs. Questions: REAL user audio from VoiceAssistant-400K (audio_q.*.tar, filenames match input_audio basenames in the manifests). Answers: synthesized in ONE fixed agent voice (LibriSpeech train-clean-100 narrator ref via ResembleAI/Chatterbox; audio_a.*.tar matching assistant_audio). Manifests carry transcripts, answer text, and GLM-4-Voice speech tokens for both sides (glm_in_tokens question /… See the full description on the dataset page: https://huggingface.co/datasets/z050209/va-spoken-qa-agentvoice-stage2.

sourceHugging Facecc-by-4.0updated 1mo agoView on Hugging Face
0likes8downloads
settings

This repository belongs to z050209 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameva-spoken-qa-agentvoice-stage2
visibilitypublic
licencecc-by-4.0
gatedno
ownerz050209
Account settings
z050209/va-spoken-qa-agentvoice-stage2 · CoolFace