tsdocode/open-vi-dialog-synthetic-100h
OpenDialog Vietnamese Synthetic Dialogue 100h Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments. 12,000 chunks 30 seconds per chunk 100.0 hours total Each item contains S1/S2 speaker labels, turn timings, target text, relationship, pronouns, environment, topic, mood, and source reference IDs. Audio renderer: vLLM-Omni VoxCPM2 Audio format: mono WAV, 48 kHz, 30 seconds per chunk This is a research dataset. Review the source/reference licensing and the… See the full description on the dataset page: https://huggingface.co/datasets/tsdocode/open-vi-dialog-synthetic-100h.
OpenDialog Vietnamese Synthetic Dialogue 100h
Synthetic Vietnamese two-speaker dialogue for ZipVoice-Dialog experiments.
- 12,000 chunks
- 30 seconds per chunk
- 100.0 hours total
- Each item contains S1/S2 speaker labels, turn timings, target text, relationship, pronouns, environment, topic, mood, and source reference IDs.
- Audio renderer: vLLM-Omni VoxCPM2
- Audio format: mono WAV, 48 kHz, 30 seconds per chunk
This is a research dataset. Review the source/reference licensing and the terms of the underlying TTS model before redistribution or commercial use.
Files:
manifest.jsonl: original local manifestmetadata.jsonl: Hugging Face-friendly relative audio metadatatrain.jsonl/validation.jsonl: deterministic 11,760 / 240 split manifestsQC_REPORT.json: full audio integrity and duplicate-hash reportchunks/: 30-second WAV files
