kadirnar/VyvoTTS-EN-Beta-DPO-samples
VyvoTTS EN-Beta — 2,000 automatic DPO pairs Exactly 2,000 unique target texts and chosen/rejected pairs, generated by Vyvo/VyvoTTS-EN-Beta at revision 70b37a5bfdbdc2f478515837081048aac63f909e. Each row embeds the actual 24 kHz reference, chosen and rejected audio, raw prompt and completion codec IDs, transcripts, WER/CER, DNSMOS P.835, sampling settings, seeds and waveform SHA-256 checksums. Both candidate waveforms are actual model outputs; no artificial corruption.… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/VyvoTTS-EN-Beta-DPO-samples.
VyvoTTS EN-Beta — 2,000 automatic DPO pairs
Exactly 2,000 unique target texts and chosen/rejected pairs, generated by Vyvo/VyvoTTS-EN-Beta at revision 70b37a5bfdbdc2f478515837081048aac63f909e. Each row embeds the actual 24 kHz reference, chosen and rejected audio, raw prompt and completion codec IDs, transcripts, WER/CER, DNSMOS P.835, sampling settings, seeds and waveform SHA-256 checksums. Both candidate waveforms are actual model outputs; no artificial corruption.
Reference-audio continuation conditioning matches the working EN-Beta inference interface. Candidates use temperatures 0.6, 1.0 and 1.2, top-p 0.95, top-k 20, repetition penalty 1.1. Preferred status is determined by metrics, never by temperature. Maximum prompt-plus-completion length is 2100 tokens; no audio tokens are truncated. The sampler is qwen3-compact-codebook-v1: valid-codebook projection with finished rows removed from the batch. Its RNG sequence differs from dense generation; backend and per-candidate batch size are retained in generation metadata.
Selection and limitations
Whisper-large-v3 FP32 transcribes every valid candidate. Selection uses Whisper English normalization after Unicode NFKC and curly-apostrophe canonicalization (whisper-english-unicode-v1). Original seed-style and original Whisper-normalized WER/CER are also retained. Chosen WER must be <=0.10 and DNSMOS OVRL >=3.0. A pair must improve WER by >=0.05 with DNSMOS loss <=0.05, or improve DNSMOS by >=0.1 without WER loss. Ties, invalid codec output and overlength completions are not preference pairs.
Labels are automatic, not human-reviewed. DNSMOS is an estimate, and ASR can make errors. Speaker similarity is not a selection metric. This dataset alone does not establish post-training improvements; evaluate trained checkpoints separately. The train split contains all 2,000 rows. Select a separate validation set before training; do not use seed-tts-eval for training or checkpoint selection.
Source and attribution
Text and reference recordings come only from the train-clean-100 split of LibriTTS-R (CC BY 4.0), by Yuma Koizumi et al., “LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus” (2023), derived from LibriTTS and LibriSpeech/LibriVox. Target speech was newly synthesized by EN-Beta; source IDs, speaker IDs and reference IDs are retained. Exact text matches and eight-word overlaps with seed-tts-eval English reference/target texts are excluded. No benchmark waveform is used as a training reference.
The source checkpoint has no declared license in its model card. This dataset does not assign a new license to the checkpoint or override source rights. See source_manifest.json, mining_report.json and validation.json for provenance and counts. This replaces the earlier 15-row demonstration in the same repo.
