VoiceHub/voicehub-arena-seed-tts-eval
VoiceHub Arena — native TTS evaluations Incrementally published generated audio and WER, CER, DNSMOS, WavLM-large ECAPA speaker SIM and UTMOS22 measurements. The full campaign is still running. Each generation method is evaluated separately using its publisher's native API. Full evaluations contain all 1,088 English Seed-TTS-Eval targets; eight-target diagnostic pilots are stored separately and must not be treated as full scores. Interactive demo · Source code Layout… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/voicehub-arena-seed-tts-eval.
VoiceHub Arena — native TTS evaluations
Incrementally published generated audio and WER, CER, DNSMOS, WavLM-large ECAPA speaker SIM and UTMOS22 measurements. The full campaign is still running. Each generation method is evaluated separately using its publisher's native API. Full evaluations contain all 1,088 English Seed-TTS-Eval targets; eight-target diagnostic pilots are stored separately and must not be treated as full scores.
Interactive demo · Source code
Layout and streaming audio
experiments/<campaign>/<pilot|full>/<method>/ contains result.json, records.json, contract.json, verification.json, and audio.tar. records.json contains a rows array. Each successful row specifies its WAV's audio_archive, audio_offset, audio_bytes, and audio_sha256. Request exactly that byte range from the archive at an immutable dataset commit; it is a complete WAV and does not require downloading the whole archive. Verify its SHA-256 before use. Failed generations have no fabricated audio.
Per-campaign inventory.json, progress.json, and native-comparison.csv record coverage and completed scores. The native-ada-full-20260916 campaign is current; native-ada-20260916 is an earlier diagnostic campaign.
Archives and offsets are verified locally, remote object hashes are checked at the returned commit, and a ranged audio download is tested before local audio is offloaded. Speaker SIM is N/A for generation methods without paired reference conditioning. Predicted MOS is not a human listening score.
Source texts and paired reference protocol: ByteDance Seed-TTS-Eval, revision 752f4297f090c46bb1a55a1f7439e5944ddefe8d. Checkpoint/source commits and generation settings are retained in each contract. Upstream model and source-data terms continue to apply.
Active user selection
The current scope is 37 arena models and 71 distinct generation methods. No additional model families, GGUF/quantized checkpoints, fast/blockwise/chunked variants, or duplicate streaming modes are evaluated. Streaming-only official APIs are used once. Voice design, reference-audio cloning and continuation remain separate where supported. NeuTTS-Air is excluded; arena NeuTTS-2E remains. The frozen original inventory is preserved; selection.json defines active methods.
