kadirnar/voicehub-arena-seed-tts-eval
VoiceHub Arena — full English Seed-TTS-Eval 35,904 synthesized WAV files: 33 model families × the same 1,088 target texts. The campaign completed on 15 September 2026 on one NVIDIA A100-SXM4 40 GB. All 198 shards and every WAV SHA256 were verified after backup. Interactive leaderboard and all audio samples · Source repository (access required). Contents audio_shards/<model>.tar: 33 WebDataset shards, each containing 1,088 original WAVs and matching JSON metadata.… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/voicehub-arena-seed-tts-eval.
VoiceHub Arena — full English Seed-TTS-Eval
35,904 synthesized WAV files: 33 model families × the same 1,088 target texts. The campaign completed on 15 September 2026 on one NVIDIA A100-SXM4 40 GB. All 198 shards and every WAV SHA256 were verified after backup.
Interactive leaderboard and all audio samples · Source repository (access required).
Contents
audio_shards/<model>.tar: 33 WebDataset shards, each containing 1,088 original WAVs and matching JSON metadata.records/<model>.json: all 1,088 original per-sample records, targets, ASR transcripts, edit counts, timing, signal diagnostics and source-run identifiers.audio-manifest.json: archive, entry, byte offset, size and original synthesis SHA256 for every WAV.audio-shards.json: size and SHA256 for each archive.leaderboard.json/leaderboard.csv: the full 33-model table.reports/: complete API results, scientific plots, metrics and backup/coverage proofs.
The per-sample JSON files inside each archive provide file_name, model, id, reference, transcript, checkpoint, seed, normalization, audio_sha256, duration, sample rate, synthesis latency, RTF, peak CUDA allocation, silence/clipping ratios and WER/CER/MER/WIL/WIP. Recordings are model outputs, not original publisher prompt recordings. No model weights or account credentials are included.
Evaluation protocol
English target texts from ByteDance Seed-TTS-Eval, publisher revision 752f4297f090c46bb1a55a1f7439e5944ddefe8d, split en/meta.lst. Full target JSONL SHA256: 8a9386efb1768a90ffba66f8931e3887b53e9c9d1bf246eb46a8cce32bc7b1c1. The publisher describes the English test material as originating from Common Voice.
Recognition: Systran/faster-whisper-large-v3, revision edaa852ec7e145841d8ffdb056a99866b5f0a478, CUDA FP16, English, beam 5, temperature 0, no VAD, no previous-text conditioning. Normalizer: whisper-normalizer==0.1.12, whisper_english. Synthesis seed 42, one repeat, identity input text. Each model uses its frozen provider voice/reference.
WER/CER aggregate edit counts over the complete corpus. Confidence intervals use 1,000 prompt-cluster bootstrap resamples, RNG seed 42. RTF is total synthesis time divided by generated audio duration and excludes model load, downloads and ASR. Memory is the measured per-sample peak CUDA allocation, not total process VRAM.
Limitations and provenance
This is fixed-voice intelligibility evaluation, not an exact reproduction of the official zero-shot speaker-identity/SIM protocol. No MOS or speaker similarity score is claimed. ASR error rates do not independently measure naturalness. Dia and Llasa have high error rates with unresolved root causes. Empty ASR output does not by itself prove silent audio. All samples, including poor outputs, remain. FishTTS and VITS execution repairs preserved verified successful samples; the source repository records their parity checks and changes.
This repository documents a benchmark experiment, not a new blanket license for upstream texts, voices, models or their outputs. Consult the linked original dataset and each checkpoint's terms before reuse; checkpoint identifiers and revisions are retained in the records and reports. No original prompt-audio collection or model weights are redistributed here.
The Space reads only the selected WAV byte range from an archive and verifies its original SHA256 in the browser before playback. WAV bytes are unchanged. Archives can also be downloaded and extracted with standard tar tools.
Quality audit — 15 September 2026
CosyVoice is Fun-CosyVoice3-0.5B-2512, base llm.pt. The archived 13.82% WER is affected by a confirmed HiFT implementation defect and is excluded from ranking. The corrected full 1,088-text evaluation is verified: WER 1.7416%, CER 0.6234%. The current table, samples and plots use this full run. The selected eight-text pilot was not used as the replacement score. Verification and preserved original data. Llasa and Dia remain under quality review after independent LM/codec checks.
Read the investigation and paired audio.
Comparison chart exports
Interactive charts and full table support nine metrics and custom model selection.
reports/bars/ contains PNG and SVG bar charts for WER, CER, synthesis RTF and peak CUDA allocation, each with best-six and all-33 views. WER/CER include measured 95% bootstrap intervals. Invalidated scores and quality reviews are marked explicitly. reports/bars/manifest.json records every plotted value, unit, interval and the exact leaderboard input SHA256. No scores or recordings were changed for these charts.
Additional high-WER audit
Investigation, plots and selected audio. reports/high-wer-audit-2026-09-15/ preserves all 96 newly generated diagnostic WAVs, 132 scored records, full-corpus integrity checks, source configurations, numerical comparisons and a confirmed ConversationTTS duration-limit repair. Each of six models uses six deliberately selected texts across multiple variants. Diagnostic scores do not replace full 1,088-text benchmark scores. Original recordings and table data are unchanged.
