simpra/xh-tts-check
ViXSD segmentation check A stratified sample of 40 clips from 3,861 produced by vixsd_segment.py. Listen to each and confirm the audio says exactly the text. A misaligned pair teaches the model a wrong sound-to-letter mapping, and it is invisible to every automated check. worst rows are the lowest-scoring clips in the whole run — if those are right, the rest almost certainly are. clip why sec score ends transcription cds_xho_079_1_0025.wav worst 1.7 0.41 sentence… See the full description on the dataset page: https://huggingface.co/datasets/simpra/xh-tts-check.
0141
