Jarbas/ovos-tts-bench
created as part of OVOS TTS plugin benchmarks Metrics RTF - Real Time Factor - real time factor, how many seconds it takes to create 1 second of audio - (lower is better) WER - Word Error Rate - this is a proxy for understandability, assuming more understandable speech scores better in STT, correlates to how many words the TTS pronounces wrong - (lower is better) DAMERAU LEVENSHTEIN SIMILARITY - this is also a proxy for understandability, assuming more understandable speech… See the full description on the dataset page: https://huggingface.co/datasets/Jarbas/ovos-tts-bench.
012
created as part of OVOS TTS plugin benchmarks
Metrics
- RTF - Real Time Factor - real time factor, how many seconds it takes to create 1 second of audio - (lower is better)
- WER - Word Error Rate - this is a proxy for understandability, assuming more understandable speech scores better in STT, correlates to how many words the TTS pronounces wrong - (lower is better)
- DAMERAU LEVENSHTEIN SIMILARITY - this is also a proxy for understandability, assuming more understandable speech scores better - (higher is better)
- Pitch Variability - Measures the variation in the pitch of the speech. Higher variability can indicate more natural, human-like speech, while low variability may suggest robotic or monotone output. (Higher is better)
NOTE: for STT google is used, and in case of failure Whisper large V3 as fallback
