CoolFace
Datasetpublic

wayu-ai/thai-aligner-bench

Thai Aligner Bench 🚧 Development in progress. How accurately can a forced aligner place Thai token and word boundaries in speech? This is a self-contained benchmark: one Python file (aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing ground truth. No Thai NLP stack or other code is needed — just numpy soundfile torch torchaudio transformers. The ground truth is what makes the dataset useful: the audio was rendered by a TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
1likes773downloads

wayu-ai/thai-aligner-bench · main · files are served by the source, never re-hosted here