typhoon-ai/gigaspeech2-typhoon
Gigaspeech2 Typhoon Project page | Paper | GitHub Gigaspeech2 Typhoon is a metadata-only reference dataset for Thai speech recognition benchmarking, specifically designed as an Accuracy Track for evaluating ASR models. The dataset contains 1,000 test samples with audio IDs and human transcriptions derived from the Gigaspeech2 corpus. Each audio_id directly links to the original Gigaspeech2 dataset, allowing users to download the corresponding audio. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/gigaspeech2-typhoon.
Fix the citation of Gigaspeech2 paper.
Add metadata, task categories and project links (#3)
Update README.md
Update README.md
Update README.md
Update README.md
Update to metadata-only dataset (removed audio data for licensing compliance)
initial commit
