CoolFace
Datasetpublic

typhoon-ai/gigaspeech2-typhoon

Gigaspeech2 Typhoon Project page | Paper | GitHub Gigaspeech2 Typhoon is a metadata-only reference dataset for Thai speech recognition benchmarking, specifically designed as an Accuracy Track for evaluating ASR models. The dataset contains 1,000 test samples with audio IDs and human transcriptions derived from the Gigaspeech2 corpus. Each audio_id directly links to the original Gigaspeech2 dataset, allowing users to download the corresponding audio. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/gigaspeech2-typhoon.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
1likes67downloads
8 commits on main
e442c6c4mo ago

Fix the citation of Gigaspeech2 paper.

Warit
c0a385d8mo ago

Add metadata, task categories and project links (#3)

Warit, nielsr
3193d838mo ago

Update README.md

Warit
deb560f8mo ago

Update README.md

Warit
8d44e498mo ago

Update README.md

Warit
22ebc878mo ago

Update README.md

Warit
ce6d0568mo ago

Update to metadata-only dataset (removed audio data for licensing compliance)

Warit
28ada248mo ago

initial commit

Warit