CoolFace
Datasetpublicgated

speechcolab/gigaspeech2

Dataset Card for GigaSpeech 2 Dataset Description GigaSpeech 2 is an evolving, large-scale, multi-domain, and multilingual ASR corpus focusing on low-resource languages. GigaSpeech 2 raw comprises about 30,000 hours of automatically transcribed speech, across Thai, Indonesian, and Vietnamese. GigaSpeech 2 refine consists of 10,000 hours of Thai, 6,000 hours each for Indonesian and Vietnamese. Repository: https://github.com/SpeechColab/GigaSpeech2 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/speechcolab/gigaspeech2.

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
71likes5.8kdownloads
settings

This repository belongs to speechcolab on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namegigaspeech2
visibilitypublic
licenceapache-2.0
gatedyes
ownerspeechcolab
Account settings
speechcolab/gigaspeech2 · CoolFace