CoolFace
Datasetpublic

Funk888/gigaspeech2-th-joined

gigaspeech2-th-joined Thai speech-transcript pairs derived from GigaSpeech 2, re-segmented into clips of a length that is convenient for training speech-language models. Built to train a Thai speech adapter for the Ultravox architecture, where very short fragments make poor training examples but long clips do not fit the context budget. Statistics Examples 40,000 Total audio ~42 hours Sample rate 16 kHz, mono Clip duration 2.0–12.0 s (mean 3.8… See the full description on the dataset page: https://huggingface.co/datasets/Funk888/gigaspeech2-th-joined.

sourceHugging Faceupdated 18d agoView on Hugging Face
1likes118downloads
settings

This repository belongs to Funk888 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namegigaspeech2-th-joined
visibilitypublic
licencenot set
gatedno
ownerFunk888
Account settings
Funk888/gigaspeech2-th-joined · CoolFace