datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mya-tiktok-asr-120h
Burmese TikTok ASR (121h)
A weakly-supervised Burmese (Myanmar, my) speech corpus: 168,852 short audio clips / 121.1 hours, segmented from 3,902 public TikTok videos and paired with the Burmese subtitles TikTok generates for those videos.
Intended for pre-training and fine-tuning Burmese ASR models (e.g. Whisper) in a language with very little open speech data.
⚠️ Read this first. The transcripts are machine-generated, not human-verified — see Labels are ASR output. Treat that… See the full description on the dataset page: https://huggingface.co/datasets/t7188409/mya-tiktok-asr-120h.TikTokDownload
