Peacockery/tajik-asr-corpus-v3
tajik-asr-corpus-v3 1,071 hours of Tajik ASR training data: 41 Tajik YouTube channels (~1,059 h, machine-labeled) plus FLEURS tg_tj (11.8 h, gold). This is the dataset behind Peacockery/omni-ctc-300m-tajik (16.9% WER on FLEURS test, 37.6% on held-out conversational speech). Layout Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=tgk_Cyrl/. Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list), and… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-corpus-v3.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face