Peacockery/farsi-asr-corpus-v4
farsi-asr-corpus-v4 985 hours of Farsi ASR training data across seven corpora: Common Voice 25 (324 h), Thomcles Farsi speech (298 h), Farsi YouTube (205 h), Mana TTS (97 h), Neyshekar (37 h), FLEURS fa_ir (13 h), WorldSpeech (11 h). This is the dataset behind Peacockery/omni-ctc-300m-farsi (8.5% WER on FLEURS test). Layout Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=fas_Arab/. Each row holds text (the normalized label)… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/farsi-asr-corpus-v4.
051
