datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
broadcast
Broadcast for 🇺🇦 Ukrainian
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Stats
Total files processed: 136736
Total duration: 300h 10m 51s
Other
Labels generated by https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
broadcast-speech
Bashkir Broadcast Speech — Radio and Television
53.3 hours of speech in 293 recordings in the Bashkir language, from television and radio programmes produced by two public broadcasters of the Republic of Bashkortostan. Audio only — no transcripts in this release — which makes the set suitable for self-supervised speech pretraining for a low-resource Turkic language.
🌐 Languages of this card: English · Башҡортса · Русский
Part of the Bashkorttele dataset series — preservation… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/broadcast-speech.broadcast-opus
Broadcast for 🇺🇦 Ukrainian (in OPUS)
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Stats
Total files processed: 136736
Total duration: 300h 10m 51s
Other
Labels generated by https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
broadcast-speech-uk
Broadcast Speech Dataset for Ukrainian
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Statistics
Total duration: 300.181 hours
Number of unique transcriptions: 127441
Duration statistics
Metrics
Value
mean
7.903199
std
3.615765
min
4.99781
25%
5.64
50%
6.65
75%
8.66
max
29.99006
Cite this work
@misc… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/broadcast-speech-uk.
