broadcast
broadcast
Broadcast for 🇺🇦 Ukrainian
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Stats
Total files processed: 136736
Total duration: 300h 10m 51s
Other
Labels generated by https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
broadcast-speech
Bashkir Broadcast Speech — Radio and Television
53.3 hours of speech in 293 recordings in the Bashkir language, from television and radio programmes produced by two public broadcasters of the Republic of Bashkortostan. Audio only — no transcripts in this release — which makes the set suitable for self-supervised speech pretraining for a low-resource Turkic language.
🌐 Languages of this card: English · Башҡортса · Русский
Part of the Bashkorttele dataset series — preservation… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/broadcast-speech.broadcast-opus
Broadcast for 🇺🇦 Ukrainian (in OPUS)
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Stats
Total files processed: 136736
Total duration: 300h 10m 51s
Other
Labels generated by https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
broadcast-speech-uk
Broadcast Speech Dataset for Ukrainian
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Statistics
Total duration: 300.181 hours
Number of unique transcriptions: 127441
Duration statistics
Metrics
Value
mean
7.903199
std
3.615765
min
4.99781
25%
5.64
50%
6.65
75%
8.66
max
29.99006
Cite this work
@misc… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/broadcast-speech-uk.BroadcastSpeechno-asr-eval-data-broadcast-sample
NRK Norwegian Speech Dataset (Sample)
Dataset Description
Note: This is a sample dataset containing a subset of chunks for demonstration and preview purposes.
The full dataset is available privately.
This dataset contains Norwegian speech data from NRK TV broadcasts (norge-rundt, supernytt), processed for automatic speech recognition (ASR) evaluation and research.
Reference text caveat: Reference text is NRK's on-air teletext subtitling, not a verbatim… See the full description on the dataset page: https://huggingface.co/datasets/NRK-KIHUB/no-asr-eval-data-broadcast-sample.
