datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
broadcast
Broadcast for 🇺🇦 Ukrainian
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Stats
Total files processed: 136736
Total duration: 300h 10m 51s
Other
Labels generated by https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
broadcast-speech
Bashkir Broadcast Speech — Radio and Television
53.3 hours of speech in 293 recordings in the Bashkir language, from television and radio programmes produced by two public broadcasters of the Republic of Bashkortostan. Audio only — no transcripts in this release — which makes the set suitable for self-supervised speech pretraining for a low-resource Turkic language.
🌐 Languages of this card: English · Башҡортса · Русский
Part of the Bashkorttele dataset series — preservation… See the full description on the dataset page: https://huggingface.co/datasets/bashkorttele/broadcast-speech.broadcast-opus
Broadcast for 🇺🇦 Ukrainian (in OPUS)
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Stats
Total files processed: 136736
Total duration: 300h 10m 51s
Other
Labels generated by https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3
broadcast-speech-uk
Broadcast Speech Dataset for Ukrainian
Community
Discord: https://bit.ly/discord-uds
Speech Recognition: https://t.me/speech_recognition_uk
Speech Synthesis: https://t.me/speech_synthesis_uk
Statistics
Total duration: 300.181 hours
Number of unique transcriptions: 127441
Duration statistics
Metrics
Value
mean
7.903199
std
3.615765
min
4.99781
25%
5.64
50%
6.65
75%
8.66
max
29.99006
Cite this work
@misc… See the full description on the dataset page: https://huggingface.co/datasets/Yehor/broadcast-speech-uk.no-asr-eval-data-broadcast-sample
NRK Norwegian Speech Dataset (Sample)
Dataset Description
Note: This is a sample dataset containing a subset of chunks for demonstration and preview purposes.
The full dataset is available privately.
This dataset contains Norwegian speech data from NRK TV broadcasts (norge-rundt, supernytt), processed for automatic speech recognition (ASR) evaluation and research.
Reference text caveat: Reference text is NRK's on-air teletext subtitling, not a verbatim… See the full description on the dataset page: https://huggingface.co/datasets/NRK-KIHUB/no-asr-eval-data-broadcast-sample.broadCast_dataset_2hat_asr_sixian_broadcast_clean
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_Sixian
客語_四縣
231.77
88,262
3,633,737
9.45
4.36
Total
-
231.77
88,262
3,633,737
9.45
4.36
broadCast_dataset_5state-of-geo-2026
The 2026 State of Generative Engine Optimization
860 AI answers scored across 85 B2B software companies in 61 categories, with every cited source traced: 5,160 citations. All citation URLs were string-matched against reddit.com, quora.com and stackoverflow.com: zero hits. Vendor-authored pages were 77.4% of citations. Volume II, The Absence Ladder, classifies the 616 answers where a company was absent and finds the shape of absence shifts with existing visibility (chi-square… See the full description on the dataset page: https://huggingface.co/datasets/Broadcastwell/state-of-geo-2026.hat_asr_sixian_broadcast_clean_r
hat_asr_sixian_broadcast_clean_r
This dataset is an enhanced -R variant of formospeech/hat_asr_sixian_broadcast_clean.
Summary
Subset: Hakka_Sixian
Dialect: 客語四縣
Train samples: 88263
Audio: enhanced 24 kHz WAV
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_Sixian
客語_四縣
231.77
88,263
8,172,269
9.45
9.79
Total
-
231.77
88,263
8,172,269
9.45
9.79
Processing
Start from the original… See the full description on the dataset page: https://huggingface.co/datasets/formospeech/hat_asr_sixian_broadcast_clean_r.broadCast_dataset_6broadCast_dataset_1groupK_project_BroadcastflowCN
A Dataset of Longitudinal Acoustic Changes for Middle-Aged Voice Aging based on 15 Years of Voice Data from Program Host in Mainland China
Abstract
This dataset provides a unique 15-year longitudinal acoustic record tracking the vocal aging of a professional Mandarin-speaking television host from age 35 to 50 (2012–2026). Comprising high-quality audio extracted from broadcast interviews and programs, the collection captures the subtle physiological and prosodic… See the full description on the dataset page: https://huggingface.co/datasets/eduhk-compling/groupK_project_BroadcastflowCN.broadCast_dataset_8032_broadcast_translation
032.방송콘텐츠 한국어-영어 번역 말뭉치 / 587,084개
qa_broadcast_conv_etbroadCast_dataset_3023_broadcast_script_summary_3sentBroadcastSpeech_part001broadCast_dataset_7023_broadcast_script_summary_20percentBroadcastSpeech_part002kms_broadcast_eval
Dataset: KMS Broadcast Eval
Description
This dataset contains financial broadcast evaluation data for compliance checks. The dataset includes labeled data to identify violations based on specific criteria.
Files
kms_broadcast_eval.csv: The main dataset containing context, violation status, reasons, and categories.
Label Definitions
0: Not violated
1: Violated
Columns
id: Unique identifier for each row
context: The actual advertisement… See the full description on the dataset page: https://huggingface.co/datasets/bonmuq/kms_broadcast_eval.LLM_Broadcaster_Song_IntroductionsBroadcastSpeech_part000BroadcastSpeech_part003broadCast_dataset_4BroadcastConsumptionPatterns
BroadcastConsumptionPatterns
tags: TVGenre, DemographicSegmentation, ConsumptionTrendAnalysis
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'BroadcastConsumptionPatterns' dataset aims to provide insights into how different demographic segments consume television content across various genres. This dataset includes information on the viewing habits, preferred genres, and trend analysis over a specified period.
CSV Content… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/BroadcastConsumptionPatterns.g-broadcastinghat_asr_sixian_broadcast
TRAIN
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
hak_sx
Hakka_Sixian
240.64
80,073
3,757,980
10.82
4.34
0
0
Total
-
240.64
80,073
3,757,980
10.82
4.34
0
0
