datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pseudolabel-science-large-v3-timestamp
Pseudolabel science context audio using Whisper Large V3
Original audio from malaysia-ai/science-context-youtube, we split every 30 seconds and pseudolabelled using Whisper Large V3.
how to prepare the dataset
huggingface-cli download --repo-type dataset \
--include 'science-chunk-*.zip' \
--local-dir './' \
--max-workers 20 \
mesolitica/pseudolabel-science-large-v3-timestamp
wget… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/pseudolabel-science-large-v3-timestamp.Pseudolabels_Whisper_Large_CSALT_FLEURDataset5-batch1_part1_subpart1_seg_20spseudo_labels
