datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
YouTube-Commons-nl-audio
YouTube Commons NL Audio
This dataset has audio files for the Dutch-language videos in Rijgersberg/YouTube-Commons-nl-transcriptions,
all under a CC BY 4.0 license.
It contains 11,669 files for a total runtime of 2493h 43m 5s, coming in at approximately 130 GB.
Source
The original source of the dataset (minus the titles, descriptions and audio files) is YouTube Commons:
YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons-nl-audio.spoken_words_en_ml_commons_filtered_splitcommon-sense-facts-audio
Common-Sense Facts Audio Dataset
A spoken fact-completion dataset for evaluating whether Speech Language Models can retrieve common-sense and factual knowledge from speech.
Each example contains three paired versions:
prompt: an incomplete factual prompt, e.g. "the capital of France is"
fact: the correct full sentence, e.g. "the capital of France is Paris"
counterfactual: an incorrect matched sentence from the same category, e.g. "the capital of France is Rome"
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/slprl/common-sense-facts-audio.Youtube-Commons-Audioyoutube_commons_vad_25_sample
