datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Speech-MASSIVE
Speech-MASSIVE
Dataset Description
Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages (Arabic, German, Spanish, French, Hungarian, Korean, Dutch, Polish, European Portuguese, Russian, Turkish, and Vietnamese) from different families and inherits from MASSIVE the annotations for the intent prediction and slot-filling tasks. MASSIVE… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE.Speech-MASSIVE-test
Speech-MASSIVE Test Split
This dataset repository is only for test split of Speech-MASSIVE.
train and dev splits are available in the separate dataset repository. https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE
Dataset Description
Speech-MASSIVE is a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages (Arabic, German, Spanish, French, Hungarian… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE-test.massive-yt-edu-queue
Massive YouTube Educational Video Queue
Full metadata and content classification for 4,489,228 YouTube educational videos totaling 3,975,157 hours.
Description
This dataset contains metadata, content categorization, and license risk assessment for ~4.5M YouTube videos identified as potentially educational. It serves as the discovery and processing queue for the massive-yt-edu-transcriptions project, which aims to create the world's largest open educational transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-queue.massive-yt-edu-transcriptions
Massive YouTube Educational Transcriptions
Large-scale educational content transcribed from YouTube using distil-whisper/distil-large-v3.5.
Stats
Videos: 59,355
Characters: 1,539,022,925 (~384M tokens)
Audio hours: 35,890
Model: faster-whisper (CTranslate2) with distil-large-v3.5
Hardware: 2x RTX 5090 + 2x RTX 4090 at 165-185x realtime
Fields
Field
Description
video_id
YouTube video ID
title
Video title
text
Full transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-transcriptions.Speech-MASSIVE_vie
Vietnamse subset of the Speech-MASSIVE dataset
extracted from:
https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE
https://huggingface.co/datasets/FBK-MT/Speech-MASSIVE-test
my extraction code: https://github.com/phineas-pta/fine-tune-whisper-vi/blob/main/misc/load-speechmassive.py
speech_MASSIVE_pt-PTpt-PT subset from FBK-MT/Speech-MASSIVE
massive-audio-transcription-pipeline
massive-audio-transcription-pipeline outputs
Transcription outputs from the
massive-audio-transcription-pipeline,
a parallel Whisper pipeline that chunks long audio into overlapping windows,
transcribes across a worker pool, merges lightweight speaker diarization, and
checkpoints every chunk for crash resume.
Generation method
Backend: faster-whisper base model (CTranslate2), 1 worker.
Audio: real public-domain speech from the Hugging Face LibriSpeech dummy… See the full description on the dataset page: https://huggingface.co/datasets/narinzar/massive-audio-transcription-pipeline.
