datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp.
Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.transcription-corpus
UN Transcription Corpus
Two splits of UN meeting audio paired with official verbatim records.
Splits
sessions — Whole meeting sessions (SC + GA plenary)
One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org.
Column
Description
symbol
UN document symbol, e.g. S/PV.9826
webtv_url
URL on UN Web TV
duration_ms
Session duration in milliseconds
num_speakers
Number of speaker turns in the verbatim record
audio_floor
Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.salt-asr-data-transcriptionsERR-transcription-to-subtitlesThis dataset is created by Ilja Samoilov. In dataset is tv show subtitles from ERR and transcriptions of those shows created with TalTech ASR.
from datasets import load_dataset, load_metric
dataset = load_dataset('csv', data_files={'train': "train.tsv", \
"validation":"val.tsv", \
"test": "test.tsv"}, delimiter='\t')
CHiME6_formatted_transcriptionsmassive-yt-edu-transcriptions
Massive YouTube Educational Transcriptions
Large-scale educational content transcribed from YouTube using distil-whisper/distil-large-v3.5.
Stats
Videos: 59,355
Characters: 1,539,022,925 (~384M tokens)
Audio hours: 35,890
Model: faster-whisper (CTranslate2) with distil-large-v3.5
Hardware: 2x RTX 5090 + 2x RTX 4090 at 165-185x realtime
Fields
Field
Description
video_id
YouTube video ID
title
Video title
text
Full transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-transcriptions.medical-transcription-instruct
About
This dataset consists of 38,924 samples of instruct-input-output data, most helpfully for training instruction-following models tailored to the medical field
Dataset Summary
Source: Original medical transcriptions with added instruction-output pairs
Size: 38,924 instruction-output pairs
Format: CSV file
Domain: Medical / Healthcare
Language: English
Last Updated: 08-20-2024
Dataset Structure
Each row in the dataset represents a unique… See the full description on the dataset page: https://huggingface.co/datasets/DataFog/medical-transcription-instruct.medical_transcription_4yt-transcriptionsghanaian-english-words-corrected-transcriptions
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Ghanaian English Transcript Corrections
A dataset of mistranscribed words and phrases from Ghanaian news media YouTube videos, corrected using Llama 3.1 405B.
Source
Extracted from YouTube transcripts of Ghanaian… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghanaian-english-words-corrected-transcriptions.voynich-transcription-mismatch
Voynich Transcription Mismatch Index
Dataset Summary
This dataset provides a line-by-line comparison across five different transcription sources of the Voynich Manuscript. It tracks agreements and disagreements between transcribers, enabling research on transcription uncertainty and consensus.
The primary comparison is between EVA-based transcriptions (ZL and IT), with additional tracking of Currier, FSG, and v101 transcription systems.
Key Statistics… See the full description on the dataset page: https://huggingface.co/datasets/Ched-ai/voynich-transcription-mismatch.Medical_Transcriptionbible-passage-transcriptionyt_transcriptionsyoutube_transcriptions_summaries_finetunned_llama3_2_3bqrl-yt-transcriptionsyoutube_transcriptions_summaries_gpt4o_vs_llama3_1_8b_vs_llama3_2_3byoutube-transcriptionsyoutube_transcriptions_summaries_gpt4o_vs_llama3_1_8byoutube_transcriptions_summaries_gpt4o_vs_llama3_1_8b_vs_llama3_2_3b_1b_v2youtube_transcriptions_summaries_2025_gpt4oyoutube_transcriptions_train_dataset_2025_gpt4.1biggest_ideas_transcriptions
Dataset Card for "biggest_ideas_transcriptions"
More Information needed
LIGHT_transcriptionsyoutube_transcriptions_summaries_gpt4_vs_llama31filled_stacks_transcriptions
Dataset Card for "filled_stacks_transcriptions"
More Information needed
Medical_Transcriptions_upsampledyoutube_transcriptions_ingestari-matti-podcast-transcriptionsyoutube_transcriptions_summaries_gpt4o
