CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jamescalam /youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp. Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.tabularquestion-answering100K<n<1M44 likes1.3k downloads4y agoHugging Face02united-nations /transcription-corpus UN Transcription Corpus Two splits of UN meeting audio paired with official verbatim records. Splits sessions — Whole meeting sessions (SC + GA plenary) One row per meeting. Audio from UN Web TV, verbatim records from documents.un.org. Column Description symbol UN document symbol, e.g. S/PV.9826 webtv_url URL on UN Web TV duration_ms Session duration in milliseconds num_speakers Number of speaker turns in the verbatim record audio_floor Floor… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-corpus.audioautomatic-speech-recognitionn<1K0 likes222 downloads7mo agoHugging Face03jq /salt-asr-data-transcriptionstabular10K<n<100K0 likes208 downloads2y agoHugging Face04IljaSamoilov /ERR-transcription-to-subtitlesThis dataset is created by Ilja Samoilov. In dataset is tv show subtitles from ERR and transcriptions of those shows created with TalTech ASR. from datasets import load_dataset, load_metric dataset = load_dataset('csv', data_files={'train': "train.tsv", \ "validation":"val.tsv", \ "test": "test.tsv"}, delimiter='\t') tabular100K<n<1M0 likes68 downloads4y agoHugging Face05AKCIT-Audio /CHiME6_formatted_transcriptionstabular10K<n<100K0 likes64 downloads1y agoHugging Face06thepowerfuldeez /massive-yt-edu-transcriptions Massive YouTube Educational Transcriptions Large-scale educational content transcribed from YouTube using distil-whisper/distil-large-v3.5. Stats Videos: 59,355 Characters: 1,539,022,925 (~384M tokens) Audio hours: 35,890 Model: faster-whisper (CTranslate2) with distil-large-v3.5 Hardware: 2x RTX 5090 + 2x RTX 4090 at 165-185x realtime Fields Field Description video_id YouTube video ID title Video title text Full transcript… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/massive-yt-edu-transcriptions.tabularautomatic-speech-recognition10K<n<100K3 likes62 downloads4mo agoHugging Face07DataFog /medical-transcription-instruct About This dataset consists of 38,924 samples of instruct-input-output data, most helpfully for training instruction-following models tailored to the medical field Dataset Summary Source: Original medical transcriptions with added instruction-output pairs Size: 38,924 instruction-output pairs Format: CSV file Domain: Medical / Healthcare Language: English Last Updated: 08-20-2024 Dataset Structure Each row in the dataset represents a unique… See the full description on the dataset page: https://huggingface.co/datasets/DataFog/medical-transcription-instruct.tabular10K<n<100K31 likes60 downloads2y agoHugging Face08rungalileo /medical_transcription_4tabular1K<n<10K5 likes45 downloads4y agoHugging Face09pinecone /yt-transcriptionsimage10K<n<100K1 likes44 downloads4y agoHugging Face10ghanaopenai /ghanaian-english-words-corrected-transcriptions This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Ghanaian English Transcript Corrections A dataset of mistranscribed words and phrases from Ghanaian news media YouTube videos, corrected using Llama 3.1 405B. Source Extracted from YouTube transcripts of Ghanaian… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghanaian-english-words-corrected-transcriptions.tabular100K<n<1M0 likes43 downloads3mo agoHugging Face11Ched-ai /voynich-transcription-mismatch Voynich Transcription Mismatch Index Dataset Summary This dataset provides a line-by-line comparison across five different transcription sources of the Voynich Manuscript. It tracks agreements and disagreements between transcribers, enabling research on transcription uncertainty and consensus. The primary comparison is between EVA-based transcriptions (ZL and IT), with additional tracking of Currier, FSG, and v101 transcription systems. Key Statistics… See the full description on the dataset page: https://huggingface.co/datasets/Ched-ai/voynich-transcription-mismatch.tabularother1K<n<10K0 likes36 downloads3mo agoHugging Face12rakesh-ai /Medical_Transcriptiontabular1K<n<10K1 likes25 downloads3y agoHugging Face13odunola /bible-passage-transcriptiontabular10K<n<100K5 likes14 downloads3y agoHugging Face14jdabello /yt_transcriptionsimage10K<n<100K0 likes13 downloads3y agoHugging Face15AndresR2909 /youtube_transcriptions_summaries_finetunned_llama3_2_3btabularn<1K0 likes11 downloads2y agoHugging Face16successor /qrl-yt-transcriptionstabular10K<n<100K0 likes10 downloads4y agoHugging Face17AndresR2909 /youtube_transcriptions_summaries_gpt4o_vs_llama3_1_8b_vs_llama3_2_3btabularn<1K0 likes10 downloads2y agoHugging Face18nepalprabin /youtube-transcriptionstabular10K<n<100K0 likes9 downloads4y agoHugging Face19AndresR2909 /youtube_transcriptions_summaries_gpt4o_vs_llama3_1_8btabularn<1K0 likes9 downloads2y agoHugging Face20AndresR2909 /youtube_transcriptions_summaries_gpt4o_vs_llama3_1_8b_vs_llama3_2_3b_1b_v2tabularn<1K0 likes9 downloads2y agoHugging Face21AndresR2909 /youtube_transcriptions_summaries_2025_gpt4otabular1K<n<10K0 likes9 downloads1y agoHugging Face22AndresR2909 /youtube_transcriptions_train_dataset_2025_gpt4.1tabular1K<n<10K0 likes9 downloads1y agoHugging Face23bejaeger /biggest_ideas_transcriptions Dataset Card for "biggest_ideas_transcriptions" More Information needed tabular10K<n<100K0 likes8 downloads4y agoHugging Face24AKCIT-Audio /LIGHT_transcriptionstabular100K<n<1M0 likes8 downloads1y agoHugging Face25AndresR2909 /youtube_transcriptions_summaries_gpt4_vs_llama31tabularn<1K0 likes7 downloads2y agoHugging Face26bejaeger /filled_stacks_transcriptions Dataset Card for "filled_stacks_transcriptions" More Information needed tabular10K<n<100K0 likes6 downloads4y agoHugging Face27bhargavi909 /Medical_Transcriptions_upsampledtabular10K<n<100K3 likes6 downloads3y agoHugging Face28AndresR2909 /youtube_transcriptions_ingesttabular1K<n<10K0 likes6 downloads2y agoHugging Face29AriMattiPodcastTranscript /ari-matti-podcast-transcriptionstabular10K<n<100K0 likes6 downloads2y agoHugging Face30AndresR2909 /youtube_transcriptions_summaries_gpt4otabular1K<n<10K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.