datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
youtube_filtered
Creative Commons YouTube
Description
YouTube is a large-scale video-sharing platform where users have the option of uploading content under a CC BY license. To collect high-quality speech-based textual content and combat the rampant license laundering on YouTube, we manually curated a set of over 2,000 YouTube channels that consistently release original openly licensed content containing speech. The resulting collection spans a wide range of genres, including lectures… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/youtube_filtered.sponsorblock-youtube-metadata-2024
SponsorBlock YouTube Metadata Dataset
A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos.
Contains the top videos from the SponsorBlock database that had data added in the year 2024.
Quick Stats
Metric
Value
Total videos
154,536
Videos with subtitles
62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.youtube
Creative Commons YouTube
Description
YouTube is large-scale video-sharing platform where users have the option of uploading content under a CC BY license.
To collect high-quality speech-based textual content and combat the rampant license laundering on YouTube, we manually curated a set of over 2,000 YouTube channels that consistently release original openly licensed content containing speech.
The resulting collection spans a wide range of genres, including lectures… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/youtube.YouTube-Commons
YouTube Commons Re-upload
This is a re-upload of PleIAs' YouTube Commons, a valuable open dataset:
YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC BY 4.0 license.
Content
The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels).
Unfortunately, there are problems with loading YouTube Commons with Hugging Face Datasets.
In order to alleviate those… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons.youtube-commons-small
📺 YouTube-Commons-Small 📺
This is a smaller subset of the YouTube-Commons dataset, which is a collection of audio transcripts from videos shared on YouTube under a CC-By license.
Dataset Description
This smaller version contains a subset of the original dataset, maintaining the same structure and features. It's designed for easier experimentation and testing purposes.
Features
The dataset includes the following information for each video:
Video ID and link… See the full description on the dataset page: https://huggingface.co/datasets/dm-petrov/youtube-commons-small.youtube_caption_corrections
Dataset Card for YouTube Caption Corrections
Dataset Summary
This dataset is built from pairs of YouTube captions where both an auto-generated and a manually-corrected caption are available for a single specified language. It currently only in English, but scripts at repo support other languages. The motivation for creating it was from viewing errors in auto-generated captions at a recent virtual conference, with the hope that there could be some way to help correct those… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/youtube_caption_corrections.Algerian-Youtube-Comments
Algerian Youtube Comments
55,365 raw YouTube comments on Algeria-related videos for Darija social-text modeling, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/Algerian-Youtube-Comments: 55,365 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/Algerian-Youtube-Comments", split="train", streaming=True): 55,365 rows).
The default config answers: how do Algerians actually write in… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/Algerian-Youtube-Comments.YouTube-Commons
YouTube Commons Re-upload
This is a re-upload of PleIAs' YouTube Commons, a valuable open dataset:
YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC BY 4.0 license.
Content
The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels).
Unfortunately, there are problems with loading YouTube Commons with Hugging Face Datasets.
In order to alleviate those… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/YouTube-Commons.youtube-comment-insights-chatml
YouTube Comment Insights - ChatML
Overview
This dataset contains instruction-tuning samples for structured YouTube comment analysis.
The dataset is formatted in ChatML conversational format and is intended for supervised fine-tuning (SFT), QLoRA, and instruction tuning of large language models.
Each sample contains:
sentiment
tone
pros
cons
Dataset Statistics
~20k training samples
~2k validation samples
Multilingual YouTube comments
Structured JSON… See the full description on the dataset page: https://huggingface.co/datasets/AnandforU/youtube-comment-insights-chatml.youtube-titles
Youtube Title & Descriptions Dataset
About
4941 videos across 50 YouTube Channels
List of sampled channels here
Splits:
Train: 4199
Validation: 493
Test: 249
Data was shuffled and sampled evenly from all channels to create splits.
Additionally, has a column ready to go for gemma-2-9b-it fine tuning formatting! Potentially more model formats to come.
About the Data:
Label
Description
channel_name
The… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/youtube-titles.samuel-and-audrey-youtube-transcripts-en
Samuel & Audrey YouTube Transcripts EN Corpus, 2012–2026
This dataset contains the English transcript archive from the Samuel and Audrey - Travel and Food Videos YouTube channel.
The corpus covers travel and food videos published between 2012 and 2026. It includes full transcript records, cue-level transcript segments, YouTube video identifiers, publication dates, titles, view counts captured at export time, tags, source URLs, transcript text, and subtitle-style payloads where… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-and-audrey-youtube-transcripts-en.samuel-y-audrey-youtube-transcripts-es-en
Samuel y Audrey Bilingual YouTube Transcript Corpus ES/EN
This dataset contains a structured bilingual transcript corpus from the Samuel y Audrey Spanish-language travel channel.
The corpus includes 643 video records with Spanish and English transcript material, video-level metadata, subtitle-style text, and cleaned transcript fields. It is intended for non-commercial research, translation analysis, retrieval workflows, language study, and media archive organization.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-y-audrey-youtube-transcripts-es-en.youtube-titles-dpoDataset to fine-tune Qwen2.5 on my YouTube title preferences via DPO. Synthetic titles were generated used Qwen2.5-7B via Together AI's API.
Video link
Blog link
GitHub Repo
Fine-tuned Model
nomadic-samuel-youtube-transcripts-corpus
Nomadic Samuel YouTube Transcripts Corpus
This dataset contains a curated corpus of full-length English transcript records from the Nomadic Samuel YouTube channel.
The corpus includes 143 video transcript records with cleaned transcript text, original subtitle-style .srt payloads, video metadata, tags, view counts captured at export time, source URLs, and caption timing information where available.
It is intended for non-commercial research, transcript search, retrieval workflows… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/nomadic-samuel-youtube-transcripts-corpus.youtube-comment-insights-clean
YouTube Comment Insights - Clean Dataset
Overview
This dataset contains structured YouTube comment analytics data designed for visualization, analytics, and machine learning workflows.
Each sample contains:
comment
sentiment
tone
pros
cons
The dataset is intended for easy readability and downstream analytics tasks.
Dataset Statistics
~20k training samples
~2k validation samples
Multilingual YouTube comments
Structured JSON format
Files… See the full description on the dataset page: https://huggingface.co/datasets/AnandforU/youtube-comment-insights-clean.leetcode_with_youtube_captionsyoutu-llm-2b-base-blind-spots
Youtu-LLM-2B-Base Blind Spots Evaluation Dataset
This dataset contains 75 evaluation prompts used to analyze the failure modes of tencent/Youtu-LLM-2B-Base,
a 1.96B parameter dense base language model released on December 31, 2025. Each row includes the input prompt, the expected answer, and the model’s
generated output obtained during inference on a Google Colab T4 GPU.
The prompts span 13 broad categories including arithmetic, logic, multilingual generation, instruction following… See the full description on the dataset page: https://huggingface.co/datasets/k-imtz/youtu-llm-2b-base-blind-spots.YouTube-Comment-Master-2024-v1
🎮 Roblox MM2 YouTube Comment Dataset (2024 Master)
A curated dataset of 27,089 clean, deduplicated, and length-filtered YouTube Short comments scraped from top Roblox Murder Mystery 2 (MM2) videos across 2024.
This dataset captures real-world internet gaming culture, short-form video engagement patterns, emoji distributions, trader slang, and brainrot banter—making it ideal for fine-tuning compact LLMs (such as Qwen2.5 or Llama 3) for casual gaming roleplay, comment generation… See the full description on the dataset page: https://huggingface.co/datasets/DinoResearch/YouTube-Comment-Master-2024-v1.snfa-youtube-videodaten
SNFA YouTube-Videodaten
Ein strukturierter Datensatz mit veröffentlichten YouTube-Videos der SNF Academy und zugehörigen Inhalten aus den Bereichen Fitness, Ernährung, Coaching, Mindset, Personal Training und Ausbildung.
Datensatzübersicht
1'423 eindeutige Videos
1'423 eindeutige YouTube-Video-IDs
1'065 Videos mit Beschreibung
472'190 erfasste Views
Veröffentlichungszeitraum: 6. September 2013 bis 10. Mai 2026
Datenprüfung: 17. Juli 2026
Sprache: überwiegend… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/snfa-youtube-videodaten.LeetCode_YouTube_CCLeetCode Information & YouTube Captions
Original data -> LimYeri/leetcode_with_youtube_captions
The original ['cc_content'] column had many repeated sentences, making the data too long.
To remove the repetitions, we used precise regular expressions to eliminate the repeated sentences. -> new column ['content']
Additionally, we also removed unnecessary strings (e.g., '[Music]').
simson-youtube-tutorials
📺 Simson YouTube Tutorial Metadata
20 kuratierte YouTube-Tutorial-Einträge für Simson-Moped Reparatur, Tuning und Restaurierung.
Inhalt
Strukturierte Metadaten der wichtigsten Simson-Tutorial-Videos auf YouTube:
Kanal-Typen: DIY-Werkstatt, Tuning-Spezialist, Restaurierungs-Kanal, Enthusiasten-Kanal, Dokumentation
Topics: Motor, Zündung, Vergaser, Elektrik, Tuning, Restaurierung, Fahrwerk, Geschichte, Wartung
Fahrzeuge: S50, S51, S70, KR51/1, KR51/2 (Schwalbe)… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/simson-youtube-tutorials.youtube
Creative Commons YouTube
Description
YouTube is large-scale video-sharing platform where users have the option of uploading content under a CC BY license.
To collect high-quality speech-based textual content and combat the rampant license laundering on YouTube, we manually curated a set of over 2,000 YouTube channels that consistently release original openly licensed content containing speech.
The resulting collection spans a wide range of genres, including… See the full description on the dataset page: https://huggingface.co/datasets/1ArmedMonkey/youtube.Test_Youtubetelugu-youtube-corpusTest_Youtube_Linksyoutu-llm-2b-base-blindspots
Blind Spots of tencent/Youtu-LLM-2B-Base
Model Tested
Model: tencent/Youtu-LLM-2B-BaseLink: https://huggingface.co/tencent/Youtu-LLM-2B-Base
This evaluation was conducted on the base (pretrained) version of the model, not an instruction-tuned variant.
Objective
The goal of this dataset is to identify systematic failure patterns ("blind spots") of the Youtu-LLM-2B-Base model through targeted probing. The evaluation focuses on arithmetic reasoning, unit… See the full description on the dataset page: https://huggingface.co/datasets/Corneille1/youtu-llm-2b-base-blindspots.youtu-llm-2b-base-blindspots
Youtu-LLM-2B-Base Blind Spots
This dataset contains 10 failure cases collected while probing tencent/Youtu-LLM-2B-Base, an open base language model on Hugging Face. The model card describes it as a Base release, lists it at 1.96B parameters, and notes support for 131,072 context length.
Model tested
Model: tencent/Youtu-LLM-2B-Base
Model type: Base model
Parameters: 1.96B
Context length: 131,072
I selected this model because it fit the assignment constraints well: it is… See the full description on the dataset page: https://huggingface.co/datasets/Candace352/youtu-llm-2b-base-blindspots.
