CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01common-pile /youtube_filtered Creative Commons YouTube Description YouTube is a large-scale video-sharing platform where users have the option of uploading content under a CC BY license. To collect high-quality speech-based textual content and combat the rampant license laundering on YouTube, we manually curated a set of over 2,000 YouTube channels that consistently release original openly licensed content containing speech. The resulting collection spans a wide range of genres, including lectures… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/youtube_filtered.texttext-generation100K<n<1M6 likes866 downloads1y agoHugging Face02ScriptSmith /sponsorblock-youtube-metadata-2024 SponsorBlock YouTube Metadata Dataset A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos. Contains the top videos from the SponsorBlock database that had data added in the year 2024. Quick Stats Metric Value Total videos 154,536 Videos with subtitles 62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.imagetext-classification10M<n<100M0 likes397 downloads2mo agoHugging Face03common-pile /youtube Creative Commons YouTube Description YouTube is large-scale video-sharing platform where users have the option of uploading content under a CC BY license. To collect high-quality speech-based textual content and combat the rampant license laundering on YouTube, we manually curated a set of over 2,000 YouTube channels that consistently release original openly licensed content containing speech. The resulting collection spans a wide range of genres, including lectures… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/youtube.texttext-generation1M<n<10M13 likes355 downloads1y agoHugging Face04Rijgersberg /YouTube-Commons YouTube Commons Re-upload This is a re-upload of PleIAs' YouTube Commons, a valuable open dataset: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC BY 4.0 license. Content The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels). Unfortunately, there are problems with loading YouTube Commons with Hugging Face Datasets. In order to alleviate those… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons.tabulartext-generation10M<n<100M6 likes336 downloads2y agoHugging Face05dm-petrov /youtube-commons-small 📺 YouTube-Commons-Small 📺 This is a smaller subset of the YouTube-Commons dataset, which is a collection of audio transcripts from videos shared on YouTube under a CC-By license. Dataset Description This smaller version contains a subset of the original dataset, maintaining the same structure and features. It's designed for easier experimentation and testing purposes. Features The dataset includes the following information for each video: Video ID and link… See the full description on the dataset page: https://huggingface.co/datasets/dm-petrov/youtube-commons-small.tabulartext-generation100K<n<1M1 likes274 downloads1y agoHugging Face06community-datasets /youtube_caption_corrections Dataset Card for YouTube Caption Corrections Dataset Summary This dataset is built from pairs of YouTube captions where both an auto-generated and a manually-corrected caption are available for a single specified language. It currently only in English, but scripts at repo support other languages. The motivation for creating it was from viewing errors in auto-generated captions at a recent virtual conference, with the hope that there could be some way to help correct those… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/youtube_caption_corrections.textother10K<n<100K8 likes177 downloads2y agoHugging Face07algerian-nlp /Algerian-Youtube-Comments Algerian Youtube Comments 55,365 raw YouTube comments on Algeria-related videos for Darija social-text modeling, from the Algerian NLP Collective. Counted 2026-09-17 via the Hub datasets-server (/info?dataset=algerian-nlp/Algerian-Youtube-Comments: 55,365 train rows) and re-counted row-by-row with datasets streaming (load_dataset("algerian-nlp/Algerian-Youtube-Comments", split="train", streaming=True): 55,365 rows). The default config answers: how do Algerians actually write in… See the full description on the dataset page: https://huggingface.co/datasets/algerian-nlp/Algerian-Youtube-Comments.texttext-generation10K<n<100K0 likes171 downloads8d agoHugging Face08amine-khelif /YouTube-Commons YouTube Commons Re-upload This is a re-upload of PleIAs' YouTube Commons, a valuable open dataset: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC BY 4.0 license. Content The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels). Unfortunately, there are problems with loading YouTube Commons with Hugging Face Datasets. In order to alleviate those… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/YouTube-Commons.tabulartext-generation10M<n<100M0 likes151 downloads10mo agoHugging Face09AnandforU /youtube-comment-insights-chatml YouTube Comment Insights - ChatML Overview This dataset contains instruction-tuning samples for structured YouTube comment analysis. The dataset is formatted in ChatML conversational format and is intended for supervised fine-tuning (SFT), QLoRA, and instruction tuning of large language models. Each sample contains: sentiment tone pros cons Dataset Statistics ~20k training samples ~2k validation samples Multilingual YouTube comments Structured JSON… See the full description on the dataset page: https://huggingface.co/datasets/AnandforU/youtube-comment-insights-chatml.texttext-classification10K<n<100K0 likes75 downloads5mo agoHugging Face10AdamLucek /youtube-titles Youtube Title & Descriptions Dataset About 4941 videos across 50 YouTube Channels List of sampled channels here Splits: Train: 4199 Validation: 493 Test: 249 Data was shuffled and sampled evenly from all channels to create splits. Additionally, has a column ready to go for gemma-2-9b-it fine tuning formatting! Potentially more model formats to come. About the Data: Label Description channel_name The… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/youtube-titles.texttext-generation1K<n<10K1 likes74 downloads2y agoHugging Face11samuelandaudreymedianetwork /samuel-and-audrey-youtube-transcripts-en Samuel & Audrey YouTube Transcripts EN Corpus, 2012–2026 This dataset contains the English transcript archive from the Samuel and Audrey - Travel and Food Videos YouTube channel. The corpus covers travel and food videos published between 2012 and 2026. It includes full transcript records, cue-level transcript segments, YouTube video identifiers, publication dates, titles, view counts captured at export time, tags, source URLs, transcript text, and subtitle-style payloads where… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-and-audrey-youtube-transcripts-en.texttext-generation1M<n<10M1 likes71 downloads4mo agoHugging Face12samuelandaudreymedianetwork /samuel-y-audrey-youtube-transcripts-es-en Samuel y Audrey Bilingual YouTube Transcript Corpus ES/EN This dataset contains a structured bilingual transcript corpus from the Samuel y Audrey Spanish-language travel channel. The corpus includes 643 video records with Spanish and English transcript material, video-level metadata, subtitle-style text, and cleaned transcript fields. It is intended for non-commercial research, translation analysis, retrieval workflows, language study, and media archive organization. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-y-audrey-youtube-transcripts-es-en.texttranslation1K<n<10K1 likes67 downloads4mo agoHugging Face13shawhin /youtube-titles-dpoDataset to fine-tune Qwen2.5 on my YouTube title preferences via DPO. Synthetic titles were generated used Qwen2.5-7B via Together AI's API. Video link Blog link GitHub Repo Fine-tuned Model texttext-generation1K<n<10K2 likes65 downloads2y agoHugging Face14samuelandaudreymedianetwork /nomadic-samuel-youtube-transcripts-corpus Nomadic Samuel YouTube Transcripts Corpus This dataset contains a curated corpus of full-length English transcript records from the Nomadic Samuel YouTube channel. The corpus includes 143 video transcript records with cleaned transcript text, original subtitle-style .srt payloads, video metadata, tags, view counts captured at export time, source URLs, and caption timing information where available. It is intended for non-commercial research, transcript search, retrieval workflows… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/nomadic-samuel-youtube-transcripts-corpus.texttext-generationn<1K1 likes57 downloads4mo agoHugging Face15AnandforU /youtube-comment-insights-clean YouTube Comment Insights - Clean Dataset Overview This dataset contains structured YouTube comment analytics data designed for visualization, analytics, and machine learning workflows. Each sample contains: comment sentiment tone pros cons The dataset is intended for easy readability and downstream analytics tasks. Dataset Statistics ~20k training samples ~2k validation samples Multilingual YouTube comments Structured JSON format Files… See the full description on the dataset page: https://huggingface.co/datasets/AnandforU/youtube-comment-insights-clean.texttext-classification10K<n<100K0 likes40 downloads5mo agoHugging Face16LimYeri /leetcode_with_youtube_captionsimagetext-classification10K<n<100K1 likes39 downloads2y agoHugging Face17k-imtz /youtu-llm-2b-base-blind-spots Youtu-LLM-2B-Base Blind Spots Evaluation Dataset This dataset contains 75 evaluation prompts used to analyze the failure modes of tencent/Youtu-LLM-2B-Base, a 1.96B parameter dense base language model released on December 31, 2025. Each row includes the input prompt, the expected answer, and the model’s generated output obtained during inference on a Google Colab T4 GPU. The prompts span 13 broad categories including arithmetic, logic, multilingual generation, instruction following… See the full description on the dataset page: https://huggingface.co/datasets/k-imtz/youtu-llm-2b-base-blind-spots.texttext-generationn<1K0 likes24 downloads7mo agoHugging Face18DinoResearch /YouTube-Comment-Master-2024-v1 🎮 Roblox MM2 YouTube Comment Dataset (2024 Master) A curated dataset of 27,089 clean, deduplicated, and length-filtered YouTube Short comments scraped from top Roblox Murder Mystery 2 (MM2) videos across 2024. This dataset captures real-world internet gaming culture, short-form video engagement patterns, emoji distributions, trader slang, and brainrot banter—making it ideal for fine-tuning compact LLMs (such as Qwen2.5 or Llama 3) for casual gaming roleplay, comment generation… See the full description on the dataset page: https://huggingface.co/datasets/DinoResearch/YouTube-Comment-Master-2024-v1.texttext-generation10K<n<100K0 likes22 downloads2mo agoHugging Face19snfacademy /snfa-youtube-videodaten SNFA YouTube-Videodaten Ein strukturierter Datensatz mit veröffentlichten YouTube-Videos der SNF Academy und zugehörigen Inhalten aus den Bereichen Fitness, Ernährung, Coaching, Mindset, Personal Training und Ausbildung. Datensatzübersicht 1'423 eindeutige Videos 1'423 eindeutige YouTube-Video-IDs 1'065 Videos mit Beschreibung 472'190 erfasste Views Veröffentlichungszeitraum: 6. September 2013 bis 10. Mai 2026 Datenprüfung: 17. Juli 2026 Sprache: überwiegend… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/snfa-youtube-videodaten.texttext-generation1K<n<10K0 likes20 downloads2mo agoHugging Face20LimYeri /LeetCode_YouTube_CCLeetCode Information & YouTube Captions Original data -> LimYeri/leetcode_with_youtube_captions The original ['cc_content'] column had many repeated sentences, making the data too long. To remove the repetitions, we used precise regular expressions to eliminate the repeated sentences. -> new column ['content'] Additionally, we also removed unnecessary strings (e.g., '[Music]'). texttext-classification10K<n<100K1 likes18 downloads2y agoHugging Face21jmp1987 /simson-youtube-tutorials 📺 Simson YouTube Tutorial Metadata 20 kuratierte YouTube-Tutorial-Einträge für Simson-Moped Reparatur, Tuning und Restaurierung. Inhalt Strukturierte Metadaten der wichtigsten Simson-Tutorial-Videos auf YouTube: Kanal-Typen: DIY-Werkstatt, Tuning-Spezialist, Restaurierungs-Kanal, Enthusiasten-Kanal, Dokumentation Topics: Motor, Zündung, Vergaser, Elektrik, Tuning, Restaurierung, Fahrwerk, Geschichte, Wartung Fahrzeuge: S50, S51, S70, KR51/1, KR51/2 (Schwalbe)… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/simson-youtube-tutorials.tabulartext-generationn<1K0 likes15 downloads4mo agoHugging Face221ArmedMonkey /youtube Creative Commons YouTube Description YouTube is large-scale video-sharing platform where users have the option of uploading content under a CC BY license. To collect high-quality speech-based textual content and combat the rampant license laundering on YouTube, we manually curated a set of over 2,000 YouTube channels that consistently release original openly licensed content containing speech. The resulting collection spans a wide range of genres, including… See the full description on the dataset page: https://huggingface.co/datasets/1ArmedMonkey/youtube.texttext-generation1M<n<10M0 likes10 downloads3mo agoHugging Face23Decre99 /Test_Youtubetexttext-generationn<1K0 likes9 downloads3y agoHugging Face24shravyakasturi /telugu-youtube-corpustexttext-generationn<1K0 likes7 downloads1y agoHugging Face25Innovina /Test_Youtube_Linkstexttext-generationn<1K0 likes5 downloads3y agoHugging Face26Corneille1 /youtu-llm-2b-base-blindspots Blind Spots of tencent/Youtu-LLM-2B-Base Model Tested Model: tencent/Youtu-LLM-2B-BaseLink: https://huggingface.co/tencent/Youtu-LLM-2B-Base This evaluation was conducted on the base (pretrained) version of the model, not an instruction-tuned variant. Objective The goal of this dataset is to identify systematic failure patterns ("blind spots") of the Youtu-LLM-2B-Base model through targeted probing. The evaluation focuses on arithmetic reasoning, unit… See the full description on the dataset page: https://huggingface.co/datasets/Corneille1/youtu-llm-2b-base-blindspots.texttext-generationn<1K0 likes5 downloads7mo agoHugging Face27Candace352 /youtu-llm-2b-base-blindspots Youtu-LLM-2B-Base Blind Spots This dataset contains 10 failure cases collected while probing tencent/Youtu-LLM-2B-Base, an open base language model on Hugging Face. The model card describes it as a Base release, lists it at 1.96B parameters, and notes support for 131,072 context length. Model tested Model: tencent/Youtu-LLM-2B-Base Model type: Base model Parameters: 1.96B Context length: 131,072 I selected this model because it fit the assignment constraints well: it is… See the full description on the dataset page: https://huggingface.co/datasets/Candace352/youtu-llm-2b-base-blindspots.texttext-generationn<1K0 likes3 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.