CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jamescalam /youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp. Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.tabularquestion-answering100K<n<1M44 likes1.3k downloads4y agoHugging Face02Youtu-Graph /AnonyRAG AnnoyRAG Dataset The AnnoyRAG dataset, introduced in Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex Reasoning, employs entity anonymization to isolate LLMs' parametric knowledge. This design enables more precise evaluation of how effectively LLMs integrate retrieved information in RAG systems. Dataset Details Dataset Description The basic statistical information of the dataset is as follows: Question Type Difficulty Level… See the full description on the dataset page: https://huggingface.co/datasets/Youtu-Graph/AnonyRAG.textquestion-answering1K<n<10K14 likes472 downloads1y agoHugging Face03samuelandaudreymedianetwork /samuel-and-audrey-youtube-transcripts-en Samuel & Audrey YouTube Transcripts EN Corpus, 2012–2026 This dataset contains the English transcript archive from the Samuel and Audrey - Travel and Food Videos YouTube channel. The corpus covers travel and food videos published between 2012 and 2026. It includes full transcript records, cue-level transcript segments, YouTube video identifiers, publication dates, titles, view counts captured at export time, tags, source URLs, transcript text, and subtitle-style payloads where… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-and-audrey-youtube-transcripts-en.texttext-generation1M<n<10M1 likes71 downloads4mo agoHugging Face04samuelandaudreymedianetwork /samuel-y-audrey-youtube-transcripts-es-en Samuel y Audrey Bilingual YouTube Transcript Corpus ES/EN This dataset contains a structured bilingual transcript corpus from the Samuel y Audrey Spanish-language travel channel. The corpus includes 643 video records with Spanish and English transcript material, video-level metadata, subtitle-style text, and cleaned transcript fields. It is intended for non-commercial research, translation analysis, retrieval workflows, language study, and media archive organization. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-y-audrey-youtube-transcripts-es-en.texttranslation1K<n<10K1 likes67 downloads4mo agoHugging Face05samuelandaudreymedianetwork /nomadic-samuel-youtube-transcripts-corpus Nomadic Samuel YouTube Transcripts Corpus This dataset contains a curated corpus of full-length English transcript records from the Nomadic Samuel YouTube channel. The corpus includes 143 video transcript records with cleaned transcript text, original subtitle-style .srt payloads, video metadata, tags, view counts captured at export time, source URLs, and caption timing information where available. It is intended for non-commercial research, transcript search, retrieval workflows… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/nomadic-samuel-youtube-transcripts-corpus.texttext-generationn<1K1 likes57 downloads4mo agoHugging Face06k-imtz /youtu-llm-2b-base-blind-spots Youtu-LLM-2B-Base Blind Spots Evaluation Dataset This dataset contains 75 evaluation prompts used to analyze the failure modes of tencent/Youtu-LLM-2B-Base, a 1.96B parameter dense base language model released on December 31, 2025. Each row includes the input prompt, the expected answer, and the model’s generated output obtained during inference on a Google Colab T4 GPU. The prompts span 13 broad categories including arithmetic, logic, multilingual generation, instruction following… See the full description on the dataset page: https://huggingface.co/datasets/k-imtz/youtu-llm-2b-base-blind-spots.texttext-generationn<1K0 likes24 downloads7mo agoHugging Face07snfacademy /snfa-youtube-videodaten SNFA YouTube-Videodaten Ein strukturierter Datensatz mit veröffentlichten YouTube-Videos der SNF Academy und zugehörigen Inhalten aus den Bereichen Fitness, Ernährung, Coaching, Mindset, Personal Training und Ausbildung. Datensatzübersicht 1'423 eindeutige Videos 1'423 eindeutige YouTube-Video-IDs 1'065 Videos mit Beschreibung 472'190 erfasste Views Veröffentlichungszeitraum: 6. September 2013 bis 10. Mai 2026 Datenprüfung: 17. Juli 2026 Sprache: überwiegend… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/snfa-youtube-videodaten.texttext-generation1K<n<10K0 likes20 downloads2mo agoHugging Face08jmp1987 /simson-youtube-tutorials 📺 Simson YouTube Tutorial Metadata 20 kuratierte YouTube-Tutorial-Einträge für Simson-Moped Reparatur, Tuning und Restaurierung. Inhalt Strukturierte Metadaten der wichtigsten Simson-Tutorial-Videos auf YouTube: Kanal-Typen: DIY-Werkstatt, Tuning-Spezialist, Restaurierungs-Kanal, Enthusiasten-Kanal, Dokumentation Topics: Motor, Zündung, Vergaser, Elektrik, Tuning, Restaurierung, Fahrwerk, Geschichte, Wartung Fahrzeuge: S50, S51, S70, KR51/1, KR51/2 (Schwalbe)… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/simson-youtube-tutorials.tabulartext-generationn<1K0 likes15 downloads4mo agoHugging Face09aiprojecom /turkce-otomotiv-yedek-parca-youtube-soru-cevap-akil-yurutme Türkçe Otomobil Yedek Parça Soru-Cevap Akıl Yürütme Veri Seti Bu veri seti, Türkçe otomobil yedek parça sorguları üzerine akıl yürütme (reasoning) yoluyla cevap vermek amacıyla tasarlanmış yapay zeka eğitim verisidir. Sorgular ağırlıklı olarak Opel marka araçlar için yedek parça sorularından oluşmakta olup, diğer markalar için de örnekler içermektedir. İçerik Özellik Değer Toplam Satır ~260 Dil Türkçe Lisans MIT Format JSONL Veri Yapısı… See the full description on the dataset page: https://huggingface.co/datasets/aiprojecom/turkce-otomotiv-yedek-parca-youtube-soru-cevap-akil-yurutme.question-answering1 likes14 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.