datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
youtube-transcriptionsThe YouTube transcriptions dataset contains technical tutorials (currently from James Briggs, Daniel Bourke, and AI Coffee Break) transcribed using OpenAI's Whisper (large). Each row represents roughly a sentence-length chunk of text alongside the video URL and timestamp.
Note that each item in the dataset contains just a short chunk of text. For most use cases you will likely need to merge multiple rows to create more substantial chunks of text, if you need to do that, this code snippet will… See the full description on the dataset page: https://huggingface.co/datasets/jamescalam/youtube-transcriptions.AnonyRAG
AnnoyRAG Dataset
The AnnoyRAG dataset, introduced in Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex Reasoning, employs entity anonymization to isolate LLMs' parametric knowledge. This design enables more precise evaluation of how effectively LLMs integrate retrieved information in RAG systems.
Dataset Details
Dataset Description
The basic statistical information of the dataset is as follows:
Question Type
Difficulty Level… See the full description on the dataset page: https://huggingface.co/datasets/Youtu-Graph/AnonyRAG.samuel-and-audrey-youtube-transcripts-en
Samuel & Audrey YouTube Transcripts EN Corpus, 2012–2026
This dataset contains the English transcript archive from the Samuel and Audrey - Travel and Food Videos YouTube channel.
The corpus covers travel and food videos published between 2012 and 2026. It includes full transcript records, cue-level transcript segments, YouTube video identifiers, publication dates, titles, view counts captured at export time, tags, source URLs, transcript text, and subtitle-style payloads where… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-and-audrey-youtube-transcripts-en.samuel-y-audrey-youtube-transcripts-es-en
Samuel y Audrey Bilingual YouTube Transcript Corpus ES/EN
This dataset contains a structured bilingual transcript corpus from the Samuel y Audrey Spanish-language travel channel.
The corpus includes 643 video records with Spanish and English transcript material, video-level metadata, subtitle-style text, and cleaned transcript fields. It is intended for non-commercial research, translation analysis, retrieval workflows, language study, and media archive organization.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/samuel-y-audrey-youtube-transcripts-es-en.nomadic-samuel-youtube-transcripts-corpus
Nomadic Samuel YouTube Transcripts Corpus
This dataset contains a curated corpus of full-length English transcript records from the Nomadic Samuel YouTube channel.
The corpus includes 143 video transcript records with cleaned transcript text, original subtitle-style .srt payloads, video metadata, tags, view counts captured at export time, source URLs, and caption timing information where available.
It is intended for non-commercial research, transcript search, retrieval workflows… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/nomadic-samuel-youtube-transcripts-corpus.youtu-llm-2b-base-blind-spots
Youtu-LLM-2B-Base Blind Spots Evaluation Dataset
This dataset contains 75 evaluation prompts used to analyze the failure modes of tencent/Youtu-LLM-2B-Base,
a 1.96B parameter dense base language model released on December 31, 2025. Each row includes the input prompt, the expected answer, and the model’s
generated output obtained during inference on a Google Colab T4 GPU.
The prompts span 13 broad categories including arithmetic, logic, multilingual generation, instruction following… See the full description on the dataset page: https://huggingface.co/datasets/k-imtz/youtu-llm-2b-base-blind-spots.snfa-youtube-videodaten
SNFA YouTube-Videodaten
Ein strukturierter Datensatz mit veröffentlichten YouTube-Videos der SNF Academy und zugehörigen Inhalten aus den Bereichen Fitness, Ernährung, Coaching, Mindset, Personal Training und Ausbildung.
Datensatzübersicht
1'423 eindeutige Videos
1'423 eindeutige YouTube-Video-IDs
1'065 Videos mit Beschreibung
472'190 erfasste Views
Veröffentlichungszeitraum: 6. September 2013 bis 10. Mai 2026
Datenprüfung: 17. Juli 2026
Sprache: überwiegend… See the full description on the dataset page: https://huggingface.co/datasets/snfacademy/snfa-youtube-videodaten.simson-youtube-tutorials
📺 Simson YouTube Tutorial Metadata
20 kuratierte YouTube-Tutorial-Einträge für Simson-Moped Reparatur, Tuning und Restaurierung.
Inhalt
Strukturierte Metadaten der wichtigsten Simson-Tutorial-Videos auf YouTube:
Kanal-Typen: DIY-Werkstatt, Tuning-Spezialist, Restaurierungs-Kanal, Enthusiasten-Kanal, Dokumentation
Topics: Motor, Zündung, Vergaser, Elektrik, Tuning, Restaurierung, Fahrwerk, Geschichte, Wartung
Fahrzeuge: S50, S51, S70, KR51/1, KR51/2 (Schwalbe)… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/simson-youtube-tutorials.turkce-otomotiv-yedek-parca-youtube-soru-cevap-akil-yurutme
Türkçe Otomobil Yedek Parça Soru-Cevap Akıl Yürütme Veri Seti
Bu veri seti, Türkçe otomobil yedek parça sorguları üzerine akıl yürütme (reasoning) yoluyla cevap vermek amacıyla tasarlanmış yapay zeka eğitim verisidir. Sorgular ağırlıklı olarak Opel marka araçlar için yedek parça sorularından oluşmakta olup, diğer markalar için de örnekler içermektedir.
İçerik
Özellik
Değer
Toplam Satır
~260
Dil
Türkçe
Lisans
MIT
Format
JSONL
Veri Yapısı… See the full description on the dataset page: https://huggingface.co/datasets/aiprojecom/turkce-otomotiv-yedek-parca-youtube-soru-cevap-akil-yurutme.
