datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Short-Storygen-v2I'd recommend you use this dataset instead because its not slopped unlike this one
Short Stories generated by Opus.
Original dataset by Sao10K
Short_Stories_ShareGPT
Short_Stories_ShareGPT
2.7 Million Stories
This is a collection of short stories generated by GPT-4.
The dataset has been cleaned of any artifacts, ASCII junk, and standardized into the ShareGPT JSON format.
Note
This dataset is NOT ready for supervised fine-tuning (SFT) "as is" and would require additional data engineering to be usable.
r_shortstories_24kFiltered and somewhat cleaned up scrape of posts from r/shortstories subreddit. Still has some reddit artifacts, but should be usable as is for training.
ShortStory-SFT-jsonl
Public Domain Short Fiction with Prompts
719 complete short stories (300 to 2500 words) by 26 authors whose work is in the public domain,
each paired with a natural-language request that could plausibly have produced it. Built for supervised
fine-tuning of small language models on fiction, where the usual sources (forum stories, model-generated
stories) lack the structural control of published short fiction.
Fields
field
description
id
stable id (hash… See the full description on the dataset page: https://huggingface.co/datasets/Travis-ML/ShortStory-SFT-jsonl.medical-shorts-silver-claims-benchmark
Medical Shorts Silver Claims Benchmark
This dataset contains silver labels for evaluating medical claim detection and verification in short-form YouTube videos.
The dataset does not include video files, audio files, thumbnails, or keyframes.
It only includes YouTube video IDs, metadata, normalized medical claims, silver labels, and evidence source references.
Dataset Size
Videos: 236
Silver claims: 470
Visual-dependent claims: 22
Videos with medical claims: 205
Videos… See the full description on the dataset page: https://huggingface.co/datasets/ArtemkaT08/medical-shorts-silver-claims-benchmark.Youtube_shorts_comments
Fine-tuned distilgpt2 on this dataset
The average amount of emojis in a YouTube short comment is 4.59 (Based on this dataset)
12 millions view 😂😂😂 good god 🤦♂️
shortstories_synthlabelsWIP
Nothing to see here, it's just some shortstories from internet with synthetic writing prompts
youtube-ai-slop-shorts-dataset
YouTube AI Slop Shorts Dataset
A dataset of 4,224 YouTube channels and 139,136 Shorts video IDs labeled for AI-generated content detection.
Overview
This dataset was created to help researchers, developers, and content moderators identify AI-generated ("slop") content on YouTube Shorts. It contains channel-level and video-level labels for AI slop detection.
Stats
Metric
Value
Channels
4,224
Shorts Videos
139,136
Channels with labels
948
Videos… See the full description on the dataset page: https://huggingface.co/datasets/AIButtonFoundation/youtube-ai-slop-shorts-dataset.short-summaries-translatedErebus-R_ShortStories-Combinedshort_stock_strategyshorts_datasetUrsa-ShortStories-Allura-Filteredallura-org_r_shortstories_24k-axo
