datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TIE_shorts
Dataset Card for TIE_Shorts
Dataset Summary
TIE_shorts is a derived version of the Technical Indian English (TIE) dataset, a large-scale speech dataset (~ 8K hours) originally consisting of approximately 750 GB of content
sourced from the NPTEL platform. The original TIE dataset contains around 9.8K technical lectures in English delivered by instructors from various regions across India,
with each lecture averaging about 50 minutes. These lectures cover a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/raianand/TIE_shorts.short_selling
Short Selling
Data Notice: This dataset provides academic research access with a 6-month data lag.
For real-time data access, please visit sov.ai to subscribe.
For market insights and additional subscription options, check out our newsletter at blog.sov.ai.
from datasets import load_dataset
df_over_shorted = load_dataset("sovai/short_selling", split="train").to_pandas().set_index(["ticker","date"])
Data is updated weekly as data arrives after market close US-EST time.
Tutorials… See the full description on the dataset page: https://huggingface.co/datasets/sovai/short_selling.Short-Storygen-v2I'd recommend you use this dataset instead because its not slopped unlike this one
Short Stories generated by Opus.
Original dataset by Sao10K
youtube-shorts-dataset
YouTube Shorts Dataset
1077 vertical YouTube Shorts from 48 creators across
19 categories. H.264/AAC, 5–60s, mean 37.3s.
Metadata only, no video files. Every row is a pointer to a public
YouTube video. Download them with the bundled scripts, see
Getting the videos.
Data Fields
field
type
description
id
string
YouTube video id
title
string
Video title
webpage_url
string
Canonical youtube.com/watch?v=... URL
source_query
string
Channel /shorts URL… See the full description on the dataset page: https://huggingface.co/datasets/mxmprhd/youtube-shorts-dataset.Short_Stories_ShareGPT
Short_Stories_ShareGPT
2.7 Million Stories
This is a collection of short stories generated by GPT-4.
The dataset has been cleaned of any artifacts, ASCII junk, and standardized into the ShareGPT JSON format.
Note
This dataset is NOT ready for supervised fine-tuning (SFT) "as is" and would require additional data engineering to be usable.
r_shortstories_24kFiltered and somewhat cleaned up scrape of posts from r/shortstories subreddit. Still has some reddit artifacts, but should be usable as is for training.
ShortStory-SFT-jsonl
Public Domain Short Fiction with Prompts
719 complete short stories (300 to 2500 words) by 26 authors whose work is in the public domain,
each paired with a natural-language request that could plausibly have produced it. Built for supervised
fine-tuning of small language models on fiction, where the usual sources (forum stories, model-generated
stories) lack the structural control of published short fiction.
Fields
field
description
id
stable id (hash… See the full description on the dataset page: https://huggingface.co/datasets/Travis-ML/ShortStory-SFT-jsonl.short-summary
Dataset Card for short_summary
Dataset Description
50000 News Articles with corresponding short summary
Languages
The text in the dataset is in English
Dataset Structure
The dataset consists of two columns namely Excerpt and Summary.
The Excerpt column consists of short text from the news article and the Summary column consists of the few words summary of the excerpt
Source Data
The dataset is scrapped from Otherweb database
Short_sentences_about_lovemedical-shorts-silver-claims-benchmark
Medical Shorts Silver Claims Benchmark
This dataset contains silver labels for evaluating medical claim detection and verification in short-form YouTube videos.
The dataset does not include video files, audio files, thumbnails, or keyframes.
It only includes YouTube video IDs, metadata, normalized medical claims, silver labels, and evidence source references.
Dataset Size
Videos: 236
Silver claims: 470
Visual-dependent claims: 22
Videos with medical claims: 205
Videos… See the full description on the dataset page: https://huggingface.co/datasets/ArtemkaT08/medical-shorts-silver-claims-benchmark.Youtube_shorts_comments
Fine-tuned distilgpt2 on this dataset
The average amount of emojis in a YouTube short comment is 4.59 (Based on this dataset)
12 millions view 😂😂😂 good god 🤦♂️
shortsents_sweepmetricsSHORT_sweepmetricsshortstories_synthlabelsWIP
Nothing to see here, it's just some shortstories from internet with synthetic writing prompts
short_synthetic_website
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
short-stories-syntheticyoutube-ai-slop-shorts-dataset
YouTube AI Slop Shorts Dataset
A dataset of 4,224 YouTube channels and 139,136 Shorts video IDs labeled for AI-generated content detection.
Overview
This dataset was created to help researchers, developers, and content moderators identify AI-generated ("slop") content on YouTube Shorts. It contains channel-level and video-level labels for AI slop detection.
Stats
Metric
Value
Channels
4,224
Shorts Videos
139,136
Channels with labels
948
Videos… See the full description on the dataset page: https://huggingface.co/datasets/AIButtonFoundation/youtube-ai-slop-shorts-dataset.Short-StoriesThis dataset was created based on https://huggingface.co/datasets/mintujupally/ROCStories
Citation
@misc{mostafazadeh_corpus_2016,
title = {A {Corpus} and {Evaluation} {Framework} for {Deeper} {Understanding} of {Commonsense} {Stories}},
copyright = {arXiv.org perpetual, non-exclusive license},
url = {https://arxiv.org/abs/1604.01696},
doi = {10.48550/ARXIV.1604.01696},
publisher = {arXiv},
author = {Mostafazadeh, Nasrin and Chambers, Nathanael and He… See the full description on the dataset page: https://huggingface.co/datasets/TOFU-SFT/Short-Stories.short-storiesshort-storiesshortStoriesshort_synthetic_website
Dataset Card
Add more information here
This dataset was produced with DataDreamer 🤖💤. The synthetic dataset card can be found here.
shorts_youtube-comments
Youtube/Youtube Shorts Comments Dataset (Russian & English)
Brief Description
A large dataset of approximately 290,000 comments collected from YouTube Shorts videos in Russian and English.
Each comment is stored on its own line.
Dataset Description
The dataset contains raw text comments, one comment per line.
It is suitable for training language models, text analysis, or other NLP tasks.
Data Format
File type: plain text (.txt)
Each comment occupies… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/shorts_youtube-comments.short_squadShortStoriesShort-Storyshort-summaries-translatedshort_slovak_sentiment
short-slovak-sentiment
Created from AIOD platform
Erebus-R_ShortStories-Combinedshort-seo-descriptions
