datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
keystroke-typing-videos
Keystroke Typing Videos of Reuters
Recordings of typing randomly sampled sentences (<= 150 characters) from nltk Reuters dataset. Keystroke data is provided too.
AVQA-videos
AVQA — Audio-Visual Question Answering (videos + annotations)
A drop-in package of the AVQA dataset (Yang et al., ACM MM 2022): real-life
audio-visual question answering over short in-the-wild clips. The original release
ships only the QA annotations and expects users to collect the source videos from
VGGSound themselves. This repository bundles the source video clips together
with the official train/val annotations, so the dataset is usable without any
YouTube scraping.… See the full description on the dataset page: https://huggingface.co/datasets/juyil/AVQA-videos.tiktok-videos-4b
TikTok Videos: 4.5 billion posts dataset
Step-by-step guide and access to the scraper code:
tiktok-api.seeksocial.io.
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/kwakuobeng/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts dataset
Step-by-step guide and access to the scraper code:
tiktok-api.seeksocial.io.
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/blaccastro/tiktok-videos-4b.Sci-VBench-Videos
Sci-VBench Videos
Sci-VBench Videos is the complete set of model outputs behind the Sci-VBench paper: 11,216 videos from 16 text-to-video models, together with the automatic and human scores computed on them. Every video was generated from the verbatim benchmark prompt under the model's default configuration — no rewriting, no prompt expansion — so the released prompts and the released videos correspond exactly.
Prompts and evaluation specifications live in the companion repo… See the full description on the dataset page: https://huggingface.co/datasets/Sci-VBench/Sci-VBench-Videos.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/dams2005/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts dataset
Step-by-step guide and access to the scraper code:
tiktok-api.seeksocial.io.
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/KOM-00/tiktok-videos-4b.repo-to-space-example-videos
Gradio Space Example Inputs — Videos
A small, curated, freely-licensed pool of videos used as gr.Examples for
Gradio Spaces that wrap video-input generation models (image-to-video,
video-to-video, motion controls, etc.). Sister dataset for images:
linoyts/repo-to-space-example-inputs.
When a Space takes video input, the agent building the Space picks 2–3 clips
whose caption + categories match the model's task, downloads them via
hf_hub_download, runs any model-specific… See the full description on the dataset page: https://huggingface.co/datasets/linoyts/repo-to-space-example-videos.test-HunyuanVideo-pixelart-videos
trojblue/test-HunyuanVideo-pixelart-images
👋 Heads up—this repository is just a PARTIAL dataset. For the full pixelart-images dataset, make sure to grab both parts:
Images Part
Video Part (this repo)
What's in the Dataset?
This dataset is all about anime-styled pixel art images that have been carefully selected to make your models shine. Here’s what makes these images special:
Rich in detail: Pixelated, yes—but still full of life and not overly simplified.… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/test-HunyuanVideo-pixelart-videos.Depth-Normal-Videos-42K
Depth and Normal Videos Dataset
42,498 videos with depth and surface normals.
Usage
from huggingface_hub import hf_hub_download
video = hf_hub_download(
repo_id="Yanbin99/Depth-Normal-Videos-42K",
filename="Depth_and_Normal_42K/group_0000/videos/00000000.mp4",
repo_type="dataset"
)
tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/alex12223322/tiktok-videos-4b.factory-manipulation-videos
Factory manipulation videos
Procedural Robotics is open sourcing a small set of our factory data so teams can assess its quality. The videos show workers performing factory tasks.
Contents
Seven continuous takes, 109 minutes in total.
Task
Station
Worker
Duration
File
cardboard manipulation
01
041
23.6 min
cardboard_manipulation_station01_worker041.mp4
cardboard manipulation
04
026
16.5 min
cardboard_manipulation_station04_worker026.mp4
defect… See the full description on the dataset page: https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/hojj/tiktok-videos-4b.wrbench-videos
WRBench Benchmark Videos
The current public dataset contains 11,100 model-output videos. The paper table
is the frozen 23-model, 9,600-output paper_main_23model_9600_20260608 surface.
The 2026-07-13 update synchronizes D5/D6 for the 2,073 applicable paper rows
and D3-D6 for 59 EasyAnimate rolling rows with the published aggregate tables.
Only the three videos_master index representations change. Video bytes,
first frames, prompts, IDs, applicability masks, and frozen paper… See the full description on the dataset page: https://huggingface.co/datasets/WRBench/wrbench-videos.tiktok-videos-users-info
TikTok Data, post + poster (user) info, ~1 Million
Dataset name: EinzzCookie/tiktok-videos-users-info
This dataset contains a large collection of TikTok video records paired with detailed creator/user information, stored in a single Parquet file (tiktok_video_user_data.parquet, ~3.87 GB). It is derived from TikTok’s internal video (“aweme”) data model and includes both post-level metadata/engagement stats and nested author profile data.
Source
Collected by… See the full description on the dataset page: https://huggingface.co/datasets/EinzzCookie/tiktok-videos-users-info.tiktok-videos-4b
Mirror of kuben-developer/tiktok-videos-4b, snapshot 2026-09-08. All credit to the original author; same research-use license applies.
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.… See the full description on the dataset page: https://huggingface.co/datasets/merway/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/mrfakename/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/seanphan/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/kkndlee/tiktok-videos-4b.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/tiktok-videos-4b.VideoScore-Bench
Overview
VideoFeedback-Bench is derived from four benchmarks or datasets: VideoFeedback, EvalCrafter, GenAI-Bench and VBench.
Examples
Citation
@article{he2024videoscore,
title = {VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation},
author = {He, Xuan and Jiang, Dongfu and Zhang, Ge and Ku, Max and Soni, Achint and Siu, Sherman and Chen, Haonan and Chandra, Abhranil and Jiang, Ziyan and Arulraj, Aaran and… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/VideoScore-Bench.tiktok-videos-sea
TikTok Videos Southeast Asia (MY / TH / ID / SG)
Country-filtered subset of kuben-developer/tiktok-videos-4b:
every row whose country field is MY (Malaysia), TH (Thailand), ID (Indonesia) or SG (Singapore).
Extracted 2026-09-07 from all 27 source parquet files; schema unchanged.
Sizes (verified against the uploaded files)
Config
Country
Rows
Distinct content_id
Ads (is_ad=1)
With caption
my
Malaysia
11,729,283
11,729,283
967,782
9,787,874
th
Thailand… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/tiktok-videos-sea.video-scissors-sessions
Coding agent session traces for kaofelix/video-scissors-sessions
This dataset contains redacted coding agent session traces collected while working on git@github.com:kaofelix/video-scissors.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line… See the full description on the dataset page: https://huggingface.co/datasets/kaofelix/video-scissors-sessions.reasoning_videos
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
✨ Highlights
|:--:|
| Examples from VideoReasonBench and three existing VideoQA benchmarks. Responses are generated by Gemini-2.5-Flash in both "Thinking" and "No Thinking" modes. |
|:--:|
| Performance of Gemini-2.5-Flash with varying thinking budgets on five benchmarks. |
Complex Reasoning
Gemini-2.5-Flash achieves over 65% accuracy on standard benchmarks but drops to… See the full description on the dataset page: https://huggingface.co/datasets/lyx97/reasoning_videos.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/abdellatifinformation/tiktok-videos-4b.fold_the_towels_videosTiktok-Videos
TikTok Video Analytics Dataset
Sample TikTok video dataset with comprehensive engagement metrics and metadata. Each row represents a single TikTok video with content and detailed analytics.
This is a sample dataset. To access the full version or request any custom dataset tailored to your needs, contact DataHive at contact@datahive.ai.
Files Included
train.csv – TikTok video analytics data
What's included
Video URLs and identifiers
Comprehensive engagement… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/Tiktok-Videos.timeseries_trending_youtube_videos_2019-04-15_to_2020-04-15Timeseries Trending YouTube Videos: 2019-04-15 to 2020-04-15
This dataset is a csv of one of the archived historical database tables queried from my non public database that contains time series data for period of 2019-04-15 to 2020-04-15. Video data was captured from the time they first appeared on trending list, and TSD exists until the video is removed from trending list.
This snapshot contains data for the 11,369 videos that appeared on trending within the timeframe, with 1,541,128 records… See the full description on the dataset page: https://huggingface.co/datasets/jettisonthenet/timeseries_trending_youtube_videos_2019-04-15_to_2020-04-15.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/grimboy/tiktok-videos-4b.videos
