datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KABR-raw-videos
Dataset Card for KABR Raw Videos: Unprocessed Drone Footage for Kenyan Animal Behavior Analysis
Dataset Summary
This dataset contains the raw, unprocessed drone video footage collected during the creation of the KABR (Kenyan Animal Behavior Recognition) dataset.
Unlike the processed KABR mini-scene dataset which contains extracted video clips with behavioral annotations,
this collection provides the original full-frame drone videos captured at the Mpala Research… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/KABR-raw-videos.gs-videos-v2gs-videos-v3keystroke-typing-videos
Keystroke Typing Videos of Reuters
Recordings of typing randomly sampled sentences (<= 150 characters) from nltk Reuters dataset. Keystroke data is provided too.
AVQA-videos
AVQA — Audio-Visual Question Answering (videos + annotations)
A drop-in package of the AVQA dataset (Yang et al., ACM MM 2022): real-life
audio-visual question answering over short in-the-wild clips. The original release
ships only the QA annotations and expects users to collect the source videos from
VGGSound themselves. This repository bundles the source video clips together
with the official train/val annotations, so the dataset is usable without any
YouTube scraping.… See the full description on the dataset page: https://huggingface.co/datasets/juyil/AVQA-videos.Repaired_videostiktok-videos-4b
TikTok Videos: 4.5 billion posts dataset
Step-by-step guide and access to the scraper code:
tiktok-api.seeksocial.io.
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/kwakuobeng/tiktok-videos-4b.LSVQ-videosThis is an unofficial copy of the videos in the LSVQ dataset (Ying et al, CVPR, 2021), the largest dataset available for Non-reference Video Quality Assessment (NR-VQA); this is to facilitate research studies on this dataset given that we have received several reports that the original links of the dataset is not available anymore.
See FAST-VQA (Wu et al, ECCV, 2022) or DOVER (Wu et al, ICCV, 2023) repo on its converted labels (i.e. quality scores for videos).
The file links to the labels in… See the full description on the dataset page: https://huggingface.co/datasets/teowu/LSVQ-videos.Videos_FL_0216tiktok-videos-4b
TikTok Videos: 4.5 billion posts dataset
Step-by-step guide and access to the scraper code:
tiktok-api.seeksocial.io.
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/blaccastro/tiktok-videos-4b.Sci-VBench-Videos
Sci-VBench Videos
Sci-VBench Videos is the complete set of model outputs behind the Sci-VBench paper: 11,216 videos from 16 text-to-video models, together with the automatic and human scores computed on them. Every video was generated from the verbatim benchmark prompt under the model's default configuration — no rewriting, no prompt expansion — so the released prompts and the released videos correspond exactly.
Prompts and evaluation specifications live in the companion repo… See the full description on the dataset page: https://huggingface.co/datasets/Sci-VBench/Sci-VBench-Videos.generated-videostiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/dams2005/tiktok-videos-4b.videosvideo-SALMONN_2_testset
video-SALMONN 2 Benchmark
Generate the caption corresponding to the video and the audio with video_salmonn2_test.json
Organize your results in the format like the following example:
[
{
"id": ["0.mp4"],
"pred": "Generated Caption"
}
]
Replace res_file in eval.py with your result file.
Run python3 eval.pytiktok-videos-4b
TikTok Videos: 4.5 billion posts dataset
Step-by-step guide and access to the scraper code:
tiktok-api.seeksocial.io.
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289… See the full description on the dataset page: https://huggingface.co/datasets/KOM-00/tiktok-videos-4b.longvideo_eval_videos
Long-RL: Scaling RL to Long Sequences (Evaluation Dataset - for research only)
Data Distribution
We strategically construct a high-quality dataset with CoT annotations for long video reasoning, named LongVideo-Reason. Leveraging a powerful VLM (NVILA-8B) and a leading open-source reasoning LLM, we develop a dataset comprising 52K high-quality Question-Reasoning-Answer pairs for long videos. We use 18K high-quality samples for Long-CoT-SFT to initialize… See the full description on the dataset page: https://huggingface.co/datasets/LongVideo-Reason/longvideo_eval_videos.Videos_FL_0225VideoSIAH-Eval
VideoSIAH-Eval
Evaluation benchmark for LongVT, containing 652 unique QA pairs across 244 long-form videos with human-in-the-loop validation.
Update (2026-03): The initial release contained 1,280 entries due to unintentional duplication during data export. This version has been cleaned to 652 unique QA pairs. Since each entry was an exact copy, all evaluation metrics reported in the paper remain unchanged.
Videos_FL_0204repo-to-space-example-videos
Gradio Space Example Inputs — Videos
A small, curated, freely-licensed pool of videos used as gr.Examples for
Gradio Spaces that wrap video-input generation models (image-to-video,
video-to-video, motion controls, etc.). Sister dataset for images:
linoyts/repo-to-space-example-inputs.
When a Space takes video input, the agent building the Space picks 2–3 clips
whose caption + categories match the model's task, downloads them via
hf_hub_download, runs any model-specific… See the full description on the dataset page: https://huggingface.co/datasets/linoyts/repo-to-space-example-videos.test-HunyuanVideo-pixelart-videos
trojblue/test-HunyuanVideo-pixelart-images
👋 Heads up—this repository is just a PARTIAL dataset. For the full pixelart-images dataset, make sure to grab both parts:
Images Part
Video Part (this repo)
What's in the Dataset?
This dataset is all about anime-styled pixel art images that have been carefully selected to make your models shine. Here’s what makes these images special:
Rich in detail: Pixelated, yes—but still full of life and not overly simplified.… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/test-HunyuanVideo-pixelart-videos.XMER-VideosDepth-Normal-Videos-42K
Depth and Normal Videos Dataset
42,498 videos with depth and surface normals.
Usage
from huggingface_hub import hf_hub_download
video = hf_hub_download(
repo_id="Yanbin99/Depth-Normal-Videos-42K",
filename="Depth_and_Normal_42K/group_0000/videos/00000000.mp4",
repo_type="dataset"
)
tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/alex12223322/tiktok-videos-4b.factory-manipulation-videos
Factory manipulation videos
Procedural Robotics is open sourcing a small set of our factory data so teams can assess its quality. The videos show workers performing factory tasks.
Contents
Seven continuous takes, 109 minutes in total.
Task
Station
Worker
Duration
File
cardboard manipulation
01
041
23.6 min
cardboard_manipulation_station01_worker041.mp4
cardboard manipulation
04
026
16.5 min
cardboard_manipulation_station04_worker026.mp4
defect… See the full description on the dataset page: https://huggingface.co/datasets/procedural-robotics/factory-manipulation-videos.tiktok-videos-4b
TikTok Videos: 4.5 billion posts with engagement metrics
4.5 billion TikTok video records with captions, engagement counts, sound
identifiers and timing. Collected from TikTok's mobile API over roughly three
weeks. Every content_id appears exactly once.
This is the largest public TikTok dataset I am aware of. It is released as-is,
for research.
What is in it
27 Parquet files, zstd compressed, about 289 GB in total. One row per video.
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/hojj/tiktok-videos-4b.wrbench-videos
WRBench Benchmark Videos
The current public dataset contains 11,100 model-output videos. The paper table
is the frozen 23-model, 9,600-output paper_main_23model_9600_20260608 surface.
The 2026-07-13 update synchronizes D5/D6 for the 2,073 applicable paper rows
and D3-D6 for 59 EasyAnimate rolling rows with the published aggregate tables.
Only the three videos_master index representations change. Video bytes,
first frames, prompts, IDs, applicability masks, and frozen paper… See the full description on the dataset page: https://huggingface.co/datasets/WRBench/wrbench-videos.youtube_videostiktok-videos-users-info
TikTok Data, post + poster (user) info, ~1 Million
Dataset name: EinzzCookie/tiktok-videos-users-info
This dataset contains a large collection of TikTok video records paired with detailed creator/user information, stored in a single Parquet file (tiktok_video_user_data.parquet, ~3.87 GB). It is derived from TikTok’s internal video (“aweme”) data model and includes both post-level metadata/engagement stats and nested author profile data.
Source
Collected by… See the full description on the dataset page: https://huggingface.co/datasets/EinzzCookie/tiktok-videos-users-info.
