datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
videos-testGoPro-Raw-Videos
Raw GoPro Videos for Four Robotic Manipulation Tasks
[Project Page]
[Paper]
[Code]
[Models]
[Processed Dataset]
This repository contains raw GoPro videos of robotic manipulation tasks collected in-the-wild using UMI, as described in the paper "Data Scaling Laws in Imitation Learning for Robotic Manipulation". The dataset covers four tasks:
Pour Water
Arrange Mouse
Fold Towel
Unplug Charger
Dataset Folders:
arrange_mouse and pour_water: Each folder contains data… See the full description on the dataset page: https://huggingface.co/datasets/Fanqi-Lin/GoPro-Raw-Videos.code-world-model-project-page-videos
Code World Model Project Page Videos
Public research-demo video assets used by the Code World Model project page.
The gallery/ directory contains aligned RGB and proxy videos for interactive comparison.
Egoverse_videosqvhighlights-videos
QVHighlights Videos
All 25124 videos from the QVHighlights benchmark (train + val + test splits).
Total size: 135.3 GB.
Layout
Files are sharded into subdirectories by the first character of the filename
(HuggingFace caps each directory at 10,000 files):
<first-char>/<youtube-id>_<start>_<end>.mp4
Source
Lei et al., "QVHighlights: Detecting Moments and Highlights in Videos via Natural
Language Queries" (NeurIPS 2021).
Original archive:… See the full description on the dataset page: https://huggingface.co/datasets/ayushsdev/qvhighlights-videos.test-videosVBench-2.0_sampled_videos
Sample Videos of VBench-2.0
This dataset is used in the paper:👉 arXiv:2503.21755
hoigen-filtered-videos
HOIGen Filtered Videos Dataset
This dataset contains 28562 filtered videos from the HOIGen-1M dataset based on the allowlist.
Dataset Structure
The videos are organized in the same structure as the original HOIGen dataset:
filtered_videos/
├── videos_part_1/
├── videos_part_2/
├── ...
└── videos_part_100/
Usage
from huggingface_hub import hf_hub_download
# Download a specific video
video_path = hf_hub_download(… See the full description on the dataset page: https://huggingface.co/datasets/charlychan123/hoigen-filtered-videos.gs-videos-v2bdd100k_videos
BDD100K videos
This repository stores the original bdd100k_videos.zip as byte-for-byte split files. The archive is not recompressed. Refer to the BDD100K license and terms before using or redistributing the data.
Restore
Download all bdd100k_videos.zip.part-* files from bdd100k_videos/, then run:
cat bdd100k_videos.zip.part-* > bdd100k_videos.zip
md5sum -c bdd100k_videos.zip.md5
Source MD5: 253d9a2f9d89d2b09d8d93f397aecdd7. There are 19 parts of up to 100.00 GiB… See the full description on the dataset page: https://huggingface.co/datasets/linxxx3/bdd100k_videos.action-atlas-rollout-videos
Action Atlas — VLA Rollout Videos
Local rollout/ablation videos for Pi0.5, OpenVLA-OFT, X-VLA, GR00T, SmolVLA, and ACT/ALOHA,
organized by model. Companion to action-atlas-{pi05,oft,xvla,groot,smolvla} (SAEs + activations + concepts).
368283 unique mp4 clips, 61.8 GB. Per-model: {'act_aloha': 990, 'groot': 163891, 'oft': 24284, 'pi05': 62468, 'smolvla': 56844, 'xvla': 59806}
manifest.jsonl: one row per clip (model, env, experiment, sha256, bytes, hf_path).
KABR-raw-videos
Dataset Card for KABR Raw Videos: Unprocessed Drone Footage for Kenyan Animal Behavior Analysis
Dataset Summary
This dataset contains the raw, unprocessed drone video footage collected during the creation of the KABR (Kenyan Animal Behavior Recognition) dataset.
Unlike the processed KABR mini-scene dataset which contains extracted video clips with behavioral annotations,
this collection provides the original full-frame drone videos captured at the Mpala Research… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/KABR-raw-videos.fixtures_videosKABR-mini-scene-raw-videos
Dataset Card for Kenyan Animal Behavior Recognition (KABR) Mini-Scene Raw Videos
Dataset Summary
This dataset is comprised of a collection of 10+ hours of drone videos focused on Kenyan wildlife that contains behaviors of giraffes, plains zebras, and Grevy's zebras.
Animals can be identified with bounding box coordinates provided, and behavior annotations can be recovered by linking the labels back to these bounding boxes from the mini-scene annotations provided in our… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/KABR-mini-scene-raw-videos.Videos4CameraBenchgs-videos-v4video-studio-toolsAVQA-videos
AVQA — Audio-Visual Question Answering (videos + annotations)
A drop-in package of the AVQA dataset (Yang et al., ACM MM 2022): real-life
audio-visual question answering over short in-the-wild clips. The original release
ships only the QA annotations and expects users to collect the source videos from
VGGSound themselves. This repository bundles the source video clips together
with the official train/val annotations, so the dataset is usable without any
YouTube scraping.… See the full description on the dataset page: https://huggingface.co/datasets/juyil/AVQA-videos.brb-traffic-videos_ORIGINAL1keystroke-typing-videos
Keystroke Typing Videos of Reuters
Recordings of typing randomly sampled sentences (<= 150 characters) from nltk Reuters dataset. Keystroke data is provided too.
brb-traffic-videos_ORIGINAL2brb-traffic-videos_ORIGINAL4youtube_videos_1brb-traffic-videos_ORIGINAL3RealEstate10K-videos
RealEstate10K-videos
Research mirror of the raw source videos for the RealEstate10K train split
(project page). The official release ships only
per-clip camera trajectories (.txt, CC-BY-4.0) with YouTube URLs; the videos themselves
must be fetched from YouTube, where a growing fraction is no longer available. This mirror
preserves what was still retrievable as of 2026-09-15.
Contents
Path
Contents
videos/<youtube_id>.mp4
6,085 of 6,559 unique train… See the full description on the dataset page: https://huggingface.co/datasets/Kunho/RealEstate10K-videos.educational_videosVideos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text
Dataset Overview
A collection of 27 domains (“topics”) and 3100 question-answer pair.
Each topic comes with average 117 QA pairs.Every QA entry comes with:
references: one or more source files the answer is extracted from
time with each reference comes the starting and ending time the answer is extracted from the reference
video_files: the video files where the answer can be found
(future) video title & description from metadata.csv
File structure
You-Are-Here!/… See the full description on the dataset page: https://huggingface.co/datasets/elmoghany/Videos-Dataset-For-LLMs-RAG-That-Require-Audio-Vidoes-And-Text.urap26-annotation-videos
URAP26 Annotation Videos
Public 1080p RGB proxies paired with YOLOMG compensated-difference videos for browser annotation.
ufc-crime-videosyoutube_videos_2
