datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Collective-Activity-Recognition
Annotation Format
Every 10th frame in all video sequences was manually annotated with the following information for each detected person:
Bounding box location
Activity class
Pose direction
Annotation Fields
Each annotation follows the format:
<frame_number> <x> <y> <width> <height> <class_id> <pose_id>
Field
Description
frame_number
Frame identifier
x
X-coordinate of the bounding box (top-left corner)
y
Y-coordinate of the bounding box (top-left… See the full description on the dataset page: https://huggingface.co/datasets/litforth/Collective-Activity-Recognition.truth-probe-activationsxiaoluo-gaming-action3000-20260910-media
Action-boundary review examples
Media for 300 selected examples from Xiaoluo (Cyberpunk 2077 and Rise of the Tomb Raider) and Gaming 500 Hours, 150 examples per dataset.
Includes 15-second review videos, observed boundary frames, and available action clips. These are visual model estimates; boundaries require human review. Source game and dataset rights remain with their respective owners.
Gallery and annotation manifests:… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/xiaoluo-gaming-action3000-20260910-media.ActiveVision
ActiveVision — An Exam for Active Observers
ActiveVision is a benchmark for iterative visual reasoning: 85 photorealistic
items across 17 tasks that cannot be solved from a single glance — the model has
to keep returning to the image to scan, trace, and compare. Every scene is
generated by a deterministic program and re-rendered photorealistically while
preserving the structure, so answers are exact by construction.
Frontier models reach about 10% with pure… See the full description on the dataset page: https://huggingface.co/datasets/activevisionai/ActiveVision.pythia-massive-activations
Hidden Dynamics of Massive Activations in Transformer Training
Dataset Description
This dataset contains comprehensive analysis data for the paper "Hidden Dynamics of Massive Activations in Transformer Training". It provides detailed measurements and mathematical characterizations of massive activation emergence patterns across the Pythia model family during training.
Massive activations are scalar values in transformer hidden states that achieve values orders of… See the full description on the dataset page: https://huggingface.co/datasets/Aimpoint-Digital/pythia-massive-activations.afrolm_active_learning_dataset
AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages
GitHub Repository of the Paper
This repository contains the dataset for our paper AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages which will appear at the third Simple and Efficient Natural Language Processing, at EMNLP 2022.
Our self-active learning framework
Languages Covered
AfroLM has been… See the full description on the dataset page: https://huggingface.co/datasets/bonadossou/afrolm_active_learning_dataset.Act2Cap_benchmarkCollected data from GUI-Action-Narrator
SAFER-Activities
SAFER-Activities
A Dataset for Smart Assessment of Fall Events and Routine Activities (ECCV 2026).
SAFER-Activities is a dataset for fall detection and physical activity monitoring from
video, with a dedicated subset for wheelchair users. It provides frame-level action
annotations (precise start/end of every action) over long, untrimmed, multi-camera
recordings.
File structure
.
├── raw/
│ ├── Safer-Activities-Full-Dataset/
│ │ ├── normal/… See the full description on the dataset page: https://huggingface.co/datasets/SAFER-Activities/SAFER-Activities.fire_actioncam
Fire Actioncam
This dataset is a collection of several real-world fire scenes, introduced by the ECCV paper "Gaussians on Fire: High-Frequency Reconstruction of Flames".
Overview
The dataset consists of 17 real-world scenes of burning paper, cardboard, wood, gasoline, ethanol, and propane. We captured each scene with three regular actioncams, synchronizing them with µs precision using a custom LED pattern.
Property
Value
Scenes
17 (two outdoor… See the full description on the dataset page: https://huggingface.co/datasets/jna-358/fire_actioncam.actionbench
🎬 ActionBench: Paired Video-3D Synthetic Benchmark
📖 Overview
ActionBench is a benchmark dataset of 128 paired video ↔ animated point-cloud samples for evaluating animated 3D mesh generation from video.
The dataset consists of synthetic scenes of animated objects from ObjaverseXL, rendered using Blender 3.5.1.
Each sample contains:
Video: 16 RGBA frames with alpha mask
Camera (camera.json): Camera parameters using Blender convention (X_cam = X @ R^T + T, camera looks… See the full description on the dataset page: https://huggingface.co/datasets/facebook/actionbench.rlbench_joint_vel_action_lerobot_trainThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 1800,
"total_frames": 375567,
"total_tasks": 202,
"total_videos": 0,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:1800"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/daixianjie/rlbench_joint_vel_action_lerobot_train.themoviedb_actorsaction-world-model-atlas-1500-media-20260914
Action World Model Atlas
Public browsing previews for 1,500 unique action clips from the completed
6,033-video bundle. OpenPixel2Play, Gaming 500 Hours, and Xiaoluo each contribute
500 examples. All 46 games in the completed bundle are represented.
Videos preserve the full five-second duration and 81 frames. They are existing
browser previews and can be smaller than the native training videos. Video and
poster checksums are verified against the source media manifests.… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/action-world-model-atlas-1500-media-20260914.robot-action-prediction-dataset
Robotic Action Prediction Dataset
Dataset Description
This dataset contains triplets of (current observation, action instruction, future observation) for training models to predict future frames of robotic actions.
Dataset Structure
Data Fields
current_frame: Input image (RGB) of the current observation
instruction: Textual description of the action to perform
future_frame: Target image (RGB) showing the expected outcome 50 frames later… See the full description on the dataset page: https://huggingface.co/datasets/bryandts/robot-action-prediction-dataset.peg_04_16_cam0and1_cam_action_no_cropThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 65,
"total_frames": 31241,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:65"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/radoolonto/peg_04_16_cam0and1_cam_action_no_crop.dual_needle_concat_action_staticThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unity",
"total_episodes": 250,
"total_frames": 97953,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:250"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/inaas/dual_needle_concat_action_static.Human_Action_Recognition
Dataset Summary
A dataset from kaggle. origin: https://dphi.tech/challenges/data-sprint-76-human-activity-recognition/233/data
Introduction
The dataset features 15 different classes of Human Activities.
The dataset contains about 12k+ labelled images including the validation images.
Each image has only one human activity category and are saved in separate folders of the labelled classes
PROBLEM STATEMENT
Human Action Recognition (HAR) aims to understand… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/Human_Action_Recognition.Active-ReconstructionActionEQA
ActionEQA: Action Interface for Embodied Question Answering
Tianwei Bao1* · Qineng Wang1* · Kangrui Wang1 · Mingkai Deng2 · Guangyi Liu5 · Jiayuan Mao3
Larry Birnbaum1 · Zhiting Hu4 · Eric P. Xing2,5 · Zhaoran Wang1 · Manling Li1
1 Northwestern University 2 Carnegie Mellon University 3 UPenn
4 UC San Diego 5 MBZUAI
* Equal contribution
ActionEQA is the first action-centric Embodied Question Answering (EQA) benchmark designed to systematically evaluate… See the full description on the dataset page: https://huggingface.co/datasets/TianweiBao/ActionEQA.drive-actionactionnet_3k_og
actionnet_3k_og — training subset, archived
Everything actionnet_grasp_2b_480 / _720 opens, plus the LeRobot data/ and meta/
that describe the same episodes: ~35.0 GB in 9 archives instead of ~14,900 loose files.
A 13th archive set, videos_15fps_768x432/, comes from a different store — read the
note below before using it.
This is still a subset of a larger tree. The source also holds videos/,
first_frames/, grasp_frames/, source_mask_merge/, rendering_videos/, gtdepth_s0/
and… See the full description on the dataset page: https://huggingface.co/datasets/seungkukim/actionnet_3k_og.video-web-actionwifi-csi-human-activities
WiFi CSI Human Activities Dataset
This repository contains WiFi Channel State Information (CSI) measurements, intermediate processing outputs, extracted features, visualizations, and experimental results used for WiFi-based Human Activity Recognition (HAR).
The dataset accompanies the Master's thesis:
Assel Ussenova
WiFi Sensing Through Digital Receive Beamforming and CSI
MSc in ICT and Internet Engineering
Università degli Studi di Roma Tor Vergata (2024/2025)… See the full description on the dataset page: https://huggingface.co/datasets/aselya9185/wifi-csi-human-activities.openp2p-action-clips-media-3000-20260911activity-diagrams-qdobr
Dataset Card for activity-diagrams-qdobr
** The original COCO dataset is stored at dataset.tar.gz**
Dataset Summary
activity-diagrams-qdobr
Supported Tasks and Leaderboards
object-detection: The dataset can be used to train a model for Object Detection.
Languages
English
Dataset Structure
Data Instances
A data point comprises an image and its object annotations.
{
'image_id': 15,
'image': <PIL.JpegImagePlugin.JpegImageFile… See the full description on the dataset page: https://huggingface.co/datasets/Francesco/activity-diagrams-qdobr.ours_joint_actionjam-actions-v0
Dataset Card for jam-actions-v0 (public subset)
Version: 0.5.1 — a documentation-only revision of the 0.5.0 record cut. No record, split or eval artifact changed; the card gained the fine-tuning evaluation banner and the "What's in a record" walkthrough, which had been added on Hugging Face and lived nowhere else.
Records built: 2026-07-11 Source tag: jam-actions-v0-0.5.0-cut-2026-07-11 (record-content correction release — Bach BWV 846 errata 001 + 002; see RELEASE_NOTES.md… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v0.sinhala-ocr-lk-acts-1010
🇱🇰 Sinhala OCR - Sri Lankan Acts Dataset
Dataset Description
This dataset contains 1,010 scanned document images of Sri Lankan legal acts (1980s-2010s) in Sinhala language with ground truth text annotations for Optical Character Recognition (OCR) training and evaluation.
Key Features
✅ High-quality scanned document images
✅ Professionally corrected ground truth text
✅ Year-wise metadata for temporal analysis
✅ Pre-split into train/eval/test sets… See the full description on the dataset page: https://huggingface.co/datasets/avishadilhara/sinhala-ocr-lk-acts-1010.dual_needle_concat_action_two_wristsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unity",
"total_episodes": 250,
"total_frames": 97953,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:250"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/inaas/dual_needle_concat_action_two_wrists.shotpath-action-diagnostic-venuslike-eval-20260709# ShotPath Action Diagnostic Venus-like Eval 20260709
This bundle contains the LLM-audited pure-operation diagnostic set for Venus-like evaluation.
Files:
action_diagnostic_pure_operation.jsonl: 908 examples after leakage audit.
images/: image files referenced by the jsonl.
scripts/eval_action_diagnostic_qwen25vl.py: Qwen2.5-VL base/LoRA evaluator.
scripts/run_action_diagnostic_venuslike_eval_server.sh: server runner for base7b, stage1 step200, stage2 step200.
Default server paths in the… See the full description on the dataset page: https://huggingface.co/datasets/purefall/shotpath-action-diagnostic-venuslike-eval-20260709.
