datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GameplayQA
GameplayQA: A Decision-Dense POV-Synced Multi-Video
Understanding Benchmark of 3D Virtual Agents
Yunzhe Wang
Runhui Xu
Kexin Zheng
Tianyi Zhang
Jayavibhav N. Kogundi
Soham Hans
Volkan Ustun
University of Southern California
ACL 2026
Corresponding Author: yunzhewa@usc.edu
Overview
GameplayQA is the first benchmark for POV-Synced Multi-Video Understanding and… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/GameplayQA.Gameplay-Walkthrough-QADiogenes_Gameplay_raw_sample_v01
Diogenes Gameplay Raw Sample v01
Formerly DiogenesLab/Diogenes_COD_sample_v01 — old links redirect here.
A sample dataset. PC gameplay recordings with frame-aligned keyboard/mouse action
annotations, in two batches:
batch
recorded
video
audio
batch 1
2026-07-25/26
1920×1080 @ 30 fps, H.264
none
batch 2
2026-07-31
1920×1080 @ 60 fps, H.264
process-loopback system audio, 48 kHz stereo s16le (zstd-compressed PCM)
This is a sample — the recording output available… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesLab/Diogenes_Gameplay_raw_sample_v01.footsies-gameplay
Footsies Dataset
Game frames from FOOTSIES game.
{episode}_{step}_{p1}_{p2}_{v1}_{v2}.png
can ignore p2, v1 and v2
Dataset Structure
data/train-00000-of-00001.parquet: Metadata
data/images/: PNG frames
Columns
Column
Type
Description
image
str
Path to image file
episode
int
Episode number (0-299)
step
int
Step within episode
p1_action
int
Player 1 action (0=none, 1=back, 2=forward)
p2_action
int
Player 2 action
Statistics… See the full description on the dataset page: https://huggingface.co/datasets/H1yori233/footsies-gameplay.2048-gameplay-dataset
2048 Expert Gameplay Dataset
State-action pairs from an expert N-Tuple Network agent playing 2048.
Can be used for imitation learning / supervised training of 2048 agents.
Stats
Source games: 10,000
Games after filtering: 9,000
Total moves: 54,010,983
Average score: 143,847
Win rate (>= 2048): 100% (losing games removed)
Score floor: 62,152 (bottom 10% removed)
Files
train.jsonl - 8,100 games (48,592,790 moves) for training
val.jsonl - 900 games (5,418,193… See the full description on the dataset page: https://huggingface.co/datasets/yethdev/2048-gameplay-dataset.action-conditioned-gameplay-10000h
Action-Conditioned Gameplay Dataset
10,000 hours of rights-cleared gameplay trajectories with synchronized video, keyboard/mouse/controller inputs, camera motion, player state, object state, events, goals, rewards, and outcomes for world models and AI agents.
This repository contains the full technical specification, annotation schema, and sample metadata files (Parquet). The production dataset is rights-cleared and delivered directly to buyers. Request access to see the full… See the full description on the dataset page: https://huggingface.co/datasets/Datoric/action-conditioned-gameplay-10000h.Diogenes_Gameplay_multigame_sample_v01
Diogenes Gameplay Multigame Sample v01
Human gameplay recordings across 19 PC titles, captured with the same
recording pipeline and the same per-frame action-annotation schema on every title.
The point of this sample is breadth: one session per title, identical schema,
per-title semantic action vocabularies (mappings/), and honest per-segment QC
columns computed from the shipped data itself.
19 sessions · 603 segments · 1.72 h · 205,091 annotated frames · 332,572 raw input… See the full description on the dataset page: https://huggingface.co/datasets/DiogenesLab/Diogenes_Gameplay_multigame_sample_v01.
