datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
standard-chess-games
[!CAUTION]
This dataset is still a work in progress and some breaking changes might occur.
Lichess Rated Standard Chess Games Dataset
Dataset Description
6,771,826,271 standard rated games, played on lichess.org, updated monthly from the database dumps.
This version of the data is meant for data analysis. If you need PGN files you can find those here. That said, once you have a subset of interest, it is trivial to convert it back to PGN as shown in the Dataset Usage… See the full description on the dataset page: https://huggingface.co/datasets/Lichess/standard-chess-games.IG-10K-Dataset
The Imitator Game - IG-10K Dataset
This dataset accompanies the paper The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction.
It contains paired human-robot demonstrations for the Imitator Game benchmark, spanning four levels of imitation difficulty (L0–L3) across real and simulated settings. The IG-10K dataset includes over 20,000 paired episodes across 50+ tasks and 6 domains, and is provided in LeRobot-0.5.0 format.
For more details, see the project… See the full description on the dataset page: https://huggingface.co/datasets/imitator-game/IG-10K-Dataset.strategic_game_chess
Chess
Recent advancements in artificial intelligence (AI) underscore the progress of reasoning and planning shown by recent generalist machine learning (ML) models. The progress can be boosted by datasets that can further boost these generic capabilities when used for training foundation models of various kind. This research initiative has generated extensive synthetic datasets from complex games — chess, Rubik's Cube, and mazes — to study facilitation and the advancement of these… See the full description on the dataset page: https://huggingface.co/datasets/laion/strategic_game_chess.game_character_skins
Game Character Skins Dataset
Summary
This comprehensive dataset contains game character skins and artwork from multiple popular mobile and PC games, providing a rich collection of character visual assets for computer vision research and game development applications. The dataset spans eight major game titles including Arknights, Azur Lane, Blue Archive, Fate/Grand Order, Genshin Impact, Girls' Frontline, Neural Cloud, Nikke, Path to Nowhere, and Honkai: Star Rail… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/game_character_skins.gamesGame compositions created by users
strategic_game_mazeNOTICE: some of the game is mistakenly label as both length and width columns are 40, they are 30 actually.
maze
This dataset contains 350,000 mazes, represents over 39.29 billion moves.Each maze is a 30x30 ASCII representation, with solutions derived using the BFS.
It has two columns:
'Maze': representation of maze in a list of string.shape is 30*30
visual example
'Path': solution from start point to end point in a list of string, each item represent a position in the maze.
course-imagesfaience-games
Faïence: human-vs-net Azul games
Every game played on Faïence, a
free browser implementation of the rules of Azul (Michael Kiesling) against
a neural net trained by self-play, unless the player switched sharing off.
This dataset is the training pile the playing page tells its players about,
and it is public precisely so that a player can read everything the project
collects. Records are anonymous by construction: moves, deals, which net
played, and the score. No names, no… See the full description on the dataset page: https://huggingface.co/datasets/RemiFabre/faience-games.rrflux-game24-game
Math Twenty Four (24s Game) Dataset
A comprehensive dataset for the classic math twenty four game (also known as the 4 numbers game / 24s game / Game of 24). This dataset of mathematical reasoning challenges was collected from 4nums.com, featuring over 1,300 unique puzzles of the Game of 24, with difficulty metrics derived from over 6.4 million human solution attempts since 2012.
In each puzzle, players must use exactly four numbers and basic arithmetic operations (+, -, ×, /) to… See the full description on the dataset page: https://huggingface.co/datasets/nlile/24-game.nemotron-3-nano-30b-20260719-spare-games-envs
Nemotron-3-Nano-30B SPARE Self-Play Environments (run_20260719_final)
This dataset packages the self-play generated game environments produced
by a live SPARE (Self-Play with Adaptive cuRriculum Extension) training run
of NVIDIA-Nemotron-3-Nano-30B-A3B. It is a raw-data export for another
agent to pick up, replay, and build its own visualization / weave log from.
Provenance
Run: run_20260719_final
Source Ray job: spare_nemotron_games_mtpg768_1784556397 (the live… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/nemotron-3-nano-30b-20260719-spare-games-envs.evaluation_logs
Evaluation logs from "Auditing Games for Sandbagging"
This dataset provides evaluation transcripts produced for the paper "Auditing Games for Sandbagging". Transcripts are provided in Inspect .eval format, see https://github.com/AI-Safety-Institute/sabotage_games for a guide to viewing them.
Dataset Details
evaluation_transcripts/handover_evals contains the transcripts provided by the red team to the blue team at the beginning of the main round of the game, showing… See the full description on the dataset page: https://huggingface.co/datasets/sandbagging-games/evaluation_logs.gamewam-vizdoom
GameWAM ViZDoom APPO Dataset
This repository contains 50,000 policy-generated gameplay trajectories from
four ViZDoom combat scenarios. The trajectories were collected for
GameWAM: A World Action Model for Video Games
by running pretrained Sample Factory APPO policies in the native ViZDoom
simulator. They are agent rollouts rather than human demonstrations.
Paper: https://huggingface.co/papers/2608.26200
Project page: https://yunncheng.github.io/GameWAM/
Code:… See the full description on the dataset page: https://huggingface.co/datasets/Yunncheng/gamewam-vizdoom.data_gameGameQA-140K
🎊 News
[2026/07] 🔥Peking University and Kuaishou Kling Team evaluate their agentic visual reasoning method Beacon on our GameQA benchmark. Beacon learns when tools are truly needed (Mode Adaptiveness) and how tool use extends capability on hard problems (Tool Effect), and achieves the highest accuracy on GameQA among open-source models of the same scale, significantly outperforming its Qwen3-VL-8B-Instruct base.
[2026/07] 🔥Peking University and WeChat AI use our Game-RL data… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/GameQA-140K.game_characters
Database of Characters in Mobile Games
All the character in the following games are supported:
Arknights (crawled from https://prts.wiki)
Fate/Grand Order (crawled from https://fgo.wiki)
Azur Lane (crawled from https://wiki.biligame.com/blhx)
Girls' Front-Line (crawled from https://iopwiki.com/)
Genshin Impact (crawled from https://genshin-impact.fandom.com/ja/wiki/%E5%8E%9F%E7%A5%9E_Wiki)
The source code and python library is hosted on narugo1992/gchar, and the scheduled job is… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/game_characters.strategic_game_cube
Cube
This dataset contains 1.64 billion Rubik's Cube solves, totaling roughly 236.39 billion moves.it is generated by Fugaku using https://github.com/trincaog/magiccube
Each solve has two columns: 'Cube' and 'Actions',
'Cube': initial scrambled states of a 3-3-3 cube in string, such as:
WOWWYOBWOOGWRBYGGOGBBRRYOGRWORBBYYORYBWRYBOGBGYGWWGRRY
the visual state of this example is
NOTICE: Crambled Cube States are spread out into the above string, row by row.
'Actions': list of… See the full description on the dataset page: https://huggingface.co/datasets/laion/strategic_game_cube.retro-games-gameplay-frames-30k-512pGamePhysicsDailyDump
GamePhysics Dataset (Daily Dump)
c2c-ai-vs-ai
C2C: Cooperate to Compete — AI vs AI Games
This dataset contains 972 fully-logged AI vs AI games from the Cooperate to Compete (C2C) benchmark — a long-horizon, mixed-motive multi-agent negotiation environment based on a four-player conquest game with private regional objectives, fog of war, and non-binding cheap-talk negotiation.
Project page: https://negotiationgame.io/c2c/
Paper: https://arxiv.org/abs/2604.25088
Play against AI agents: https://negotiationgame.io
Github:… See the full description on the dataset page: https://huggingface.co/datasets/negotiation-games/c2c-ai-vs-ai.steam-games-dataset
Steam Games Dataset
Information of 141,900 games published on Steam.
This dataset has been created with this code (MIT) and use the API provided by Steam, the largest gaming platform on PC. Data is also collected from Steam Spy. Only published games, no DLCs, episodes, music, videos, etc.
Maintained by Fronkon Games.
mind-games-dataSPADE-Environments-Qwen3-30B-Games
SPADE generated environments: games
Paper | Code | All artifacts
Executable game environments written by the SPADE Environment Designer during the paper's 30B games self-play run. One Python file per environment; manifest.json records the generation checkpoint, training step, skill, and difficulty of each.
Environments
3310
Training steps covered
113 (step 0 to 396)
With skill label
3119
Designer / agent model
Qwen/Qwen3-30B-A3B-Instruct-2507… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environments-Qwen3-30B-Games.lsat_logic_games-analytical_reasoningNovel annotated evaluation dataset of LSAT logic games associated with paper:
Lost in the Logic: An Evaluation of Large Language Models’ Reasoning Capabilities on LSAT Logic Games
Arxiv: http://arxiv.org/pdf/2409.19012
If you find this dataset useful, please cite the paper!
@misc{malik2024lostlogicevaluationlarge,
title={Lost in the Logic: An Evaluation of Large Language Models' Reasoning Capabilities on LSAT Logic Games},
author={Saumya Malik},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saumyamalik/lsat_logic_games-analytical_reasoning.Nexora-Game-Datagame-data-anomaly-samples
Game-data quality — CORRECTED analysis (controller / uncaptured-input finding)
TL;DR
Many sessions that the first pass called "completely idle" are not idle. They were
played with a controller/gamepad (or are cutscenes / auto-path), which the
keyboard+mouse capture tool never recorded. The video shows full gameplay while the
action labels are empty — poison for keyboard+mouse behaviour cloning.
Proof (胡宸 / Monster Hunter World)
parquet actions: 18… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuanhucs/game-data-anomaly-samples.NBA_Games
NBA Full-Game Video Dataset
This dataset provides metadata, official statistics, and official play-by-play annotations for full-length NBA game videos available on YouTube. Instead of redistributing video files, we provide YouTube video IDs and URLs so users can download videos independently when their use case and local policies allow it.
The dataset links long-form basketball videos with structured NBA.com game data. Each retained game has a verified… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/NBA_Games.gamewam-minecraft
GameWAM Minecraft Datasets
This repository contains three Minecraft gameplay datasets used by
GameWAM: A World Action Model for Video Games.
All three use LeRobot v2.1 format and
share a 22-D keyboard/mouse action space, a 6-D raw state (pitch, yaw, cursor
x/y, hotbar, and GUI-open state), and a 15-D canonical proprioceptive
representation.
Project page: https://yunncheng.github.io/GameWAM/Code: https://github.com/yunncheng/GameWAM
Dataset
Directory
Episodes
Frames… See the full description on the dataset page: https://huggingface.co/datasets/Yunncheng/gamewam-minecraft.nba-games
NBA Games Data
This data is an updated version of the original NBA
Games by Nathan Lauga.
Data source
Code
Updated to: 2025-02-13
The dataset retains the original format and includes the following files:
games.csv – Summary of NBA games, including scores and team details.
games_details.csv – Detailed player statistics for each game.
players.csv – Player information.
ranking.csv – Daily NBA team rankings.
teams.csv – List of all NBA teams.
SPADE-Environment-Pool-GPT5.5-Games
SPARE GPT-5.5 Grounded Cognitive Multi-Turn Games
This public dataset contains 7,872 validated Python game environments for actor-only SPARE training.
Six cognitive skills, exactly 1,312 environments per skill
Generated with GPT-5.5 and grounded by spice_megascience_15k.jsonl
Grounding corpus SHA-256: a36a928b4940b5b5d9e3f4cb5804a94c69462360943adb3be14613c82f0f72c0
Maximum 25 turns and 32K generation context
Every environment passes load, reset, step, and replay validation with… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environment-Pool-GPT5.5-Games.
