datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
faience-games
Faïence: human-vs-net Azul games
Every game played on Faïence, a
free browser implementation of the rules of Azul (Michael Kiesling) against
a neural net trained by self-play, unless the player switched sharing off.
This dataset is the training pile the playing page tells its players about,
and it is public precisely so that a player can read everything the project
collects. Records are anonymous by construction: moves, deals, which net
played, and the score. No names, no… See the full description on the dataset page: https://huggingface.co/datasets/RemiFabre/faience-games.NBA_Games
NBA Full-Game Video Dataset
This dataset provides metadata, official statistics, and official play-by-play annotations for full-length NBA game videos available on YouTube. Instead of redistributing video files, we provide YouTube video IDs and URLs so users can download videos independently when their use case and local policies allow it.
The dataset links long-form basketball videos with structured NBA.com game data. Each retained game has a verified… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/NBA_Games.car-game-pi-tracesrummlu
ruMMLU
Task Description
Russian Massive Multitask Language Understanding (ruMMLU) is a dataset designed to measure model professional knowledge acquired during pretraining in various fields . The task covers 57 subjects (subdomains) across different topics (domains): HUMANITIES; SOCIAL SCIENCE; SCIENCE, TECHNOLOGY, ENGINEERING, AND MATHEMATICS (STEM); OTHER. The dataset was created based on the English MMLU dataset proposed in the original paper and follows its… See the full description on the dataset page: https://huggingface.co/datasets/gametwix/rummlu.game-design-pattern-core-collection
Game Design Patterns Dataset
Original Source Attribution
This dataset is derived from the work of Staffan Björk and Jussi Holopainen. The original content comes from "Patterns in Game Design," published by Charles River Media in 2005.
Original authors: Jussi Kuittinen, Staffan Björk and Jussi Holopainen
Original format: HTML documents publicly available at: https://www.researchgate.net/publication/379683418_collection.zip
Citation: Bjork, S., & Holopainen, J. (2005).… See the full description on the dataset page: https://huggingface.co/datasets/HughXuechen/game-design-pattern-core-collection.FIFA_World_Cup_Games
FIFA World Cup Full-Match Video Dataset
This dataset provides YouTube full-match references and structured match annotations for classic FIFA World Cup games. Instead of redistributing video files, it stores YouTube video IDs and URLs, then links each retained match to structured football data including match metadata, team statistics, lineups, formations, text event timelines, and pre-match prediction polls.
The dataset is designed for long-form sports… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/FIFA_World_Cup_Games.game_commentary_sftgames-catalog
Council of AI — games catalog
Catalog door. Games load into Council Space. No contest engine on this card.
Live: https://councilof.ai/gspc-arena
Council OS: https://councilof.ai/os
Council Space: https://councilof.ai/gspc-arena
Measurement, not certification. Empty slots are not for sale. No scores on this card.
Jail is a measured floor, not a 16th pane.
The live board is the authority
GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub… See the full description on the dataset page: https://huggingface.co/datasets/csoai/games-catalog.qwen3-4b-0701-tooluse-glory-kl0-spare-games-envs
qwen3-4B-Instruct-0701-tooluse-glory-kl0 — generated environments
Environments generated by the SPARE proposer during training run
223t1pws (qwen3-4B-Instruct-0701-tooluse-glory-kl0), recovered from the spare-viz durable cache.
The run's scratch directory no longer exists; this dataset is the surviving copy.
Games
350
Steps covered
16 (step 0–384)
With recovered skill
350
With hint
0
Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/qwen3-4b-0701-tooluse-glory-kl0-spare-games-envs.threejs-gamecode-instruct-v3-ultra
Three.js GameCode Instruct v3 Ultra
This is a large synthetic/original instruction dataset for training or testing LLM behavior around Three.js, browser game development, gameplay programming, debugging, optimization, architecture, and general coding.
Important note
This dataset is synthetic and programmatically generated from original templates. It is designed as a useful starting point for experiments, not as a fully hand-curated gold-standard benchmark.
No… See the full description on the dataset page: https://huggingface.co/datasets/agagasf123123/threejs-gamecode-instruct-v3-ultra.osworld-realtime-gamesplaydate-games
Coding agent session traces for aaaaliou/playdate-games
This dataset contains redacted coding agent session traces exported with pi-share-hf from a local pi workspace. The traces were filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured session entry. Entries include session headers, user and assistant… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/playdate-games.beat-the-game-minecraft
Mine AI MCP — the run that beat Minecraft
An LLM agent played Minecraft 1.21.4 from an empty world to a defeated Ender Dragon,
autonomously, in a single unbroken session. No human input after the prompt, no
scripted behaviour trees, no save-scumming. This dataset is the complete record of
that run.
📺 Watch the run: https://www.youtube.com/watch?v=ZjtwWEfFVFY
💻 Code: https://github.com/aibengineering/mine-ai-mcp (MIT)
What it cost
Beating Minecraft took $93.17 of… See the full description on the dataset page: https://huggingface.co/datasets/aibengineering/beat-the-game-minecraft.qwen3-30b-0617-6skill-regretonly-spare-games-envs
qwen3-30B-A3B-Instruct-0617-6skill-regretonly — generated environments
Environments generated by the SPARE proposer during training run
a1s4s63z (qwen3-30B-A3B-Instruct-0617-6skill-regretonly), recovered from the spare-viz durable cache.
The run's scratch directory no longer exists; this dataset is the surviving copy.
Games
158
Steps covered
6 (step 0–138)
With recovered skill
128
With hint
0
Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/qwen3-30b-0617-6skill-regretonly-spare-games-envs.SPADE-Grounding-Corpus-Games-15K
SPADE grounding corpus — games (15k)
Reference documents the SPADE proposer is grounded on when generating cognitive-skill
game environments. 15,000 documents: 10k drawn from a mathematics corpus and 5k from a
science corpus.
Documents
15,000
Setting
games
Fields
Field
Description
text
The document, exactly as embedded in the generation prompt
metadata
domain (mathematics / science) and url (source provenance)
Each generation… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Grounding-Corpus-Games-15K.SKYLENAGE-GameCodeGym
V-GameGym: Visual Game Generation for Code Large Language Models
Abstract
Code large language models have demonstrated remarkable capabilities in programming tasks, yet current benchmarks primarily focus on single modality rather than visual game development. Most existing code-related benchmarks evaluate syntax correctness and execution accuracy, overlooking critical game-specific metrics such as playability, visual aesthetics, and user engagement that are… See the full description on the dataset page: https://huggingface.co/datasets/alibabagroup/SKYLENAGE-GameCodeGym.photon47-game-introduction
Photon 47 game-introduction documents
Verified bilingual release, 2026-09-02.
Photon47_Game_Introduction_EN-ZH.pdf is the canonical 16-page A4
submission document. Its cover, pagination, images, English text and Chinese
text were visually reviewed after a browser-native PDF render.
Photon47_Game_Introduction_EN-ZH.docx is the editable companion. It was
reopened and round-tripped through LibreOffice; English, Chinese and the
final provenance section remained readable.… See the full description on the dataset page: https://huggingface.co/datasets/ryan-superman/photon47-game-introduction.gamedev_alpacaSpurious-Token-Game
🧩 Spurious Token Game (STG)
The Spurious Token Game (STG) dataset contains two subtasks designed for evaluating models under spurious correlations.
Subtask: STG_E
Training splits: STG_S, STG_M, and STG_L (representing different data sizes or difficulty levels)
Test splits: IID (in-distribution) and OOD (out-of-distribution)
Subtask: STG_H
Training split: single training dataset
Test splits: IID and OOD
💡 Example Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Kairong-Han/Spurious-Token-Game.RNGBench-Game-Trajectories
RNGBench-Game-Trajectories
SFT (supervised fine-tuning) trajectory data accompanying RNGBench · Reconstructive Non-Markov
Games — an evaluation framework that tests whether multimodal language models can reconstruct
hidden state from memory and act on it in closed-loop environments (the "remember-to-act" setting,
where the current observation alone is not enough and the model must recall relevant history before
deciding).
📄 Paper: https://arxiv.org/abs/2606.19338
🌐 Project… See the full description on the dataset page: https://huggingface.co/datasets/internlm/RNGBench-Game-Trajectories.Unity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam
Unity Code and GPT-Generated GDD Pairs Dataset
This dataset contains paired samples of Unity game mechanic scripts and their corresponding GPT-4 generated Game Design Documents (GDDs). It is intended for training and benchmarking LLMs in game code generation from design specifications.
Format
Each entry is stored as a .jsonl file with:
"input": GPT-4 generated GDD describing a specific game and its mechanics
"output": Unity C# scripts implementing the described mechanic… See the full description on the dataset page: https://huggingface.co/datasets/AmnaHassan/Unity-Engine-CSharp-Code-and-Game-Design-Document-Code-Pairs-Mix-and-Jam.spaceship-game-leaderboard
Spaceship Game - Leaderboard
This dataset contains leaderboard entries for the Spaceship Game on Reachy Mini.
Stats
Entries: 2
Top Score: 600 by Pilot
Last Updated: 2026-08-25
Published by: DarthBinks
Format
The leaderboard.json file contains an array of entries:
Field
Type
Description
score
int
Final game score
name
string
Player name
date
string
ISO 8601 timestamp
waves_completed
int?
Number of waves completed… See the full description on the dataset page: https://huggingface.co/datasets/DarthBinks/spaceship-game-leaderboard.spaceship-game-leaderboard
Spaceship Game - Leaderboard
This dataset contains leaderboard entries for the Spaceship Game on Reachy Mini.
Stats
Entries: 8
Top Score: 9100 by ri7
Last Updated: 2026-09-21
Published by: ri7-5
Format
The leaderboard.json file contains an array of entries:
Field
Type
Description
score
int
Final game score
name
string
Player name
date
string
ISO 8601 timestamp
waves_completed
int?
Number of waves completed
Top 10… See the full description on the dataset page: https://huggingface.co/datasets/ri7-5/spaceship-game-leaderboard.board_games
Board Game Datasets
This directory contains board-game position datasets, produced for generating verified Q&A
about (1) interpreting a board position given its standard notation, and (2) advising a good
next move. Each game lives in its own self-contained project (own extract.py,
pyproject.toml, .venv) and produces one JSON file. Ground truth (legal moves, best move,
evaluation) always comes from a real rules library / game engine — never guessed by an LLM.
This is the… See the full description on the dataset page: https://huggingface.co/datasets/nlp-and-reasoning/board_games.gamedev-nocode
Gamedev Data Set
The Gamedev Data Set was created with focus on the many aspects of game development that are not coding. As such, there are only a few code-specific entries, and the rest focus on everything from marketing to design theory and project management. It is meant as a companion set to coding-specific data.
This data set consists of 76,114 questions and answer sets. The data is generated by AI, guided by humans, and based on notes taken from publicly available sources.… See the full description on the dataset page: https://huggingface.co/datasets/theprint/gamedev-nocode.lemonseed-games-r2
lemonseed-games-r2
LemonSeed — games round 2 (Go/Sudoku atari + constraint reasoning).
Contents
games_r2.jsonl (12000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
game-theory-business-strategy
⚠️ DEPRECATED — Game Theory Business Strategy
This dataset has been superseded by GameTheory-Formulator, which contains 1,215 real-world scenarios across 6 domains (including 220 business scenarios) with full formulation steps, solutions, and interpretations.
Migration
Please use the newer, more comprehensive dataset:
from datasets import load_dataset
# New dataset (recommended)
ds = load_dataset("Alogotron/GameTheory-Formulator", split="train")
# Filter for… See the full description on the dataset page: https://huggingface.co/datasets/Alogotron/game-theory-business-strategy.ru-word-games
Dataset Summary
Dataset contains more than 100k examples of pairs word-description, where description is kind of crossword question. It could be useful for models that generate some description for a word, or try to a guess word from a description.
Source code for parsers and example of project are available here
Key stats:
Number of examples: 133223
Number of sources: 8
Number of unique answers: 35024
subset
count
350_zagadok
350
bashnya_slov
43522
crosswords
39290… See the full description on the dataset page: https://huggingface.co/datasets/artemsnegirev/ru-word-games.repro-doubly-exponential-lower-bounds-for-follow-the-regularized-leader-in-potential-games
Doubly Exponential Lower Bounds for Follow-the-Regularized-Leader
This is a reproduction logbook for ICML 2026.
OpenReview ID: l6KZJO7w48
Paper Abstract
This logbook reproduces theoretical results about FTRL convergence in potential games.
See logbook.json for full claim verification details.
spaceship-game-leaderboard
Spaceship Game - Leaderboard
This dataset contains leaderboard entries for the Spaceship Game on Reachy Mini.
Stats
Entries: 3
Top Score: 8150 by Romeo
Last Updated: 2026-08-22
Published by: EdRedon
Format
The leaderboard.json file contains an array of entries:
Field
Type
Description
score
int
Final game score
name
string
Player name
date
string
ISO 8601 timestamp
waves_completed
int?
Number of waves completed
Top… See the full description on the dataset page: https://huggingface.co/datasets/EdRedon/spaceship-game-leaderboard.
