datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NBA_Games
NBA Full-Game Video Dataset
This dataset provides metadata, official statistics, and official play-by-play annotations for full-length NBA game videos available on YouTube. Instead of redistributing video files, we provide YouTube video IDs and URLs so users can download videos independently when their use case and local policies allow it.
The dataset links long-form basketball videos with structured NBA.com game data. Each retained game has a verified… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/NBA_Games.FIFA_World_Cup_Games
FIFA World Cup Full-Match Video Dataset
This dataset provides YouTube full-match references and structured match annotations for classic FIFA World Cup games. Instead of redistributing video files, it stores YouTube video IDs and URLs, then links each retained match to structured football data including match metadata, team statistics, lineups, formations, text event timelines, and pre-match prediction polls.
The dataset is designed for long-form sports… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/FIFA_World_Cup_Games.threejs-gamecode-instruct-v3-ultra
Three.js GameCode Instruct v3 Ultra
This is a large synthetic/original instruction dataset for training or testing LLM behavior around Three.js, browser game development, gameplay programming, debugging, optimization, architecture, and general coding.
Important note
This dataset is synthetic and programmatically generated from original templates. It is designed as a useful starting point for experiments, not as a fully hand-curated gold-standard benchmark.
No… See the full description on the dataset page: https://huggingface.co/datasets/agagasf123123/threejs-gamecode-instruct-v3-ultra.gamedev-nocode
Gamedev Data Set
The Gamedev Data Set was created with focus on the many aspects of game development that are not coding. As such, there are only a few code-specific entries, and the rest focus on everything from marketing to design theory and project management. It is meant as a companion set to coding-specific data.
This data set consists of 76,114 questions and answer sets. The data is generated by AI, guided by humans, and based on notes taken from publicly available sources.… See the full description on the dataset page: https://huggingface.co/datasets/theprint/gamedev-nocode.lemonseed-games-r2
lemonseed-games-r2
LemonSeed — games round 2 (Go/Sudoku atari + constraint reasoning).
Contents
games_r2.jsonl (12000 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
infinite-game-os
Dataset Card for Infinite Game OS for Sovereign Creators
Dataset Description
This dataset packages the Infinite Game OS framework as an instruction-tuning corpus for AI training pipelines. The framework was created by Lane Belone. It gives Creators a philosophical and structural foundation for building a business from authentic expression rather than market optimization.
The central organizing pattern is the Creator Flywheel: live the life, share the breadcrumbs… See the full description on the dataset page: https://huggingface.co/datasets/lanebelone/infinite-game-os.game_translate
