datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gaming-500-hours
Gaming Dataset (gaming-1) — 494.7 Hours
Native PC/console gameplay screen-recordings, organized by game. Each workflow
is one play session, trimmed to pure gameplay — login screens, launchers,
desktop, collection-app references, and any watching/streaming are removed.
In-game menus, lobbies, loading, and cutscenes are retained as part of the session.
Workflows: 776
Total gameplay: 494.7 hours
Distinct games: 168
Clip duration (min): median 24.0, p90 90.9, max 457.7
Platforms:… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/gaming-500-hours.cqadupstack-gaming
CQADupstackGamingRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
CQADupStack: A Benchmark Data Set for Community Question-Answering Research
Task category
t2t
Domains
Web, Written
Reference
http://nlp.cis.unimelb.edu.au/resources/cqadupstack/
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["CQADupstackGamingRetrieval"])
evaluator = mteb.MTEB(task)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/cqadupstack-gaming.my-shared-filesxiaoluo-gaming-action3000-20260910-media
Action-boundary review examples
Media for 300 selected examples from Xiaoluo (Cyberpunk 2077 and Rise of the Tomb Raider) and Gaming 500 Hours, 150 examples per dataset.
Includes 15-second review videos, observed boundary frames, and available action clips. These are visual model estimates; boundaries require human review. Source game and dataset rights remain with their respective owners.
Gallery and annotation manifests:… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/xiaoluo-gaming-action3000-20260910-media.gaming-500-hours
Gaming Dataset (gaming-1) — 494.7 Hours
Native PC/console gameplay screen-recordings, organized by game. Each workflow
is one play session, trimmed to pure gameplay — login screens, launchers,
desktop, collection-app references, and any watching/streaming are removed.
In-game menus, lobbies, loading, and cutscenes are retained as part of the session.
Workflows: 776
Total gameplay: 494.7 hours
Distinct games: 168
Clip duration (min): median 24.0, p90 90.9, max 457.7
Platforms:… See the full description on the dataset page: https://huggingface.co/datasets/WillowVoiceAI/gaming-500-hours.gaming-500-hours
Gaming Dataset (gaming-1) — 494.7 Hours
Native PC/console gameplay screen-recordings, organized by game. Each workflow
is one play session, trimmed to pure gameplay — login screens, launchers,
desktop, collection-app references, and any watching/streaming are removed.
In-game menus, lobbies, loading, and cutscenes are retained as part of the session.
Workflows: 776
Total gameplay: 494.7 hours
Distinct games: 168
Clip duration (min): median 24.0, p90 90.9, max 457.7
Platforms:… See the full description on the dataset page: https://huggingface.co/datasets/Naya2004/gaming-500-hours.gaming-500-hours
Gaming Dataset (gaming-1) — 494.7 Hours
Native PC/console gameplay screen-recordings, organized by game. Each workflow
is one play session, trimmed to pure gameplay — login screens, launchers,
desktop, collection-app references, and any watching/streaming are removed.
In-game menus, lobbies, loading, and cutscenes are retained as part of the session.
Workflows: 776
Total gameplay: 494.7 hours
Distinct games: 168
Clip duration (min): median 24.0, p90 90.9, max 457.7
Platforms:… See the full description on the dataset page: https://huggingface.co/datasets/ApertureQA/gaming-500-hours.gaming-500-hours
Gaming Dataset (gaming-1) — 494.7 Hours
Native PC/console gameplay screen-recordings, organized by game. Each workflow
is one play session, trimmed to pure gameplay — login screens, launchers,
desktop, collection-app references, and any watching/streaming are removed.
In-game menus, lobbies, loading, and cutscenes are retained as part of the session.
Workflows: 776
Total gameplay: 494.7 hours
Distinct games: 168
Clip duration (min): median 24.0, p90 90.9, max 457.7
Platforms:… See the full description on the dataset page: https://huggingface.co/datasets/xsghh/gaming-500-hours.gaming-500-hours
Gaming Dataset (gaming-1) — 494.7 Hours
Native PC/console gameplay screen-recordings, organized by game. Each workflow
is one play session, trimmed to pure gameplay — login screens, launchers,
desktop, collection-app references, and any watching/streaming are removed.
In-game menus, lobbies, loading, and cutscenes are retained as part of the session.
Workflows: 776
Total gameplay: 494.7 hours
Distinct games: 168
Clip duration (min): median 24.0, p90 90.9, max 457.7
Platforms:… See the full description on the dataset page: https://huggingface.co/datasets/manrajs/gaming-500-hours.gaming-input-review-100-media-20260910100 Gaming input-only preview clips, each 15 seconds. No visual filtering or semantic action labeling was performed for this sample. The sample is stratified by game and includes 30 GTA V clips. Use samples.jsonl for source intervals, checksums and input timing. Video paths are relative to this repo.
Companion review: https://huggingface.co/spaces/mikusama99/sunain-gaming-review-200-20260910
gaming-ai-corpus
🎮 Gaming AI Corpus & Dataset Community
A massive, structured collection of multi-game wiki text and scripting data designed for training, fine-tuning Large Language Models (LLMs), building Retrieval-Augmented Generation (RAG) systems, and developing custom video game AI agents.
🌟 Support & Project Resources
If you find this dataset useful for your AI or RAG projects, please support the project by leaving a star on GitHub!
GitHub Repository:… See the full description on the dataset page: https://huggingface.co/datasets/hsosa/gaming-ai-corpus.gaming-video-text-clean98
Gaming Video Text Data Notes
Dataset summary
Preparation notes and schema examples for Gaming tasks using Video Text data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/BINTANGPURN/gaming-video-text-clean98.gaming-audio-video
Gaming Audio Video Data Notes
Dataset summary
Preparation notes and schema examples for Gaming tasks using Audio Video data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
build_dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/niko-laev/gaming-audio-video.gaming-image-audio-clean
Gaming Image Audio Data Notes
Dataset summary
A documented Gaming data-preparation workflow for Image Audio records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
preprocess.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md —… See the full description on the dataset page: https://huggingface.co/datasets/smithjoshua1982/gaming-image-audio-clean.phd-gaming
Gaming Text Tabular Data Notes
Dataset summary
Preparation notes and schema examples for Gaming tasks using Text Tabular data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
clean.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/Imnataliamorozov/phd-gaming.gaming500-720p-hdf5gaming-sensor-fusion
Gaming Sensor Fusion Data Notes
Dataset summary
A documented Gaming data-preparation workflow for Sensor Fusion records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md —… See the full description on the dataset page: https://huggingface.co/datasets/ALLENYOON1988/gaming-sensor-fusion.toy-gaming
Gaming Sensor Fusion Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Gaming work with Sensor Fusion inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/pdxgreen/toy-gaming.gaming-pointcloud-text50
Gaming Pointcloud Text Data Notes
Dataset summary
This data card accompanies a lightweight Gaming loader for Pointcloud Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
preprocess.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/thaparbiolab/gaming-pointcloud-text50.gaming-video-text
Gaming Video Text Data Notes
Dataset summary
Preparation notes and schema examples for Gaming tasks using Video Text data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/VICTORWVW/gaming-video-text.cqadupstack-gaming-trgaming-image-text-benchmark56
Gaming Image Text Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Gaming work with Image Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
dataset.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable… See the full description on the dataset page: https://huggingface.co/datasets/Jprob-inson/gaming-image-text-benchmark56.random-gaming
Gaming Image Depth Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Gaming work with Image Depth inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/yuchenhuang/random-gaming.gaming-collection
Gaming Image Text Data Notes
Dataset summary
Preparation notes and schema examples for Gaming tasks using Image Text data. Full source material is intentionally not bundled, so provenance and licensing remain explicit.
Included material
loader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/sant-oso39/gaming-collection.beir-cqadupstack-gaming
CQADupstackGamingRetrieval — BEIR, unified schema
A normalised copy of the dataset behind the mteb task CQADupstackGamingRetrieval, one of the tasks of the BEIR benchmark as mteb defines it (a member of the aggregate task CQADupstackRetrieval). Same queries, documents
and relevance judgements as the benchmark evaluates — reshaped into one strict schema shared by every dataset
in this collection.
Source
mteb/cqadupstack-gaming @ 4885aa143210 (the revision pinned in… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/beir-cqadupstack-gaming.gaming-image-text-benchmark-2023
Gaming Image Text Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Gaming work with Image Text inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
dataloader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/Eymenk-ara8526/gaming-image-text-benchmark-2023.gaming-audio-video
Gaming Audio Video Data Notes
Dataset summary
This data card accompanies a lightweight Gaming loader for Audio Video metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
preprocess.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/Vellorephotonics/gaming-audio-video.test-gaming-2024
Gaming Video Text Data Notes
Dataset summary
This data card accompanies a lightweight Gaming loader for Video Text metadata. It is meant for pipeline inspection, source adaptation, and reproducible split preparation.
Included material
dataloader.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data card and… See the full description on the dataset page: https://huggingface.co/datasets/connoredw/test-gaming-2024.cv-gaming
Gaming Image Depth Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Gaming work with Image Depth inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/ritsumeineurolab/cv-gaming.trial-gaming
Gaming Text Tabular Data Notes
Dataset summary
A documented Gaming data-preparation workflow for Text Tabular records. The bundled rows demonstrate the schema and validation path rather than pretending to be a full training corpus.
Included material
clean.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small, human-readable records for checking the schema.
README.md — data… See the full description on the dataset page: https://huggingface.co/datasets/bjmachado/trial-gaming.
