CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01p-doom /atari-breakout-dataset Atari-Breakout Dataset This is a large dataset of 10M video frames and actions collected from the Breakout atari environment (Bellemare et al., 2012) in order to train world models.The dataset enables reproducible, large-scale experiments in action-conditioned video prediction. It is meant to be used with Jasmine, our JAX-based world modeling codebase. Dataset Summary Environment: Atari Learning Environment Frames: 10 million Resolution: 84 × 84 Format:… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/atari-breakout-dataset.0 likes1.4k downloads11mo agoHugging Face02p-doom /doom-dataset Doom Dataset This is a large dataset of 10M video frames and actions collected from the Doom environment for training world models.The dataset enables reproducible, large-scale experiments in action-conditioned video prediction. It is meant to be used with Jasmine, our JAX-based world modeling codebase. Dataset Summary Environment: VizDoom Frames: 10 million Resolution: 60 × 80 Format: ArrayRecord (for fast I/O) Splits: train / val / test License: CC0 1.0… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/doom-dataset.4 likes836 downloads11mo agoHugging Face03invocation02 /doom-dataset-largetext100K<n<1M0 likes799 downloads10mo agoHugging Face04doomskeletonpr /denemedatabaseOptimized binary datasets for Project X. Status: Confidential / Work in Progress Since this is a confidential project, the files are encrypted. Please do not delete the files. Thank you :) 1 likes737 downloads9mo agoHugging Face05invocation02 /doom-rnd-largetext1M<n<10M0 likes719 downloads10mo agoHugging Face06p-doom /crowd-code-dataset-1.0 Install crowd-code 2.0 to help crowd-source the next-generation coding dataset. crowd-code-dataset-1.0 is an anonymized dataset of fine-grained IDE interactions crowd-sourced across 25 people over the last 6 months using crowd-code 1.0, a VS Code/Cursor extension capturing large parts of the software engineering workflow. The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-1.0.text1M<n<10M5 likes629 downloads8mo agoHugging Face07p-doom /common-craft common-craft common-craft is a dataset of Minecraft survival gameplay videos curated for research on world modeling. It can be combined with scalable repositories such as Jasmine. Overview Total duration: ~5,000 hours Content type: Survival-mode Minecraft Let's Plays Source: YouTube videos and playlists (handpicked) Format: Raw videos (.mp4 and .webm) at a resolution of 640x360 and 30 FPS along with full metadata Structure common-craft/ ├──… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/common-craft.2 likes575 downloads11mo agoHugging Face08invocation02 /doom-dataset-1text10M<n<100M0 likes530 downloads10mo agoHugging Face09p-doom /atari-crazy_climber-dataset Atari-Crazy Climber Dataset This is a large dataset of 10M video frames and actions collected from the Crazy Climber atari environment (Bellemare et al., 2012) in order to train world models.The dataset enables reproducible, large-scale experiments in action-conditioned video prediction. It is meant to be used with Jasmine, our JAX-based world modeling codebase. Dataset Summary Environment: Atari Learning Environment Frames: 10 million Resolution: 84 × 84… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/atari-crazy_climber-dataset.0 likes470 downloads11mo agoHugging Face10RohanNaga /doom-dense-arnold-latents DoomDiT per-tic latents Derived from the public raw dataset RohanNaga/doom-dense-arnold: every tic of each episode encoded once through a frozen Stable Diffusion autoencoder, stored as one .npy per episode plus a .npz sidecar with the per-tic metadata. Released so that a training run can start without re-encoding, and so that latents are reproducible bit for bit: latent bytes depend on the autoencoder, its scaling, the precision and the encode batch size, all recorded in… See the full description on the dataset page: https://huggingface.co/datasets/RohanNaga/doom-dense-arnold-latents.0 likes434 downloads4m agoHugging Face11p-doom /atari-assault-dataset Atari-Assault Dataset This is a large dataset of 10M video frames and actions collected from the Assault atari environment (Bellemare et al., 2012) in order to train world models.The dataset enables reproducible, large-scale experiments in action-conditioned video prediction. It is meant to be used with Jasmine, our JAX-based world modeling codebase. Dataset Summary Environment: Atari Learning Environment Frames: 10 million Resolution: 84 × 84 Format:… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/atari-assault-dataset.0 likes431 downloads11mo agoHugging Face12p-doom /atari-battle_zone-dataset Atari-Battle Zone Dataset This is a large dataset of 10M video frames and actions collected from the Battle Zone atari environment (Bellemare et al., 2012) in order to train world models.The dataset enables reproducible, large-scale experiments in action-conditioned video prediction. It is meant to be used with Jasmine, our JAX-based world modeling codebase. Dataset Summary Environment: Atari Learning Environment Frames: 10 million Resolution: 84 × 84 Format:… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/atari-battle_zone-dataset.0 likes421 downloads11mo agoHugging Face13p-doom /atari-amidar-dataset Atari-Amidar Dataset This is a large dataset of 10M video frames and actions collected from the Amidar atari environment (Bellemare et al., 2012) in order to train world models.The dataset enables reproducible, large-scale experiments in action-conditioned video prediction. It is meant to be used with Jasmine, our JAX-based world modeling codebase. Dataset Summary Environment: Atari Learning Environment Frames: 10 million Resolution: 84 × 84 Format:… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/atari-amidar-dataset.0 likes379 downloads11mo agoHugging Face14p-doom /atari-boxing-dataset Atari-Boxing Dataset This is a large dataset of 10M video frames and actions collected from the Boxing atari environment (Bellemare et al., 2012) in order to train world models.The dataset enables reproducible, large-scale experiments in action-conditioned video prediction. It is meant to be used with Jasmine, our JAX-based world modeling codebase. Dataset Summary Environment: Atari Learning Environment Frames: 10 million Resolution: 84 × 84 Format:… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/atari-boxing-dataset.0 likes357 downloads11mo agoHugging Face15p-doom /atari-alien-dataset Atari-Alien Dataset This is a large dataset of 10M video frames and actions collected from the Alien atari environment (Bellemare et al., 2012) in order to train world models.The dataset enables reproducible, large-scale experiments in action-conditioned video prediction. It is meant to be used with Jasmine, our JAX-based world modeling codebase. Dataset Summary Environment: Atari Learning Environment Frames: 10 million Resolution: 84 × 84 Format: ArrayRecord… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/atari-alien-dataset.0 likes323 downloads11mo agoHugging Face16p-doom /atari-bank_heist-dataset Atari-Bank Heist Dataset This is a large dataset of 10M video frames and actions collected from the Bank Heist atari environment (Bellemare et al., 2012) in order to train world models.The dataset enables reproducible, large-scale experiments in action-conditioned video prediction. It is meant to be used with Jasmine, our JAX-based world modeling codebase. Dataset Summary Environment: Atari Learning Environment Frames: 10 million Resolution: 84 × 84 Format:… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/atari-bank_heist-dataset.0 likes322 downloads11mo agoHugging Face17p-doom /atari-asterix-dataset Atari-Asterix Dataset This is a large dataset of 10M video frames and actions collected from the Alien atari environment (Bellemare et al., 2012) in order to train world models.The dataset enables reproducible, large-scale experiments in action-conditioned video prediction. It is meant to be used with Jasmine, our JAX-based world modeling codebase. Dataset Summary Environment: Atari Learning Environment Frames: 10 million Resolution: 84 × 84 Format:… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/atari-asterix-dataset.0 likes295 downloads11mo agoHugging Face18p-doom /coinrun-dataset CoinRun Dataset This is a large dataset of 50M video frames and actions collected from the CoinRun environment (Cobbe et al., 2020) for training world models.The dataset enables reproducible, large-scale experiments in action-conditioned video prediction. It is meant to be used with Jasmine, our JAX-based world modeling codebase. Dataset Summary Environment: CoinRun (Procgen Benchmark) Frames: 50 million Resolution: 64 × 64 Format: ArrayRecord (for fast I/O)… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/coinrun-dataset.1 likes287 downloads11mo agoHugging Face19doomon /RADAR_datasetimage10K<n<100K0 likes285 downloads5mo agoHugging Face20DoomAI /doomaiimage10K<n<100K0 likes277 downloads2y agoHugging Face21p-doom /idm-eval IDM Eval Set A validation set for evaluating Inverse Dynamics Models on macOS screen recordings. Each sample is a 5-second clip of real productivity desktop usage (browser, IDE, terminal, docs, dashboards) paired with a ground-truth action log captured at the OS level. The task: given a short screen recording, predict the sequence of user input actions (keypresses, mouse clicks, scrolls, cursor moves) that produced the observed screen changes. Code Training… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/idm-eval.video-classificationn<1K0 likes220 downloads3mo agoHugging Face22p-doom /openai-minecraft-datasetThis dataset is a preprocessed version of the OpenAI Minecraft Video Dataset VPT, reformatted into array record files compatible with the Grain dataloader for efficient streaming. 9 likes218 downloads1y agoHugging Face23p-doom /AGI-CAST-0.6k AGI-CAST: Behaviour-Cloning Knowledge Work AGI-CAST-0.6k is a >600-hour dataset of screencasts capturing raw workflows of researchers at p(doom). This is the first major release in our effort to collect months-long fine-grained expert trajectories of human reasoning in their day-to-day work. Please refer to the blog post for more details. Dataset Summary Contemporary frontier model development has saturated the internet and… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/AGI-CAST-0.6k.100K<n<1M24 likes218 downloads6mo agoHugging Face24p-doom /crowd-code-dataset-0.1The crowd-code-dataset-0.1 is a raw, unfiltered dataset of fine-grained IDE interactions collected during the development of Jasmine using crowd-code, a VS Code/Cursor extension capturing large parts of the software engineering workflow. The dataset captures real research engineering workflows (character-level edits, navigation, terminal use, iterative debugging). The crowd-code-dataset-0.1 only includes data from the Jasmine authors. We are actively working on cleaning and curating the full… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/crowd-code-dataset-0.1.tabular100K<n<1M5 likes199 downloads8mo agoHugging Face25p-doom /atari-demon_attack-dataset Atari-Demon Attack Dataset This is a large dataset of 10M video frames and actions collected from the Demon Attack atari environment (Bellemare et al., 2012) in order to train world models.The dataset enables reproducible, large-scale experiments in action-conditioned video prediction. It is meant to be used with Jasmine, our JAX-based world modeling codebase. Dataset Summary Environment: Atari Learning Environment Frames: 10 million Resolution: 84 × 84… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/atari-demon_attack-dataset.0 likes196 downloads11mo agoHugging Face26lucrbrtv /doom-e1-internet-gameplayI've used the Inverse Dynamic Model, I've previously trained on manually recorded gameplay, on pure gameplay YouTube videos. This dataset is in public domain, use it however you want. image100K<n<1M0 likes137 downloads6mo agoHugging Face27RohanNaga /doom-dense-arnold DoomDiT dense Arnold recordings Lossless, per-tic recordings of the Arnold agent (Lample and Chaplot, AAAI 2017) playing ViZDoom deathmatch with 8 bots on Freedoom assets, made for training action-conditioned world models. Every engine tic (35 per second) is stored with the executed control vector, so the data can be used at any frame stride. Recorded September 2026 at CMU for the DoomDiT project (Rohan Nagabhirava, Keerthana Chirumamilla). What is here… See the full description on the dataset page: https://huggingface.co/datasets/RohanNaga/doom-dense-arnold.tabularimage-to-image1M<n<10M0 likes122 downloads2d agoHugging Face28ScoobyBaby1999 /doomalay-superpowers doomalay-superpowers The obra/superpowers agent-skills corpus, hub-native for the doomalay public library (v3): every skill is a whole-directory BUNDLE (SKILL.md + scripts + references + prompts riding one {"v":1,"entry":"SKILL.md","files":[…]} manifest), the docs ride as hidden doc items inside the superpowers-obra bunch, and the maintainer shell scripts ride the script library. items/index.json is the item list; each item's payload lives at its file path. What's… See the full description on the dataset page: https://huggingface.co/datasets/ScoobyBaby1999/doomalay-superpowers.tabularn<1K0 likes111 downloads3h agoHugging Face29BharathK333 /DOOMGAN-Ocular-Morphs DOOMGAN: Ocular Morph Dataset This repository contains the official public dataset for the paper: "DOOMGAN: High-Fidelity Dynamic Identity Obfuscation Ocular Generative Morphing" funded by the NSF award no. 2345561. The dataset consists of 10,000 high-fidelity morphed ocular images generated by the DOOMGAN model. These images are intended to facilitate research and development of Morph Attack Detection (MAD) systems for visible-spectrum ocular biometrics. Paper: IJCB 2025 DOOMGAN… See the full description on the dataset page: https://huggingface.co/datasets/BharathK333/DOOMGAN-Ocular-Morphs.imagefeature-extraction10K<n<100K1 likes103 downloads1y agoHugging Face30Karajan42 /doom-dungeon-55 doom-dungeon-55 Checkpoint archive for the doom_v55_hg world-model lineage. The canonical, fully documented serving copy of the robot-campaign checkpoint (v6sf_phase2_ep26/) lives at alakazamworld/doom-dungeon-hg; start there. 0 likes93 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.