datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cad-environments
CAD Environments
CAD Environments is a multimodal dataset of complete, human-performed workflows in desktop CAD software. The current release contains 51 task workflows totaling 99.03 hours, covering eight software groups across mechanical design, architecture, MEP, structural design, and general 3D modeling.
Each workflow preserves the full task context—not just the final model—including the problem statement, reference and input files, a gold output, evaluation rubrics, a… See the full description on the dataset page: https://huggingface.co/datasets/markov-ai/cad-environments.gaia2
Gaia2
Paper | Code | Project Page
Dataset Summary
Gaia2 is a benchmark dataset for evaluating AI agent capabilities in simulated environments. The dataset contains 800 scenarios that test agent performance in environments where time flows continuously and events occur dynamically.
The dataset evaluates seven core capabilities: Execution (multi-step planning and state changes), Search (information gathering and synthesis), Adaptability (dynamic response to environmental… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2.gaia2_filesystem
GAIA2 Filesystem
This is a dataset containing files for the GAIA2 benchmark. You should not use this dataset on its own, but instead use the Meta Agents Research Environments framework to execute scenarios from that GAIA2 dataset.
Dataset Link
https://huggingface.co/datasets/meta-agents-research-environments/gaia2
Contact Details
Publishing POC: Meta AI Research Team
Affiliation: Meta Platforms, Inc.
Website:… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2_filesystem.HydroGym-environmentsSPADE-Environments-Qwen3-30B-Games
SPADE generated environments: games
Paper | Code | All artifacts
Executable game environments written by the SPADE Environment Designer during the paper's 30B games self-play run. One Python file per environment; manifest.json records the generation checkpoint, training step, skill, and difficulty of each.
Environments
3310
Training steps covered
113 (step 0 to 396)
With skill label
3119
Designer / agent model
Qwen/Qwen3-30B-A3B-Instruct-2507… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environments-Qwen3-30B-Games.SPADE-Environments-ToolUse
SPADE generated environments: tool use
Paper | Code | All artifacts
Multi-turn tool-use environments written by the SPADE designer during training, pooled
across every captured run. 2,231 environments across 7 runs and two model scales (30B-A3B and 4B).
Source run
Scale
Environments
qwen3-30b-0617-tooluse-regen32-mixed
30B-A3B
41
qwen3-30b-0624-tooluse-blend
30B-A3B
243
qwen3-30b-0703-tooluse-glory-kl005
30B-A3B
260
qwen3-4b-0630-tooluse-eval-aligned-r32
4B
456… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environments-ToolUse.proc-gen-environments
geodesic-research/proc-gen-environments
Local-pipeline snapshot published via --push-from-local (GH #52). All configs below were built locally (Hub-independent) and uploaded in a single commit at one snapshot revision.
Pipeline run params hash: dd13a8843fce63fefa4e70c743f85c6cba30904523b6097a96bfe98100c298eb
Configs in this snapshot: case-elaboration-conversation, cases-conversation, cases-fields-conversation, cases-intermediate-conversation, contracts-conversation
Per-run… See the full description on the dataset page: https://huggingface.co/datasets/geodesic-research/proc-gen-environments.gaia2-cli
GAIA2 CLI
Benchmark dataset for gaia2-cli, the CLI-based agent evaluation harness.
Schema
Each row has two columns:
Column
Type
Description
scenario_id
string
Unique scenario identifier (e.g. scenario_universe_21_1qgjj6)
scenario
string
Complete scenario as a JSON string
Usage
from datasets import load_dataset
import json
# Load a specific config (160 scenarios)
ds = load_dataset("meta-agents-research-environments/gaia2-cli", "adaptability"… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2-cli.DROID-sim-environmentsaria-synthetic-environments
Aria Synthetic Environments (ASE) Dataset
Introduction
The Aria synthetic environments (ASE) dataset is a large-scale synthetic dataset featuring 100,000 procedurally-generated indoor scenes. It is designed for research on 3D scene understanding, object detection, and tracking. Each scene is populated with realistic 3D objects and simulated sensor data that mirrors Project Aria glasses' characteristics.
Unlike prior datasets for 3D scene understanding, which are… See the full description on the dataset page: https://huggingface.co/datasets/projectaria/aria-synthetic-environments.penbioscience-environmentsprocessrl-terminal-environments
ProcessRL Terminal Environments
ProcessRL is a collection of behavior-conditioned terminal environments for training and evaluating agent process control. The tasks are designed around failures that appear in interactive terminal work: stopping after a misleading successful command, repeating an unproductive action, failing to pivot after a dead end, losing track of migrated state, and leaving partial progress unfinished.
This release contains the first public train/heldout… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/processrl-terminal-environments.KOR-RE-natures-and-environments
Dataset Card for [KOR-RE-natures-and-environments]
You can find relation map, guidelines(written in Korean), short technical papers in this github repo. This work is done by as part of project for Boostcamp AI Tech supported by Naver Connect Foundation.
Main Data Fields
Sentences: sentences
Subject_entity: infos for subject entity in the sentence including words, start index, end index, type of entity
object_entity: infos for object entity in the sentence including words… See the full description on the dataset page: https://huggingface.co/datasets/kimcando/KOR-RE-natures-and-environments.SPADE-Environments-Qwen3-30B-ToolUse
qwen3-30B-A3B-Instruct-0703-tooluse-glory-kl005 — generated environments
Environments generated by the SPARE proposer during training run
2hjdrbeh (qwen3-30B-A3B-Instruct-0703-tooluse-glory-kl005), recovered from the spare-viz durable cache.
The run's scratch directory no longer exists; this dataset is the surviving copy.
Games
260
Steps covered
7 (step 0–192)
With recovered skill
260
With hint
0
Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/SPADE-Environments-Qwen3-30B-ToolUse.quadruped-environmentsclaimcheck-rl-environments
ClaimCheck RL Environments
Public environment artifacts for the ClaimCheck evidence-verification project.
Prime Hub Environment
Environment
Prime Hub
Version
Contents
claimcheck-sft
https://app.primeintellect.ai/dashboard/environments/mohammedalshehri-77/claimcheck-sft
0.1.4
Static single-turn evidence-verification environment packaged for Prime eval/RL runs.
The Prime package is mirrored under prime_hub/claimcheck_sft/. It includes the environment… See the full description on the dataset page: https://huggingface.co/datasets/mohammed8284/claimcheck-rl-environments.vibeworlding-environments-10
VibeWorlding 场景准备:10 个场景
独立场景源码与验证记录,方便协作下载。 不是作者 V2 原生训练集,也不是统一通过完整物理验证的数据集。场景未与新增资产库提前绑定。
下载和打开
下载完整环境包(67.5 MB)
解压后在包根目录运行 python3 -m http.server 8766 --bind 127.0.0.1,打开 http://127.0.0.1:8766/batch-001-ten/gallery.html 看场景,或 http://127.0.0.1:8766/v2-physics-pilot/index.html 看物理测试回放。无需部署 AI 模型;HF 数据集页面供下载,HTML 预览需上述本地静态服务。
现有验证状态
10/10 通过静态准备检查;8/10 通用 Three.js JSON 重载与单视角画面对照通过。10 个场景都实际运行了 Rapier 20 秒、240 Hz 刚体仿真,5/10 满足本轮全部保守判据。保存了全部失败和逐物体轨迹。… See the full description on the dataset page: https://huggingface.co/datasets/anon123312/vibeworlding-environments-10.scalewob-environments
ScaleWoB Environments
Private, versioned browser-runtime assets corresponding to the scalewob-verl dataset package.
Contents
scalewob-env-v0.1.0.tar.zst: deterministic archive containing the scalewob-env/ directory.
manifest.json: archive checksum, file counts, indexed environment count, and compatible runtime versions.
THIRD_PARTY_NOTICES.md: preliminary redistribution audit notes; not a complete license inventory.
The source tree is approximately 355 MB and… See the full description on the dataset page: https://huggingface.co/datasets/hysi-lab/scalewob-environments.processrl-terminal-environments
ProcessRL Terminal Environments
ProcessRL is a collection of behavior-conditioned terminal environments for training and evaluating agent process control. The tasks are designed around failures that appear in interactive terminal work: stopping after a misleading successful command, repeating an unproductive action, failing to pivot after a dead end, losing track of migrated state, and leaving partial progress unfinished.
This release contains the first public train/heldout… See the full description on the dataset page: https://huggingface.co/datasets/poolside-laguna-hackathon/processrl-terminal-environments.environments
Chapaty Environments
This dataset repository hosts the pre-compiled financial environments for Chapaty, a Rust library for building quantitative trading agents in financial markets.
Inspired by OpenAI Gymnasium, these datasets provide standardized, easily reproducible simulation states for training and backtesting agents.
What are these files?
The datasets here are serialized using .postcard (a highly efficient binary format for Rust). They contain environment… See the full description on the dataset page: https://huggingface.co/datasets/chapaty/environments.Environment_Service_Projects
Environment_Service_Projects_Dataset
Description
This dataset contains information about various service learning projects related to environment that teaches students how to carry out these projects in their communities and schools and overall make nature a better place to live.
Structure
Structure of all data samples in dataset is : Context,Request and Response. Some only have Request and Response.
Example:
{
"Context": "Reduce plastic waste at… See the full description on the dataset page: https://huggingface.co/datasets/Moodyspider266/Environment_Service_Projects.trace-environmentsenvironmentsenvironmentsKletterMix-Ablations-Environments
KletterMix runtime environments
Pinned container images used for training, conversion, and evaluation. Runtime logs are intentionally not included.
Branch
Contents
main
See the files and pinned artifact manifest in the repository.
The branch names are part of the reproducibility interface. The private
KletterMix_Ablations
repository records the local source paths, exact branch commits, and the
submodule mapping. The tokenized data can be consumed by Megatron… See the full description on the dataset page: https://huggingface.co/datasets/RuHae/KletterMix-Ablations-Environments.cwm-benchmarks-dl4c-environmentsDesktop-Environmentsenvironment-state-description-datasetEnvironment State Description Dataset
A small dataset describing environmental conditions and corresponding
interpretive responses.
proc-gen-environments-vllm
