datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EO-Gym
EO Gym
EO Gym provides a local Earth-observation tool server and trainer environment
adapter. It exposes remote-sensing tools for cropping imagery, loading
multispectral bands, computing masks and indices, inspecting metadata, and
running EO Gym rollouts through a trainer-facing API.
Croissant metadata
EO-Gym provides two Croissant representations:
Hugging Face-generated Croissant metadata:
/api/datasets/paperuploadacount/EO-Gym/croissant
This is generated… See the full description on the dataset page: https://huggingface.co/datasets/paperuploadacount/EO-Gym.excelforum-nemo-gymgauss_gym_dataSWE-Factory-Gymnemo-gym-indian-bankingThis is NPCI/nemo-gym-indian-banking — the dataset for the indian_banking
resources server in NVIDIA NeMo Gym: 300 synthetic
multi-turn Indian retail-banking customer-support tasks (250 train / 50 validation), the
197-customer synthetic bank database and the 59-article knowledge base the environment
loads at startup.
NeMo Gym Indian Banking Agent Tasks
Tool-calling customer-service tasks for an Indian retail-banking assistant, in the
NVIDIA NeMo Gym agent-input JSONL format.… See the full description on the dataset page: https://huggingface.co/datasets/NPCI/nemo-gym-indian-banking.six-gym-sqlite
📢 Update 2026-03-23
We release BIRD-Critic-SQLite, a dataset containing 500 high-quality user issues focused on real-world SQLite database applications. This dataset is the train split of BIRD-Critic-SQLite, comprising 5,000 data instances for model training and development. Along with the dataset, we also release three RL-trained models: BIRD-Talon-14B, BIRD-Talon-7B, and BIRD-Zeno-7B.
📋 Dataset Structure
Below is a description of the dataset fields and additional… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/six-gym-sqlite.six-gym-pg-1.5
Update 2026-03-27
We release Six-Gym-PG-1.5, a train split of BIRD-Critic, comprising 5,000 data instances for model training and development.
Dataset Structure
Below is a description of the dataset fields and additional information about the structure:
instance_id: Unique identifier for each task.
issue_sql: The buggy SQL query written by the user.
dialect: The SQL dialect (PostgreSQL).
version: The dialect version (14.12).
db_id: The name of the database.
clean_up_sql:… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/six-gym-pg-1.5.Gym_Salesman_Dataset
Gym Salesman Dataset
11,997 synthetic gym-membership sales conversations, each labelled SUCCESS or FAILURE.
🔗 Project Links
| Live App — practice against an AI customer | Hugging Face Space |
| Telegram Bot — practice on the go | @ido_salescoach_bot |
| Dataset — 11,997 labelled conversations | elg4/Gym_Salesman_Dataset |
| Data Generation — how the data was built | notebook |
| Recommendation — the embedding retriever | notebook |
Every conversation is a… See the full description on the dataset page: https://huggingface.co/datasets/elg4/Gym_Salesman_Dataset.MoatlessTools-Agent-Verifier-Train-Datagym-aqa-beam-finegym
FineGym-AQA Balance Beam Scoring Setup
Prepared data for fine-tuning an AQA (Action Quality Assessment) model on balance-beam
routines, using the FineGym-AQA annotations (official competition D/E/ND/total scores).
Commercial-intent note (2026-09-21): the end product is commercial. This
FineGym-derived dataset and prototype are research-only (CC BY-NC 4.0); the
commercial model will be trained exclusively on owned, consent-gated data — see
data-collection/README.md for the full… See the full description on the dataset page: https://huggingface.co/datasets/dschultz0404/gym-aqa-beam-finegym.wiki_hotpotqavi-gym-causal-ascii
Vi-Gym Causal ASCII Trajectories
This dataset contains autoregressive trajectories of a Large Language Model (LLM) agent learning spatial reasoning and geometric drawing within a simulated Vi (Vim) editor environment.
Warning
This dataset is a direct derivation of the source material, it might therefore also contain content not suitable for all audiences. All authors of the original artwork have full ownership.
Dataset Structure
Each record is a discrete step… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/vi-gym-causal-ascii.chess_nemo_gymppl-gym-rollouts
ppl-gym rollouts
Model rollouts from ppl-gym, a benchmark of probabilistic-programming
problems: language-neutral problem statements with per-language ground-truth
realizations. Each rollout is one LLM-generated solution (program) for a
(model, language, problem) triple, executed and scored against ground truth
under a single answer comparator.
Code & methodology: https://github.com/shepardxia/ppl-gym
Browser: https://pplgym.kingdomofends.org
Configs
rollouts… See the full description on the dataset page: https://huggingface.co/datasets/Sheppp/ppl-gym-rollouts.D3-Gym-Trajectories
D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery
D3-Gym is the first automatically constructed dataset of verifiable environments for Data-Driven Discovery. It contains 565 tasks derived from 239 real-world multi-disciplinary scientific repositories.
The present dataset contains all training trajectories used in our paper, with each split representing the trajectories sampled from a model among the Qwen3 family.
Citation
If you find… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/D3-Gym-Trajectories.grabbeaiDani_01mo_tap_gym_nhung_van_beoOPSD-PI-SWE-Gym-512
OPSD-PI SWE-Gym Stage PI 512
Qwen3.5-9B stage-adaptive OPSD-PI 的公开 512-row 数据与 Weak from-scratch
训练包。Public 512-row data and Weak from-scratch training bundle.
Files
data/train.jsonl: 512 deterministic SWE-Gym rows with Weak, Medium, and
Strong PI for EXPLORE, REPRODUCE, DIAGNOSE, EDIT, and VERIFY.
data/manifest.json: source selection and integrity metadata.
release/OPSD_pi-opsd-pi-weak-from-scratch-20260818.tar.gz: immutable
source release containing launchers… See the full description on the dataset page: https://huggingface.co/datasets/LSW142857/OPSD-PI-SWE-Gym-512.hybrid_gym_func_gen_raw
Hybrid Gym: Function Generation Dataset
Dataset for the Function Generation benchmark task. Each instance contains a function with its signature and docstring, where the agent must implement the function body.
See the benchmark README for usage instructions.
Benchmark-R2E-Gym-Easy
Benchmark R2E-Gym Easy subset
This is a subset of R2E-Gym/R2E-Gym-Subset.
Licensing
This dataset is licensed under the Apache License 2.0. See
ATTRIBUTION.md for source-project attribution and
LICENSES/ for their license notices.
gauss_gym_arkitsmall-model-schema-gym
Small Model Schema Gym Dataset
Deterministically generated chat examples for first-pass JSON compliance on
the project Dream brief and Safety plan contracts.
Files
train.jsonl: 500 training examples.
validation.jsonl: 200 held-out examples.
manifest.json: counts, seed, provenance, and overlap check.
Each row contains:
{
"id": "stable example identifier",
"messages": [
{"role": "system", "content": "..."},
{"role": "user", "content": "..."}… See the full description on the dataset page: https://huggingface.co/datasets/KwabsHug/small-model-schema-gym.gym
Overview
This dataset, Gym Preference Dataset, is designed to help train models to generate high-quality, contextually appropriate responses to common gym-related questions. It contains pairs of prompts, preferred ("chosen") responses, and less preferred ("rejected") responses. The dataset is ideal for tasks such as preference learning, instruction tuning, and Reinforcement Learning from Human Feedback (RLHF).
Author: Yasin Arafat
License: MIT
Repository: YasinArafat05/gym… See the full description on the dataset page: https://huggingface.co/datasets/YasinArafat05/gym.hybrid_gym_func_localize_raw
Hybrid Gym: Function Localization Dataset
Dataset for the Function Localization benchmark task. Each instance contains a function description (without file path or name), and the agent must locate the function in the repository and add a docstring.
See the benchmark README for usage instructions.
Fields
instance_id: Unique identifier
repo: GitHub repository (owner/repo)
base_commit: Commit hash to checkout
file_path: Path to the file containing the target function… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-gym/hybrid_gym_func_localize_raw.D3-Gym-Trajectorieshybrid_gym_dep_search_raw
Hybrid Gym: Dependency Search Dataset
Dataset for the Dependency Search benchmark task. Each instance contains a target function and its ground-truth dependencies (functions/classes directly called by that function) within a repository.
See the benchmark README for usage instructions.
Fields
instance_id: Unique identifier
repo: GitHub repository (owner/repo)
base_commit: Commit hash to checkout
target_function_name: Name of the target function
target_function_file: Path… See the full description on the dataset page: https://huggingface.co/datasets/hybrid-gym/hybrid_gym_dep_search_raw.mujoco-gymnasium-v5-benchmark
MuJoCo Gymnasium v5 Benchmark
PPO evaluation benchmark on MuJoCo Gymnasium v5 locomotion and manipulation environments.
Tasks
Task ID
Name
hopper-v5
Hopper-v5
halfcheetah-v5
HalfCheetah-v5
walker2d-v5
Walker2d-v5
ant-v5
Ant-v5
humanoid-v5
Humanoid-v5
reacher-v5
Reacher-v5
Usage
Submit evaluation results to this benchmark by adding .eval_results/*.yaml files
to your model repo, referencing this dataset's task IDs.
© 2026 ParamTatva.org
gym_datasetswe_gym_raw_localize_no_exec_claude37_441i_no_str_replacesgym_claims
