rl-environment
data_agent_rl_environment_train_multireward
AdithyaSK/data_agent_rl_environment_train_multireward
Multi-reward variant of AdithyaSK/data_agent_rl_environment_train (2238 tasks, identical
data/instructions). The only change: each task's verifier now emits a reward.json with
three named rewards instead of a single float:
reward
meaning
range
correctness
graded answer matches gold (exact / numeric / LLM-judge)
0 or 1
submission
a non-empty answer was written to /workdir/answer.txt
0 or 1
tool_efficiency
fewer… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent_rl_environment_train_multireward.SPADE-Environments-Qwen3-30B-Games
SPADE generated environments: games
Paper | Code | All artifacts
Executable game environments written by the SPADE Environment Designer during the paper's 30B games self-play run. One Python file per environment; manifest.json records the generation checkpoint, training step, skill, and difficulty of each.
Environments
3310
Training steps covered
113 (step 0 to 396)
With skill label
3119
Designer / agent model
Qwen/Qwen3-30B-A3B-Instruct-2507… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environments-Qwen3-30B-Games.SPADE-Environment-Pool-GPT5.5-Games
SPARE GPT-5.5 Grounded Cognitive Multi-Turn Games
This public dataset contains 7,872 validated Python game environments for actor-only SPARE training.
Six cognitive skills, exactly 1,312 environments per skill
Generated with GPT-5.5 and grounded by spice_megascience_15k.jsonl
Grounding corpus SHA-256: a36a928b4940b5b5d9e3f4cb5804a94c69462360943adb3be14613c82f0f72c0
Maximum 25 turns and 32K generation context
Every environment passes load, reset, step, and replay validation with… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environment-Pool-GPT5.5-Games.SPADE-Environments-ToolUse
SPADE generated environments: tool use
Paper | Code | All artifacts
Multi-turn tool-use environments written by the SPADE designer during training, pooled
across every captured run. 2,231 environments across 7 runs and two model scales (30B-A3B and 4B).
Source run
Scale
Environments
qwen3-30b-0617-tooluse-regen32-mixed
30B-A3B
41
qwen3-30b-0624-tooluse-blend
30B-A3B
243
qwen3-30b-0703-tooluse-glory-kl005
30B-A3B
260
qwen3-4b-0630-tooluse-eval-aligned-r32
4B
456… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environments-ToolUse.data_agent_rl_environment_eval
data_agent_rl_environment_eval
The official verified eval suite for the data-agent RL pipeline. 366 Harbor-format
data-analysis tasks, each with an LLM-assigned difficulty label (L1–L5), a Kaggle
dataset dependency, and a tested reward function.
💡 Browse this dataset in your browser — click the badge above or open
AdithyaSK/harbor-visualiser
to inspect every task's spec, instruction, environment, tests, and difficulty.
Reproduce the eval — end to end
The… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent_rl_environment_eval.SPADE-Environment-Pool-GPT5.5-ToolUse
SPARE GPT-5.5 Multi-Turn Tool-Use Games v1
A public static pool of 11,039 validated multi-turn tool-use environments generated by GPT-5.5 for SPARE actor training.
Training alignment
Source recipe: Qwen3-30B-A3B 0624 tool-use GAMES configuration
400 rollouts x 24 games/rollout = 9,600 no-reuse games required
11,039 validated games provide 1,439 games of headroom
Six balanced skills: API orchestration, data retrieval, state modification, error recovery, tool… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environment-Pool-GPT5.5-ToolUse.
