CoolFace
17 results

rl-environment

AdithyaSK /data_agent_rl_environment_train_multireward AdithyaSK/data_agent_rl_environment_train_multireward Multi-reward variant of AdithyaSK/data_agent_rl_environment_train (2238 tasks, identical data/instructions). The only change: each task's verifier now emits a reward.json with three named rewards instead of a single float: reward meaning range correctness graded answer matches gold (exact / numeric / LLM-judge) 0 or 1 submission a non-empty answer was written to /workdir/answer.txt 0 or 1 tool_efficiency fewer… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent_rl_environment_train_multireward.1K<n<10K1 likes1.8k downloads2mo agoHugging Facespade-rl /SPADE-Environments-Qwen3-30B-Games SPADE generated environments: games Paper | Code | All artifacts Executable game environments written by the SPADE Environment Designer during the paper's 30B games self-play run. One Python file per environment; manifest.json records the generation checkpoint, training step, skill, and difficulty of each. Environments 3310 Training steps covered 113 (step 0 to 396) With skill label 3119 Designer / agent model Qwen/Qwen3-30B-A3B-Instruct-2507… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environments-Qwen3-30B-Games.0 likes1.7k downloads27d agoHugging Facespade-rl /SPADE-Environment-Pool-GPT5.5-Games SPARE GPT-5.5 Grounded Cognitive Multi-Turn Games This public dataset contains 7,872 validated Python game environments for actor-only SPARE training. Six cognitive skills, exactly 1,312 environments per skill Generated with GPT-5.5 and grounded by spice_megascience_15k.jsonl Grounding corpus SHA-256: a36a928b4940b5b5d9e3f4cb5804a94c69462360943adb3be14613c82f0f72c0 Maximum 25 turns and 32K generation context Every environment passes load, reset, step, and replay validation with… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environment-Pool-GPT5.5-Games.reinforcement-learning1K<n<10K1 likes1.5k downloads27d agoHugging Facespade-rl /SPADE-Environments-ToolUse SPADE generated environments: tool use Paper | Code | All artifacts Multi-turn tool-use environments written by the SPADE designer during training, pooled across every captured run. 2,231 environments across 7 runs and two model scales (30B-A3B and 4B). Source run Scale Environments qwen3-30b-0617-tooluse-regen32-mixed 30B-A3B 41 qwen3-30b-0624-tooluse-blend 30B-A3B 243 qwen3-30b-0703-tooluse-glory-kl005 30B-A3B 260 qwen3-4b-0630-tooluse-eval-aligned-r32 4B 456… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environments-ToolUse.2 likes1.1k downloads27d agoHugging FaceAdithyaSK /data_agent_rl_environment_eval data_agent_rl_environment_eval The official verified eval suite for the data-agent RL pipeline. 366 Harbor-format data-analysis tasks, each with an LLM-assigned difficulty label (L1–L5), a Kaggle dataset dependency, and a tested reward function. 💡 Browse this dataset in your browser — click the badge above or open AdithyaSK/harbor-visualiser to inspect every task's spec, instruction, environment, tests, and difficulty. Reproduce the eval — end to end The… See the full description on the dataset page: https://huggingface.co/datasets/AdithyaSK/data_agent_rl_environment_eval.n<1K3 likes782 downloads2mo agoHugging Facespade-rl /SPADE-Environment-Pool-GPT5.5-ToolUse SPARE GPT-5.5 Multi-Turn Tool-Use Games v1 A public static pool of 11,039 validated multi-turn tool-use environments generated by GPT-5.5 for SPARE actor training. Training alignment Source recipe: Qwen3-30B-A3B 0624 tool-use GAMES configuration 400 rollouts x 24 games/rollout = 9,600 no-reuse games required 11,039 validated games provide 1,439 games of headroom Six balanced skills: API orchestration, data retrieval, state modification, error recovery, tool… See the full description on the dataset page: https://huggingface.co/datasets/spade-rl/SPADE-Environment-Pool-GPT5.5-ToolUse.tabularreinforcement-learning10K<n<100K1 likes672 downloads27d agoHugging Face