datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
recursive-tasktrove-out
AnshKetchum/tasktrove-recursive-task-synthesis
TaskTrove
TaskTrove
v5.1 (current) — independent-review source retirement — moves 15 sources with majority or unanimous REJECT verdicts out of the default config and into deprecated/. Three blinded reviewers each sampled 10 tasks per source from all 50 v5.0 source-drop candidates, read the instructions and packaged tests, and issued independent KEEP or REJECT verdicts. The 15 retired sources received at least two REJECT votes. The active catalog changes from 93 sources and 1,674,033… See the full description on the dataset page: https://huggingface.co/datasets/open-thoughts/TaskTrove.terminal_bench_2_tasktrove_dq_stack_pytest_step25_30b_a3b_20260730_053956
TaskTrove stack-pytest — training rollout traces (Qwen3-Coder-30B-A3B, step 25)
Terminus-2/Harbor rollouts recorded while training
laion/tasktrove-dq-stack-pytest-step25-30b-a3b
with SkyRL on the TaskTrove stack-pytest source.
One row per trial, holding that trial's last episode as an OpenAI-style conversations list, the
task instruction, the reward the verifier assigned (result), and the verifier's own stdout
(verifier_output).
Source run… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_stack_pytest_step25_30b_a3b_20260730_053956.terminal_bench_2_tasktrove_dq_unix_step10_30b_a3b_20260730_014756
terminal_bench_2_tasktrove_dq_unix_step10_30b_a3b
OpenCode agent trajectories from the TaskTrove DQ unix arm of a Qwen3-Coder-30B-A3B
agentic RL sweep, exported from the complete Harbor rollout artifact set.
Coverage
Built from the full trace_jobs prefix of run rl-tasktrove-dq-sweep-30b-qwen3-coder-30-20260727-082204-e42f1d
(12034 trial directories, 11937 of them scored).
quantity
value
scored trials (result.json)
11937
rows published
11937
coverage… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_unix_step10_30b_a3b_20260730_014756.tasktrove-hparam-optimization
TaskTrove hyperparameter optimization artifacts
This dataset is the campaign-level artifact archive for the TaskTrove hyperparameter and backend experiment. It
contains experiment configurations, reward and timing curves, evaluation records, checkpoint-cleanup records, the
A0 same-source base evaluation workspace, and the generated analysis website.
The experiment report, policy, and tracker remain in the parent experiment directory. See
PUBLISHED_ARTIFACTS.md for the separately… See the full description on the dataset page: https://huggingface.co/datasets/penfever/tasktrove-hparam-optimization.terminal_bench_2_tasktrove_dq_taco_step15_30b_a3b_20260729_222705
Agent trace dataset
OpenCode/Harbor rollout traces from the MarinSkyRL run
rl-tasktrove-dq-sweep-30b-qwen3-coder-30-20260726-235656-574ba8, exported with
make_and_upload_trace_dataset --episodes last (the last episode of each trial — the rollouts
the policy was trained on).
Coverage
Built from the complete trial set on durable object storage, not from a local evidence bundle.
quantity
value
trial directories on object storage
21711
trials with a… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_taco_step15_30b_a3b_20260729_222705.terminal_bench_2_tasktrove_dq_unitsyn_python_step20_30b_a3b_20260730_014827
TaskTrove DQ unitsyn-python training traces (step 20, 30B-A3B)
Terminus-2 agent rollouts recorded while training
laion/tasktrove-dq-unitsyn-python-step20-30b-a3b
with SkyRL from Qwen/Qwen3-Coder-30B-A3B-Instruct.
Each row is the last episode of one trial: the full agent transcript, the task instruction, the
scalar reward, and the verifier's output.
Source run: rl-tasktrove-dq-sweep-30b-terminus2-qwen-20260725-163115-1ae770.
Coverage
This dataset is the complete… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_unitsyn_python_step20_30b_a3b_20260730_014827.task-trove
TaskTrove Clean
TaskTrove Clean is a normalized release of
open-thoughts/TaskTrove at revision
0292300. It contains 861,848 retained Harbor tasks from
1,739,326 input rows. The graders were built from Marin commit
b76d03131c.
How it was made
The conversion pipeline applies these stages:
Pin the upstream Hugging Face revision and inventory each source's task templates.
Keep sources with recoverable task contracts and record every source decision.
Convert each… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/task-trove.terminal_bench_2_tasktrove_dq_tulu3_personas_math_step13_30b_a3b_20260730_054033
terminal_bench_2_tasktrove_dq_tulu3_personas_math_step13_30b_a3b
OpenCode agent traces from the Iris RL run
rl-tasktrove-dq-sweep-30b-qwen3-coder-30-20260727-143750-b2bcd7, exported from the
run's Harbor trace_jobs artifacts (last episode per trial).
Coverage is complete for this run: all 15,740 trial directories were enumerated and every
trial that produced a result.json is present. The 89 trials without a result.json
never completed a scoreable episode and contribute no rows.… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_tulu3_personas_math_step13_30b_a3b_20260730_054033.jupiter-tasktrove-dapo-artifacts
Jupiter TaskTrove DAPO — campaign artifacts
Standalone artifact archive for the jupiter-tasktrove-dapo campaign: a six-arm
objective ablation (GRPO control vs DAPO variants vs GSPO) training
Qwen/Qwen3-Coder-30B-A3B-Instruct with agentic RL (MarinSkyRL / SkyRL fully-async GRPO
trainers, Harbor + Daytona sandboxed terminus-2 rollouts) on competitive-programming
tasks, run on JSC Jupiter (GH200) 2026-08-20 → 2026-08-31.
The campaign closed inconclusive (platform degradation, a… See the full description on the dataset page: https://huggingface.co/datasets/penfever/jupiter-tasktrove-dapo-artifacts.tasktrove-bugsinpy-v4-oracle17-apptainer-v1
TaskTrove BugsInPy v4: oracle-verified Apptainer subset v1
This derived Harbor release contains 17 of 479 upstream tasks. Every included reference solution is grounded in the official BugsInPy patch and selected by verifier execution. The remaining tasks are retained in the exclusion ledger; they are not silently discarded.
The release targets offline, rootless Apptainer on aarch64. Validation evidence is stored under validation/. Do not describe the full 479-task source as… See the full description on the dataset page: https://huggingface.co/datasets/laion/tasktrove-bugsinpy-v4-oracle17-apptainer-v1.terminal_bench_2_tasktrove_dq_pymethods2test_step75_30b_a3b_20260730_054052
terminal_bench_2_tasktrove_dq_pymethods2test_step75_30b_a3b
Terminus-2 agent traces from the Iris RL run
rl-tasktrove-dq-sweep-30b-terminus2-qwen-20260725-231748-4fe0b4, exported from the
run's Harbor trace_jobs artifacts (last episode per trial).
Coverage is complete for this run: all 948 trial directories were enumerated and every
trial that produced a result.json is present. The 94 trials without a result.json
never completed a scoreable episode and contribute no rows.… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_pymethods2test_step75_30b_a3b_20260730_054052.tasktrove-curriculum-medium-v2-rootless-apptainer
TaskTrove Curriculum Medium v2: rootless Apptainer runtime release
This release contains 489 runnable Harbor tasks from open-thoughts/TaskTrove/DCAgent__exp_rpt_curriculum-medium-v2@946884702046be1a7dcea2638186ad6b0d2ea103 and bundles the exact aarch64 SIF images used on Jupiter with rootless Apptainer 1.4.5. Task contents are byte-identical to the selected upstream task archives.
0 source tasks are excluded for missing required task/verifier files. The source does not provide… See the full description on the dataset page: https://huggingface.co/datasets/laion/tasktrove-curriculum-medium-v2-rootless-apptainer.tasktrove-curriculum-easy-rootless-apptainer
TaskTrove Curriculum Easy: rootless Apptainer runtime release
This release contains 505 runnable Harbor tasks from open-thoughts/TaskTrove/DCAgent__exp_rpt_curriculum-easy@96567362fa3c41208e0954317c53767023b420eb and bundles the exact aarch64 SIF images used on Jupiter with rootless Apptainer 1.4.5. Task contents are byte-identical to the selected upstream task archives.
0 source tasks are excluded for missing required task/verifier files. The source does not provide reference… See the full description on the dataset page: https://huggingface.co/datasets/laion/tasktrove-curriculum-easy-rootless-apptainer.tasktrove-repro-x10-x2
TaskTrove reproducibility package — X10 fsdp2-fa2 and X2 clip 0.2
Artifacts for two arms of the TaskTrove hyperparameter campaign
(marin-community/marin#7785). The two arms run on different clusters with different
launchers, so each has its own section.
Both arms train Qwen/Qwen3-Coder-30B-A3B-Instruct on DCAgent/exp_rpt_multifile
through the SkyRL terminal-bench entrypoint with the Harbor agent harness.
Everything here is a record of what ran. Where a fact could not be… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/tasktrove-repro-x10-x2.TaskTrove
TaskTrove
TaskTrove is an open-source collection of agentic task datasets, released by the OpenThoughts-Agent team. It is the task complement to AgentTrove — the agent traces in AgentTrove were generated by running models against these task datasets using the Harbor framework.
v3.2 (current) — replaced the old swegym task dataset (laion__swegym-tasks-patched-validated-v2, 989 tasks) with laion/swegym-tasks-patched-validated-v5 (2,438 tasks, patched + validated). No other… See the full description on the dataset page: https://huggingface.co/datasets/Artificial-Production-Units/TaskTrove.snowball-tasktrove-v342-together-verifiers
