datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TaskTrove
TaskTrove
v5.1 (current) — independent-review source retirement — moves 15 sources with majority or unanimous REJECT verdicts out of the default config and into deprecated/. Three blinded reviewers each sampled 10 tasks per source from all 50 v5.0 source-drop candidates, read the instructions and packaged tests, and issued independent KEEP or REJECT verdicts. The 15 retired sources received at least two REJECT votes. The active catalog changes from 93 sources and 1,674,033… See the full description on the dataset page: https://huggingface.co/datasets/open-thoughts/TaskTrove.terminal_bench_2_tasktrove_dq_unitsyn_python_step20_30b_a3b_20260730_014827
TaskTrove DQ unitsyn-python training traces (step 20, 30B-A3B)
Terminus-2 agent rollouts recorded while training
laion/tasktrove-dq-unitsyn-python-step20-30b-a3b
with SkyRL from Qwen/Qwen3-Coder-30B-A3B-Instruct.
Each row is the last episode of one trial: the full agent transcript, the task instruction, the
scalar reward, and the verifier's output.
Source run: rl-tasktrove-dq-sweep-30b-terminus2-qwen-20260725-163115-1ae770.
Coverage
This dataset is the complete… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_unitsyn_python_step20_30b_a3b_20260730_014827.task-trove
TaskTrove Clean
TaskTrove Clean is a normalized release of
open-thoughts/TaskTrove at revision
0292300. It contains 861,848 retained Harbor tasks from
1,739,326 input rows. The graders were built from Marin commit
b76d03131c.
How it was made
The conversion pipeline applies these stages:
Pin the upstream Hugging Face revision and inventory each source's task templates.
Keep sources with recoverable task contracts and record every source decision.
Convert each… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/task-trove.TaskTrove
TaskTrove
TaskTrove is an open-source collection of agentic task datasets, released by the OpenThoughts-Agent team. It is the task complement to AgentTrove — the agent traces in AgentTrove were generated by running models against these task datasets using the Harbor framework.
v3.2 (current) — replaced the old swegym task dataset (laion__swegym-tasks-patched-validated-v2, 989 tasks) with laion/swegym-tasks-patched-validated-v5 (2,438 tasks, patched + validated). No other… See the full description on the dataset page: https://huggingface.co/datasets/Artificial-Production-Units/TaskTrove.
