datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
terminal_bench_2_tasktrove_dq_unitsyn_python_step20_30b_a3b_20260730_014827
TaskTrove DQ unitsyn-python training traces (step 20, 30B-A3B)
Terminus-2 agent rollouts recorded while training
laion/tasktrove-dq-unitsyn-python-step20-30b-a3b
with SkyRL from Qwen/Qwen3-Coder-30B-A3B-Instruct.
Each row is the last episode of one trial: the full agent transcript, the task instruction, the
scalar reward, and the verifier's output.
Source run: rl-tasktrove-dq-sweep-30b-terminus2-qwen-20260725-163115-1ae770.
Coverage
This dataset is the complete… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_unitsyn_python_step20_30b_a3b_20260730_014827.cfa_extracted_exercise_sup_sample_from_policy_v1.1_stepwise_dpo_binarized_chunk_20flare_finqa_sup_sample_from_policy_v1.1_stepwise_dpo_chunk_20qwmathbase_full_raft_step20_minerva_mathDAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data-all33296-n4
DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data-all33296-n4
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-12b-pt-from-step60-seed43, subfolder step_000020
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 133,184
Unique prompts: 33,296
Responses per prompt: 4
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data-all33296-n4.DAPO-Gemma3-27B-PT-warmup20-step80-SFT-DataDAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data
DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data
Teacher-generated SFT/distillation data for Gemma 3 math distillation.
Source
Teacher: JWei05/dapo-gemma3-12b-pt-from-step60-seed43, subfolder step_000020
Prompts: JWei05/DAPO-OpenMathInstruct2-34k, train split
Rows: 66,592
Unique prompts: 33,296
Responses per prompt: 2
Sampling: temperature=1.0, top_p=1.0, top_k=-1, max_tokens=20480
Columns
Column
Description
messages
User prompt and teacher assistant… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/DAPO-Gemma3-12B-PT-RL-step20-seed43-SFT-Data.terminal_bench_2_Qwen3_32B_SweSmith_20step_20260317_215519sonnet_combined_all_traj_20stepterminal_bench_2_Qwen3_32B_SweSmith_20step_20260317_223743terminal_bench_2_Qwen3_32B_SweSmith_20step_20260318_064735terminal_bench_2_Qwen3_32B_SweSmith_20step_20260318_023437terminal_bench_2_Qwen3_32B_SweSmith_20step_20260501_180019terminal_bench_2_Qwen3_32B_SweSmith_20step_20260317_223400traindec10_stepsleft_binsof20_10ktraindec10_stepsleft_binsof20_5kterminal_bench_2_Qwen3_32B_SweSmith_20step_20260316_071051traindec10_stepsleft_binsof20qwmathbase_raftpp_bz128_step20_with_score_passncfa_extracted_exercise_sup_sample_from_policy_v1_1_rpo_stepwise_iter_1_stepwise_dpo_chunk_203k-forcing-400-022325-from20-to25-step3qwmathbase_raw_raft_step20_minerva_mathtraindec13_stepsleft_binsof20_irsub_10kqwmathbase_full_raft_step20_amc23dev_set_v2_r2egym_nl2bash_stack_bugsseq_lr3e_5_exp_rpt_stack_php_v2_step20_2026c99c20693k-forcing-400-022325-mask20-step2dev_set_71_tasks_nl2bash_nl2bash_bugsseq_Qwen3_8B_maxEps24_112925harbor_step20_80448101dev_set_v2_Qwen3_32B_SweSmith_20step_20260321_1658183k_forcing_400_mask20_step3_022525qwmathbase_full_raft_step20_olympiadbench
