datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
exp_rpt_pymethods2test-large-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/exp_rpt_pymethods2test-large-qwen3.5-122b-131k-opencode-traces.a3-rl-DCAgent_exp_rpt_pymethods2test-largeablation-pymethods2test-seqnormablation-pymethods2test-seqnorm-tisablation-pymethods2test-seqnorm-tis-part2ablation-pymethods2test-shapedablation-pymethods2test-seqmean-arm0rl__24GPU_base__exp_rpt_pymethods2test-large__qwen3base-GLM-4_7-swglm52-datagen-r11-30-pymethods2test-large-tracesterminal_bench_2_a3_rl_DCAgent_exp_rpt_pymethods2test_v3_10_8B_20260830_021230rl__24GPU_base_excl_timeouts__exp_rpt_pymethods2test-large__GLM-4_7-swesmith-san__40-0-tracesrl__56GPU_base_staleclip__exp_rpt_pymethods2test-large__GLM-4_7-swesmith-san-tracesrl__24GPU_base__exp_rpt_pymethods2test-large__Qwen3-8B-Baserl__24GPU_shaped__exp_rpt_pymethods2test-large__exp_tas_optimal_combterminal_bench_2_a3_rl_DCAgent_exp_rpt_pymethods2test_v3_10_8B_20260830_174151swebench_verified_random_100_folders_a3_rl_DCAgent_exp_rpt_pymethods2test_v3_10_8Bc8b3394ddev_set_v2_a3_rl_DCAgent_exp_rpt_pymethods2test_v3_10_8B_20260826_040036terminal_bench_2_tasktrove_dq_pymethods2test_step75_30b_a3b_20260730_054052
terminal_bench_2_tasktrove_dq_pymethods2test_step75_30b_a3b
Terminus-2 agent traces from the Iris RL run
rl-tasktrove-dq-sweep-30b-terminus2-qwen-20260725-231748-4fe0b4, exported from the
run's Harbor trace_jobs artifacts (last episode per trial).
Coverage is complete for this run: all 948 trial directories were enumerated and every
trial that produced a result.json is present. The 94 trials without a result.json
never completed a scoreable episode and contribute no rows.… See the full description on the dataset page: https://huggingface.co/datasets/laion/terminal_bench_2_tasktrove_dq_pymethods2test_step75_30b_a3b_20260730_054052.exp_rpt_pymethods2test-v3-qwen3.5-122b-131k-opencode-traces
Agent trace dataset
Decoding the literal token IDs
The prompt_token_ids / completion_token_ids / logprobs columns are the
verbatim tokens the serving engine emitted, stored PER AGENT STEP as a
list-of-lists (one inner list per turn). To turn them back into text you MUST
use the exact tokenizer the model was served with — a generic same-family
tokenizer will decode word tokens to garbage.
Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/exp_rpt_pymethods2test-v3-qwen3.5-122b-131k-opencode-traces.maxgn09_hint__exp_rpt_pymethods2test-large__GLM-4_7-swesmith-san__OpenThoughts-TB-devrl__24GPU_base__exp_rpt_pymethods2test-large__r2egym-nl2bash-stack__40-0swebench_verified_random_100_folders_a3_rl_DCAgent_exp_rpt_pymethods2test_v3_10_8B42a94b69terminal_bench_2_a3_rl_DCAgent_exp_rpt_pymethods2test_v3_10_8B_20260829_151827dev_set_v2_a3_rl_DCAgent_exp_rpt_pymethods2test_v3_10_8B_20260826_223541swebench_verified_random_100_folders_a3_rl_DCAgent_exp_rpt_pymethods2test_v3_10_8B7ca492a3rl__56GPU_base_zclip__exp_rpt_pymethods2test-large__GLM-4_7-swesmith-san-tracesrl__24GPU_base_lr5e-6__exp_rpt_pymethods2test-large__GLM-4_7-swesmith-san__40-0-tracesgaia_127_rl__24GPU_shaped__exp_rpt_pymethods2test_large__GLM_4_7_swesmith_san_3ab1d39d3pymethods2test-large-qwen3.5-122b-32k-tracesgaia_127_rl_swesmith_fixthink_pymethods2test_45_20260526_201917
