CoolFace
Datasetpublic

violetxi/tb2-eval-qwen3-8b-wm-summary-klanchor

Terminal-Bench 2.0 eval results - Qwen3-8B WM summary KL-anchor SFT Compact Harbor eval results for violetxi/qwen3-8b-terminal-wm-summary-klanchor on Terminal-Bench 2.0 using the terminus-2 agent harness. Each dataset split is one checkpoint step. Rows are Harbor trial directories and include binary reward, per-test-case pass/fail data from verifier/ctrf.json, exception text when present, and run metadata. Raw terminal recordings, panes, and completion logs are not included.… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb2-eval-qwen3-8b-wm-summary-klanchor.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes149downloads
Dataset Card

Terminal-Bench 2.0 eval results - Qwen3-8B WM summary KL-anchor SFT

Compact Harbor eval results for violetxi/qwen3-8b-terminal-wm-summary-klanchor on Terminal-Bench 2.0 using the terminus-2 agent harness.

Each dataset split is one checkpoint step. Rows are Harbor trial directories and include binary reward, per-test-case pass/fail data from verifier/ctrf.json, exception text when present, and run metadata. Raw terminal recordings, panes, and completion logs are not included.

Splits

splitrowsscoredresolvedmean pct tests passedrun finished
step_2508970021.0916yes
step250traj1900-yes
step_5008969017.4046yes
step500traj20100.0yes
step_7508970113.7073yes
step750traj1900-yes
step_10008970015.9616yes
step1000traj1900-yes
step_12508970014.6423yes
step1250traj1900-yes
step_15008969018.2079yes
step1500traj201050.0yes
step_17508968120.0184yes
step1750traj2100-yes
step_20008970018.7158yes
step2000traj1900-yes
step_22508970118.1959yes
step2250traj1900-yes
step_25008970018.9422yes
step2500traj1900-yes
step_27508970018.1441yes
step2750traj1900-yes
step_30008970119.113yes
step_32508968018.8385yes
step_35008968017.5788yes
step_37508970116.6844yes
step_40008970019.113yes
step_42508970017.6628yes
step_45008969016.903yes
step_47508970016.0049yes
step_50008970117.641yes

Columns

  • —reward: 1.0 iff all hidden tests passed; null if the trial did not reach verifier scoring.
  • —n_passed, n_tests, pct_tests_passed, test_cases: partial-credit verifier details.
  • —exception: truncated Harbor exception text for errored trials.
  • —run_*: metadata copied from the parent Harbor result.json.
  • —local_trial_dir: original path in the source cluster checkout for provenance.