violetxi/tb2-eval-qwen3-8b-wm-summary-klanchor
Terminal-Bench 2.0 eval results - Qwen3-8B WM summary KL-anchor SFT Compact Harbor eval results for violetxi/qwen3-8b-terminal-wm-summary-klanchor on Terminal-Bench 2.0 using the terminus-2 agent harness. Each dataset split is one checkpoint step. Rows are Harbor trial directories and include binary reward, per-test-case pass/fail data from verifier/ctrf.json, exception text when present, and run metadata. Raw terminal recordings, panes, and completion logs are not included.… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb2-eval-qwen3-8b-wm-summary-klanchor.
Terminal-Bench 2.0 eval results - Qwen3-8B WM summary KL-anchor SFT
Compact Harbor eval results for violetxi/qwen3-8b-terminal-wm-summary-klanchor on Terminal-Bench 2.0 using the terminus-2 agent harness.
Each dataset split is one checkpoint step. Rows are Harbor trial directories and include binary reward, per-test-case pass/fail data from verifier/ctrf.json, exception text when present, and run metadata. Raw terminal recordings, panes, and completion logs are not included.
Splits
Columns
reward:1.0iff all hidden tests passed; null if the trial did not reach verifier scoring.n_passed,n_tests,pct_tests_passed,test_cases: partial-credit verifier details.exception: truncated Harbor exception text for errored trials.run_*: metadata copied from the parent Harborresult.json.local_trial_dir: original path in the source cluster checkout for provenance.
