violetxi/tb21-eval-qwen35-rewritten-w005-10k-thinking-32k-timeout2x
qwen35-rewritten-w005-10k — Terminal-Bench 2.1 Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-rewritten-obs-wm-weight-0p05-10k-tacc through the served model ID qwen35-rewritten-w005-10k with Terminus-2. Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs. Result Trials: 445 Tasks / attempts: 89 × 5 Errored trials scored as zero:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-rewritten-w005-10k-thinking-32k-timeout2x.
qwen35-rewritten-w005-10k — Terminal-Bench 2.1
Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-rewritten-obs-wm-weight-0p05-10k-tacc through the served model ID qwen35-rewritten-w005-10k with Terminus-2.
Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs.
Result
- Trials: 445
- Tasks / attempts: 89 × 5
- Errored trials scored as zero: 228
- Agent timeouts / context-length events / output-cap events: 227 / 0 / 0
- Mean reward / Pass@1: 0.098876
- Pass@5: 0.213483
- Total input/output/cache tokens: 80106804 / 4169889 / 0
- Proactive summaries: 181
Protocol
Temperature 0.6, top-p
- Trials / assistant turns with incomplete thinking markup: 29 / 59
- Assistant-turn thinking-markup completeness: 0.992574
0.95, top-k omitted (disabled),max_tokens=32768,enable_thinking=true, max turns30, concurrency164.
The server context length was 65536; Terminus proactive summarization began at 40960 tokens. The run used timeout_multiplier=2, and all errored trials remain in the denominator with reward zero.
Viewer columns
messages and messages_without_thinking are aligned, strictly alternating user/assistant conversations. Terminal observations are user messages. steps and steps_without_thinking retain the detailed Harbor trajectory. For thinking mode, Qwen's prompt-supplied opening <think> marker is reconstructed while generated reasoning and </think> are preserved. Zero-tag generations are preserved byte-for-byte in both views without inference and are explicitly marked by thinking_markup_complete, untagged_assistant_turn_count, and thinking_markup_policy. trajectory_integrity reports which raw continuation segments were available; missing segments are never fabricated. Harbor's literal technical-difficulties fallback is marked with an empty <think></think> block only in the normalized full-thinking view; its count is recorded separately and the raw fallback text is retained.
Only verifier test names/statuses are published. Hidden verifier traces and source, authorization headers, API keys, and credential values are excluded or redacted.
