violetxi/tb21-eval-qwen35-action-only-40k-c164-max32k-timeout2x
qwen35-action-only-40k — Terminal-Bench 2.1 Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-action-only-40k-tacc through the served model ID qwen35-action-only-40k with Terminus-2. Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs. Result Recorded trials: 445 Tasks / attempts: 89 × 5 Errored trials scored as zero: 211 Exception… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-action-only-40k-c164-max32k-timeout2x.
qwen35-action-only-40k — Terminal-Bench 2.1
Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-action-only-40k-tacc through the served model ID qwen35-action-only-40k with Terminus-2.
Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs.
Result
- Recorded trials: 445
- Tasks / attempts: 89 × 5
- Errored trials scored as zero: 211
- Exception counts: `{"AgentTimeoutError": 211}`
- Agent timeouts / context-length events / output-cap events: 211 / 0 / 0
- Mean reward / Pass@1: 0.110112
- Pass@5: 0.213483
- Total input/output/cache tokens: 76827439 / 4205908 / 0
- Proactive summaries: 211
Protocol
Temperature 0.6, top-p 0.95, top-k omitted (disabled), max_tokens=32768, enable_thinking=true, max turns 30, concurrency 164.
- Trials / untagged / malformed assistant turns with incomplete thinking markup: 42 / 91 / 0
- Assistant-turn thinking-markup completeness: 0.988199
The server context length was 65536; Terminus proactive summarization began at 40960 tokens. Task-native TB2.1 timeouts were scaled (timeout_multiplier=2.0), so this is a noncanonical evaluation. All errored trials remain in the denominator with reward zero.
Viewer columns
messages and messages_without_thinking are aligned, strictly alternating user/assistant conversations. Terminal observations are user messages. steps and steps_without_thinking retain the detailed Harbor trajectory. For thinking mode, Qwen's prompt-supplied opening <think> marker is reconstructed while generated reasoning and </think> are preserved. Zero-tag generations are preserved byte-for-byte in both views without inference and are explicitly marked by thinking_markup_complete, untagged_assistant_turn_count, and thinking_markup_policy. trajectory_integrity reports which raw continuation segments were available; missing segments are never fabricated. Harbor's literal technical-difficulties fallback is marked with an empty <think></think> block only in the normalized full-thinking view; its count is recorded separately and the raw fallback text is retained.
Only verifier test names/statuses are published. Hidden verifier traces and source, authorization headers, API keys, and credential values are excluded or redacted.
