violetxi/tb21-eval-qwen35-action-only-40k-c164-max32k-timeout8x
qwen35-action-only-40k — Terminal-Bench 2.1 Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-action-only-40k-tacc through the served model ID qwen35-action-only-40k with Terminus-2. Noncanonical run: timeout_multiplier=8 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs. Result Recorded trials: 445 Tasks / attempts: 89 × 5 Errored trials scored as zero: 111 Exception… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-action-only-40k-c164-max32k-timeout8x.
qwen35-action-only-40k — Terminal-Bench 2.1
Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-action-only-40k-tacc through the served model ID qwen35-action-only-40k with Terminus-2.
Noncanonical run: timeout_multiplier=8 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs.
Result
- Recorded trials: 445
- Tasks / attempts: 89 × 5
- Errored trials scored as zero: 111
- Exception counts: `{"AgentTimeoutError": 107, "Timeout": 4}`
- Agent timeouts / context-length events / output-cap events: 107 / 0 / 0
- Mean reward / Pass@1: 0.125843
- Pass@5: 0.235955
- Total input/output/cache tokens: 106941251 / 6683245 / 0
- Proactive summaries: 469
Protocol
Temperature 0.6, top-p 0.95, top-k omitted (disabled), max_tokens=32768, enable_thinking=true, max turns 30, concurrency 164.
- Trials / untagged / malformed assistant turns with incomplete thinking markup: 64 / 261 / 0
- Assistant-turn thinking-markup completeness: 0.974263
The server context length was 65536; Terminus proactive summarization began at 40960 tokens. Task-native TB2.1 timeouts were scaled (timeout_multiplier=8.0), so this is a noncanonical evaluation. All errored trials remain in the denominator with reward zero.
Viewer columns
messages and messages_without_thinking are aligned, strictly alternating user/assistant conversations. Terminal observations are user messages. steps and steps_without_thinking retain the detailed Harbor trajectory. For thinking mode, Qwen's prompt-supplied opening <think> marker is reconstructed while generated reasoning and </think> are preserved. Zero-tag generations are preserved byte-for-byte in both views without inference and are explicitly marked by thinking_markup_complete, untagged_assistant_turn_count, and thinking_markup_policy. trajectory_integrity reports which raw continuation segments were available; missing segments are never fabricated. Harbor's literal technical-difficulties fallback is marked with an empty <think></think> block only in the normalized full-thinking view; its count is recorded separately and the raw fallback text is retained.
Only verifier test names/statuses are published. Hidden verifier traces and source, authorization headers, API keys, and credential values are excluded or redacted.
