CoolFace
Datasetpublic

violetxi/tb21-eval-qwen35-original-w005-10k-c164-max32k-timeout2x

qwen35-original-w005-10k — Terminal-Bench 2.1 Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-original-obs-wm-weight-0p05-10k-tacc through the served model ID qwen35-original-w005-10k with Terminus-2. Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs. Operator-directed failure: build-pov-ray__rpB5FQS was manually terminated and counted as… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-original-w005-10k-c164-max32k-timeout2x.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes19downloads
Dataset Card

qwen35-original-w005-10k — Terminal-Bench 2.1

Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-original-obs-wm-weight-0p05-10k-tacc through the served model ID qwen35-original-w005-10k with Terminus-2.

Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs.
Operator-directed failure: build-pov-ray__rpB5FQS was manually terminated and counted as reward zero.

Result

  • —Recorded trials: 445
  • —Tasks / attempts: 89 × 5
  • —Errored trials scored as zero: 213
  • —Exception counts: `{"AgentTimeoutError": 212, "CancelledError": 1}`
  • —Agent timeouts / context-length events / output-cap events: 212 / 0 / 0
  • —Mean reward / Pass@1: 0.110112
  • —Pass@5: 0.235955
  • —Total input/output/cache tokens: 81042805 / 4255476 / 0
  • —Proactive summaries: 246

Protocol

Temperature 0.6, top-p 0.95, top-k omitted (disabled), max_tokens=32763, enable_thinking=true, max turns 30, concurrency 164.

  • —Trials / untagged / malformed assistant turns with incomplete thinking markup: 28 / 67 / 0
  • —Assistant-turn thinking-markup completeness: 0.991775

The server context length was 65536; Terminus proactive summarization began at 40960 tokens. Task-native TB2.1 timeouts were scaled (timeout_multiplier=2.0), so this is a noncanonical evaluation. All errored trials remain in the denominator with reward zero.

Viewer columns

messages and messages_without_thinking are aligned, strictly alternating user/assistant conversations. Terminal observations are user messages. steps and steps_without_thinking retain the detailed Harbor trajectory. For thinking mode, Qwen's prompt-supplied opening <think> marker is reconstructed while generated reasoning and </think> are preserved. Zero-tag generations are preserved byte-for-byte in both views without inference and are explicitly marked by thinking_markup_complete, untagged_assistant_turn_count, and thinking_markup_policy. trajectory_integrity reports which raw continuation segments were available; missing segments are never fabricated. Harbor's literal technical-difficulties fallback is marked with an empty <think></think> block only in the normalized full-thinking view; its count is recorded separately and the raw fallback text is retained.

Only verifier test names/statuses are published. Hidden verifier traces and source, authorization headers, API keys, and credential values are excluded or redacted.