CoolFace
Datasetpublic

violetxi/tb21-eval-qwen35-rewritten-w005-10k-thinking-32k-timeout2x

qwen35-rewritten-w005-10k — Terminal-Bench 2.1 Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-rewritten-obs-wm-weight-0p05-10k-tacc through the served model ID qwen35-rewritten-w005-10k with Terminus-2. Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs. Result Trials: 445 Tasks / attempts: 89 × 5 Errored trials scored as zero:… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/tb21-eval-qwen35-rewritten-w005-10k-thinking-32k-timeout2x.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes34downloads
Dataset Card

qwen35-rewritten-w005-10k — Terminal-Bench 2.1

Noncanonical Terminal-Bench 2.1 evaluation of violetxi/qwen35-4b-offline-echo-rewritten-obs-wm-weight-0p05-10k-tacc through the served model ID qwen35-rewritten-w005-10k with Terminus-2.

Noncanonical run: timeout_multiplier=2 instead of 1.0; concurrency=164 exceeds 30. Do not compare this score directly with canonical TB2.1 leaderboard runs.

Result

  • —Trials: 445
  • —Tasks / attempts: 89 × 5
  • —Errored trials scored as zero: 228
  • —Agent timeouts / context-length events / output-cap events: 227 / 0 / 0
  • —Mean reward / Pass@1: 0.098876
  • —Pass@5: 0.213483
  • —Total input/output/cache tokens: 80106804 / 4169889 / 0
  • —Proactive summaries: 181

Protocol

Temperature 0.6, top-p

  • —Trials / assistant turns with incomplete thinking markup: 29 / 59
  • —Assistant-turn thinking-markup completeness: 0.992574 0.95, top-k omitted (disabled), max_tokens=32768, enable_thinking=true, max turns 30, concurrency 164.

The server context length was 65536; Terminus proactive summarization began at 40960 tokens. The run used timeout_multiplier=2, and all errored trials remain in the denominator with reward zero.

Viewer columns

messages and messages_without_thinking are aligned, strictly alternating user/assistant conversations. Terminal observations are user messages. steps and steps_without_thinking retain the detailed Harbor trajectory. For thinking mode, Qwen's prompt-supplied opening <think> marker is reconstructed while generated reasoning and </think> are preserved. Zero-tag generations are preserved byte-for-byte in both views without inference and are explicitly marked by thinking_markup_complete, untagged_assistant_turn_count, and thinking_markup_policy. trajectory_integrity reports which raw continuation segments were available; missing segments are never fabricated. Harbor's literal technical-difficulties fallback is marked with an empty <think></think> block only in the normalized full-thinking view; its count is recorded separately and the raw fallback text is retained.

Only verifier test names/statuses are published. Hidden verifier traces and source, authorization headers, API keys, and credential values are excluded or redacted.