CoolFace
Datasetpublic

whywhywhyyy/qwen36-35b-a3b-clawbench

Qwen3.6-35B-A3B ClawBench Full Results This repository contains organized ClawBench run artifacts from bench/ClawBench/test-output/qwen36-35b-a3b-full-20260510-2140. The formal result directories are preserved under results/. Quarantined infrastructure-failure directories and top-level runtime control files such as .pid, .proc, .logpath, and .monitor-hermes-state/ are intentionally excluded from this organized export. Contents… See the full description on the dataset page: https://huggingface.co/datasets/whywhywhyyy/qwen36-35b-a3b-clawbench.

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
2likes210downloads
Dataset Card

Qwen3.6-35B-A3B ClawBench Full Results

This repository contains organized ClawBench run artifacts from bench/ClawBench/test-output/qwen36-35b-a3b-full-20260510-2140.

The formal result directories are preserved under results/. Quarantined infrastructure-failure directories and top-level runtime control files such as .pid, .proc, .logpath, and .monitor-hermes-state/ are intentionally excluded from this organized export.

Contents

  • —results/<group>/Qwen3.6-35B-A3B/<run>/: run artifacts including run-meta.json, data/actions.jsonl, data/requests.jsonl, data/agent-messages.jsonl, data/interception.json, screenshots, and recording.mp4 when present.
  • —results/<group>/batch-logs/: original batch logs for that group.
  • —metadata/runs.csv: one row per run with parsed run-meta.json fields and relative artifact paths.
  • —metadata/runs.jsonl: one JSON object per parsed run, including the full run-meta.json payload.
  • —metadata/aggregate.json: aggregate counts, pass rates, duplicate-case notes, and byte/file totals.
  • —metadata/missing-run-meta.csv: run directories found without run-meta.json.
  • —metadata/file_manifest.csv: relative paths and file sizes for the organized export.
  • —metadata/batch-summaries/: original batch-summary.json files copied into one place.
  • —source-logs/: top-level experiment and restart logs from the source directory.

Summary

GroupSuiteRun dirsParsed runsMissing metadataPassInterceptedPass rateBytes
openclaw-v1v11591545202012.99%7528939435
openclaw-v2v21321302545441.54%8483743682
hermes-v1v11551532181811.76%6038334965
hermes-v2v21301300454534.62%5080753019
Total-576567913713724.16%27131771101

Notes:

  • —Pass/intercept counts are computed from parsed run-meta.json files, not from batch-summary.json, because some batches were resumed and mark previously completed jobs as skipped in the batch summary.
  • —openclaw-v1 contains one duplicate parsed case (736-home-services-maintenance-plumbing-ace-hardware) and five run directories without run-meta.json; see metadata/aggregate.json and metadata/missing-run-meta.csv.
  • —The run artifacts can include synthetic user profile traces and disposable email addresses generated by ClawBench.
whywhywhyyy/qwen36-35b-a3b-clawbench · CoolFace