CoolFace
Datasetpublic

DCAgent2/eval-glm-base-swe100-debug-tpu-20260704-174259

GLM-4.7-swesmith base — swebench-verified-random-100 (TPU-reproduced) Per-trial agent traces (full terminus-2 trajectory + verifier report) for the TPU-reproduced clean base evaluation used in the reproducibility check for Marin issue #6958. Model: laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink Benchmark: swebench-verified-random-100-folders (100 tasks × 3 reps = 300 trials) Score: 0.240 resolved (72/300) Rows: 300 (one per trial;… See the full description on the dataset page: https://huggingface.co/datasets/DCAgent2/eval-glm-base-swe100-debug-tpu-20260704-174259.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes16downloads
Dataset Card

GLM-4.7-swesmith base — swebench-verified-random-100 (TPU-reproduced)

Per-trial agent traces (full terminus-2 trajectory + verifier report) for the TPU-reproduced clean base evaluation used in the reproducibility check for Marin issue #6958.

  • —Model: laion/GLM-4_7-swesmith-sandboxes-with_tests-oracle_verified_120s-maxeps-131k-fixthink
  • —Benchmark: swebench-verified-random-100-folders (100 tasks × 3 reps = 300 trials)
  • —Score: 0.240 resolved (72/300)
  • —Rows: 300 (one per trial; episodes=last)
  • —Runtime: Marin Iris TPU cluster (re-run of the original SLURM/H100 reference eval)

This is the TPU-reproduced parity run. It agrees with the original reference eval (`DCAgent2/swebench_verified_random_100_folders_GLM_4_7_swesmith_sandboxes_with_tests_orac315f1e89`, 0.237 = 71/300) at 24%, backing the reproducibility claim. The 228/300 non-resolved trials are legitimate AgentTimeouts (0 summarization/harness errors), so 0.24 is a clean pass rate.

Schema mirrors the reference set: conversations, agent, model, model_provider, date, task, episode, run_id, trial_name, result, verifier_output.