CoolFace
Datasetpublic

violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think

harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think 4,000 evaluation attempts: 250 tasks × 16 samples. Original samples 0–3 and twelve additional seeded samples 4–15. The train split contains evaluation records, not training data. The evaluated model is violetxi/qwen35-9b-harvey-v4-notes-conditioned-1m at revision 21c40ef032e9cf968482584191c0b8c269f70001. Cohort Attempts All-criteria-pass rate ± task-level SEM Original four 1,000 3.100% ±… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think.

sourceHugging Facemitupdated 2d agoView on Hugging Face
0likes108downloads
Dataset Card

harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think

4,000 evaluation attempts: 250 tasks × 16 samples. Original samples 0–3 and twelve additional seeded samples 4–15. The train split contains evaluation records, not training data. The evaluated model is violetxi/qwen35-9b-harvey-v4-notes-conditioned-1m at revision 21c40ef032e9cf968482584191c0b8c269f70001.

CohortAttemptsAll-criteria-pass rate ± task-level SEM
Original four1,0003.100% ± 0.914 pp
Additional twelve3,0004.267% ± 0.963 pp
Combined sixteen4,0003.975% ± 0.931 pp

An attempt scores 1 only if every rubric criterion passes. The headline averages the sixteen attempts per task, then the 250 task means. Error bars are one SEM across task means, not confidence intervals or variation across training seeds. All failures, context-limit outcomes and empty answers remain in the denominator. The original four scores and responses were preserved without regrading.

Responses and scores

  • —Default config, train: all 4,000 attempts with exact messages, full thinking, tool observations, final answers, criterion verdicts and judge explanations.
  • —sample_batch distinguishes original_four from new_twelve; (task_id, sample) uniquely identifies an attempt. task_mean_score uses all sixteen samples; original_task_mean_score preserves the previous four-sample aggregate.
  • —Merged scores: sixteen-sample aggregate, per-task and per-attempt results, with original-four and new-twelve summaries.
  • —original_four config and existing data/, raw/, judge/, scores.json and manifest.json retain the original four-sample publication unchanged.
  • —resampling16/raw/new-twelve.jsonl and resampling16/judge/new-twelve.jsonl contain the exact new generation and per-criterion judge records.
  • —Original four-sample documentation records historical generation, grading and training provenance.

Preserved evaluation protocol

Historical 20-turn agent loop with glob, grep, read; thinking enabled, temperature 0.6, 8,192 output tokens per turn, 65,536 context. New collection uses native vLLM DP=8/TP=1 and 512 concurrent tasks, capped at 64 requests per engine. Inline thinking and tool tags remain in the saved assistant text. Prompts and task definitions match the original evaluation. Request seeds are sample*1000000 + integer(task_id)*1000 + turn; original request seeds were not recorded. Original and new cohorts use different runtime installations, so the separate cohort metrics are retained for comparison.

Judge: GPT-5.6-Sol, original per-criterion rubric prompt, 16,384-token judge output cap. Existing four-sample judgments remain exact; the extension uses the frozen adapter archived under resampling16/code/. No success filtering. Tasks and benchmark material are from Harvey AI's MIT-licensed synthetic legal benchmark. Full task definitions remain in tasks/.