violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think 4,000 evaluation attempts: 250 tasks × 16 samples. Original samples 0–3 and twelve additional seeded samples 4–15. The train split contains evaluation records, not training data. The evaluated model is violetxi/qwen35-9b-harvey-v4-notes-conditioned-1m at revision 21c40ef032e9cf968482584191c0b8c269f70001. Cohort Attempts All-criteria-pass rate ± task-level SEM Original four 1,000 3.100% ±… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think.
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think
4,000 evaluation attempts: 250 tasks × 16 samples. Original samples 0–3 and twelve additional seeded samples 4–15. The train split contains evaluation records, not training data. The evaluated model is violetxi/qwen35-9b-harvey-v4-notes-conditioned-1m at revision 21c40ef032e9cf968482584191c0b8c269f70001.
An attempt scores 1 only if every rubric criterion passes. The headline averages the sixteen attempts per task, then the 250 task means. Error bars are one SEM across task means, not confidence intervals or variation across training seeds. All failures, context-limit outcomes and empty answers remain in the denominator. The original four scores and responses were preserved without regrading.
Responses and scores
- Default config,
train: all 4,000 attempts with exact messages, full thinking, tool observations, final answers, criterion verdicts and judge explanations. sample_batchdistinguishesoriginal_fourfromnew_twelve;(task_id, sample)uniquely identifies an attempt.task_mean_scoreuses all sixteen samples;original_task_mean_scorepreserves the previous four-sample aggregate.- Merged scores: sixteen-sample aggregate, per-task and per-attempt results, with original-four and new-twelve summaries.
original_fourconfig and existingdata/,raw/,judge/,scores.jsonandmanifest.jsonretain the original four-sample publication unchanged.resampling16/raw/new-twelve.jsonlandresampling16/judge/new-twelve.jsonlcontain the exact new generation and per-criterion judge records.- Original four-sample documentation records historical generation, grading and training provenance.
Preserved evaluation protocol
Historical 20-turn agent loop with glob, grep, read; thinking enabled, temperature 0.6, 8,192 output tokens per turn, 65,536 context. New collection uses native vLLM DP=8/TP=1 and 512 concurrent tasks, capped at 64 requests per engine. Inline thinking and tool tags remain in the saved assistant text. Prompts and task definitions match the original evaluation. Request seeds are sample*1000000 + integer(task_id)*1000 + turn; original request seeds were not recorded. Original and new cohorts use different runtime installations, so the separate cohort metrics are retained for comparison.
Judge: GPT-5.6-Sol, original per-criterion rubric prompt, 16,384-token judge output cap. Existing four-sample judgments remain exact; the extension uses the frozen adapter archived under resampling16/code/. No success filtering. Tasks and benchmark material are from Harvey AI's MIT-licensed synthetic legal benchmark. Full task definitions remain in tasks/.
