violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 1.3000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 1.3000%. The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent runs. These are not new native-tool 200-turn Harvey runs. Source: violetxi/wmrl-v4-scale70-3m-agentic-eval-20t-think. Every original attempt, including failures and empty answers, is retained. messages, transcript, final answers and row ordering are unchanged. full_trajectories preserves non-system user/assistant messages; action_only_trajectories keeps observations and parsed commands/task_complete. The exact system prompt is retained in canonical messages.
The full saved final answer, or the original unwrapped fallback for non-done attempts, is supplied as response.md to the original Harvey scoring function. No 6,000-character truncation is applied. An empty saved answer remains empty. Task JSONs and original generation code are included. The rubric and task title are taken from the verified historical task archive. Judge source: Harvey revision 1dd81403b2fbb60596f7aea3fcecafad7bf73143; profile harvey-single-gpt-5.6-sol-v1. The API rejects temperature=0 for this model; the compatibility bridge retries without that parameter. Output schema, 16,384-token cap, prompt and default reasoning effort match the pinned judge. There is one judge and no judge averaging.
score is binary per attempt: 1 only when every criterion passes. Task scores average the four attempts; the headline averages the 250 task scores. criterion_pass_rate is a separate diagnostic. historical_score and historical_task_mean_score preserve old rubric-fraction scores and are not the same metric. Full criterion verdicts and reasoning are in criterion_results, judge_responses, and judge/results.jsonl.
Files
data/*.parquet: Dataset Viewer records, tasks, exact messages and new judgments.raw/*.jsonl: byte-for-byte original generations.tasks/*/task.json: all 250 complete task definitions and criteria.historical_code/: original generation/scoring source for provenance.judge/rubric_criterion.txt: exact current judge prompt.scores.json: new aggregate and per-task results.manifest.json: immutable source revision, hashes, and validation receipt.
