datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 5.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 1.3000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 4.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 2.2000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 3.1000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-1m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 7.0000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-30m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-5m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-5m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 3.4000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-5m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 8.0000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-100m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-10m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-10m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), graded
with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and
binary all-criteria-pass rule. Mean all-pass rate: 6.2000%.
The train split contains held-out evaluation records, not training examples.
Model and training mixture
The evaluated checkpoint is Qwen3.5-9B trained for two epochs on the… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-notecondtraj30-10m-historical-20t-think.
