CoolFace
Datasetpublic

violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think

harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 4.0000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.

sourceHugging Facemitupdated 6d agoView on Hugging Face
0likes283downloads
5 commits on main
f7e50ba6d ago

Update title and manifest for one-time historical mixture label

violetxi
b95eb416d ago

Update dataset title and manifest for Harvey evaluation naming

violetxi
f24af256d ago

Update dataset title and manifest for Harvey evaluation naming

violetxi
ba2b2ea8d ago

Regrade historical answers with gpt-5.6-sol and preserve exact generations

violetxi
d2f1fcc8d ago

initial commit

violetxi