CoolFace
Datasetpublic

violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think

harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 2.2000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-1m-historical-20t-think.

sourceHugging Facemitupdated 6d agoView on Hugging Face
0likes102downloads
5 commits on main
6b35e1b6d ago

Update title and manifest for one-time historical mixture label

violetxi
59dbe5c6d ago

Update dataset title and manifest for Harvey evaluation naming

violetxi
77cf1956d ago

Update dataset title and manifest for Harvey evaluation naming

violetxi
d2df3598d ago

Regrade historical answers with gpt-5.6-sol and preserve exact generations

violetxi
3aacba68d ago

initial commit

violetxi