evo-hq/autoresearch-novelty-bench
Autoresearch Novelty Bench A benchmark for testing whether an autonomous AI research agent proposes novel, mechanism-distinct hypotheses that anticipate breakthroughs later found by other researchers. By Evo. Built on Prime Intellect's autonomous-speedrunning archive — two AI agents (Claude Code and Codex) competing on modded-nanogpt's optimization speedrun. What's in this dataset table rows description experiments.parquet 10,380 One row per training… See the full description on the dataset page: https://huggingface.co/datasets/evo-hq/autoresearch-novelty-bench.
This repository belongs to evo-hq on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
