CoolFace
Datasetpublic

evo-hq/autoresearch-novelty-bench

Autoresearch Novelty Bench A benchmark for testing whether an autonomous AI research agent proposes novel, mechanism-distinct hypotheses that anticipate breakthroughs later found by other researchers. By Evo. Built on Prime Intellect's autonomous-speedrunning archive — two AI agents (Claude Code and Codex) competing on modded-nanogpt's optimization speedrun. What's in this dataset table rows description experiments.parquet 10,380 One row per training… See the full description on the dataset page: https://huggingface.co/datasets/evo-hq/autoresearch-novelty-bench.

sourceHugging Faceapache-2.0updated 4mo agoView on Hugging Face
10likes396downloads
3 commits on main
14cc65f4mo ago

Add judge-backend note to usage block

alok97
27b775d4mo ago

initial dataset push: experiments, snapshots, embeddings, recovered variants, current recipes

alok97
47799ae4mo ago

initial commit

alok97