datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
autoresearch-fr
Autoresearch French Dataset
Source: wikimedia/wikipedia (20231101.fr) · 105 shards · val: shard_00104.parquet
autoresearch-novelty-bench
Autoresearch Novelty Bench
A benchmark for testing whether an autonomous AI research agent proposes
novel, mechanism-distinct hypotheses that anticipate breakthroughs
later found by other researchers.
By Evo. Built on Prime Intellect's
autonomous-speedrunning archive
— two AI agents (Claude Code and Codex) competing on modded-nanogpt's
optimization speedrun.
What's in this dataset
table
rows
description
experiments.parquet
10,380
One row per training run —… See the full description on the dataset page: https://huggingface.co/datasets/evo-hq/autoresearch-novelty-bench.autoresearch-nl
Autoresearch Dutch Dataset
Source: wikimedia/wikipedia (20231101.nl) · 47 shards · val: shard_00046.parquet
perovskite-solar-cell-efficiency-autoresearch
🔬 Perovskite Solar Cell Text Corpus for Karpathy's autoresearch
A 98.9 MB text corpus of perovskite solar cell scientific literature formatted for direct use with karpathy/autoresearch — the autonomous LLM-driven hyperparameter search framework that trains a GPT from scratch and has an AI agent iteratively modify train.py to minimize val_bpb (bits per byte).
📊 Dataset Stats
Metric
Value
Total documents
19,730
Total text
98.9 MB (~103M characters)… See the full description on the dataset page: https://huggingface.co/datasets/CollinL/perovskite-solar-cell-efficiency-autoresearch.autoresearch-es
Autoresearch Spanish Dataset
Source: wikimedia/wikipedia (20231101.es) · 91 shards · val: shard_00090.parquet
autoresearch-zh
Autoresearch Chinese Dataset
Source: wikimedia/wikipedia (20231101.zh) · 50 shards · val: shard_00049.parquet
autoresearch-de
Autoresearch German Dataset
Source: wikimedia/wikipedia (20231101.de) · 108 shards · val: shard_00107.parquet
autoresearch-manim
Autoresearch Manim
Curated Manim code-generation examples exported from the autoresearch_manim_finetune pipeline.
Preview Gallery
Preview
Preview
Preview
Machine learning: attention plus residual mixing
Physics: boundary layer flow near a surface
Biology: neuron structure and signal direction
Finance: compound growth over time
Economics: production frontier tradeoff
Neuroscience: action potential phases
Summary
Focus:… See the full description on the dataset page: https://huggingface.co/datasets/sebastianboehler/autoresearch-manim.autoresearch-ja
Autoresearch Japanese Dataset
Source: wikimedia/wikipedia (20231101.ja) · 131 shards · val: shard_00130.parquet
autoresearch-gu
Autoresearch Gujarati Dataset
Source: wikimedia/wikipedia (20231101.gu) · 3 shards · val: shard_00002.parquet
autoresearch-experiments
Autoresearch Cross-Platform Experiments
Dataset Description
This dataset contains 2,637 hyperparameter optimization experiments from an autonomous LLM-driven ML research project. An LLM agent (Claude Sonnet) autonomously proposes hyperparameter modifications, trains a small language model for 5 minutes, evaluates validation bits-per-byte (val_bpb), and iterates.
Experiments span 3 hardware platforms, 5 GPU models, and 7 text datasets, making this a unique resource for… See the full description on the dataset page: https://huggingface.co/datasets/davegraham/autoresearch-experiments.autoresearch-hitl-annotations
Autoresearch × Prolific HITL dataset
Annotation dataset from the study "When does autoresearch need a human?" — a case study running Karpathy's autoresearch on a DPO task and evaluating the resulting models with 300 Prolific participants. Full interactive report covers per-pair stats, Bradley-Terry ranking, LLM-clustered comment themes, and methodology.
What's in this dataset
Two configs:
annotations (default, annotations.parquet) — 1,507 rows. Each row is one… See the full description on the dataset page: https://huggingface.co/datasets/ProlificAI/autoresearch-hitl-annotations.autoresearch-or
Autoresearch Odia Dataset
Source: wikimedia/wikipedia (20231101.or) · 2 shards · val: shard_00001.parquet
DeepScaleR-Autoresearch-responses-ver0DeepScaleR-Autoresearch-codebook-ver1DeepScaleR-Autoresearch-responses-ver1autoresearch-hi
Autoresearch Hindi Dataset
Source: wikimedia/wikipedia (20231101.hi) · 13 shards · val: shard_00012.parquet
adaevolve-autoresearch-smoketest-3iterautoresearchpp
autoresearch-cpp
C++20 / LibTorch port of karpathy/autoresearch. An AI agent modifies train source files, builds, runs a fixed-budget experiment, and keeps changes only when val_bpb improves.
Requirements
CMake >= 3.25
C++20 compiler (GCC 12+, Clang 15+, MSVC 19.34+)
LibTorch (CPU, CUDA, or MPS build)
NVIDIA GPU optional — CPU and Apple MPS are fully supported
Setup
Linux + CUDA
1. Download LibTorch:
wget… See the full description on the dataset page: https://huggingface.co/datasets/mendax0110/autoresearchpp.DeepScaleR-Autoresearch-codebook-ver0lambda-text-autoresearchAutoResearch-paper-md
AutoResearch Paper Markdown
This repository contains MinerU-converted markdown files for the AutoResearch paper corpus.
The markdown files are stored inside tar shards to make upload/download reliable on Hugging Face.
Only the following content is uploaded by the accompanying script:
markdown_shards/markdown-*.tar
paper_index.json
README.md
Original PDF files are intentionally not uploaded to this repository.
Summary
Item
Count / Size
Markdown files… See the full description on the dataset page: https://huggingface.co/datasets/Yy245/AutoResearch-paper-md.adaevolve-autoresearch-run1
