CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Sovesh /autoresearch-fr Autoresearch French Dataset Source: wikimedia/wikipedia (20231101.fr) · 105 shards · val: shard_00104.parquet text1M<n<10M0 likes434 downloads7mo agoHugging Face02evo-hq /autoresearch-novelty-bench Autoresearch Novelty Bench A benchmark for testing whether an autonomous AI research agent proposes novel, mechanism-distinct hypotheses that anticipate breakthroughs later found by other researchers. By Evo. Built on Prime Intellect's autonomous-speedrunning archive — two AI agents (Claude Code and Codex) competing on modded-nanogpt's optimization speedrun. What's in this dataset table rows description experiments.parquet 10,380 One row per training run —… See the full description on the dataset page: https://huggingface.co/datasets/evo-hq/autoresearch-novelty-bench.tabular10K<n<100K10 likes303 downloads4mo agoHugging Face03Sovesh /autoresearch-nl Autoresearch Dutch Dataset Source: wikimedia/wikipedia (20231101.nl) · 47 shards · val: shard_00046.parquet text1M<n<10M0 likes178 downloads7mo agoHugging Face04CollinL /perovskite-solar-cell-efficiency-autoresearch 🔬 Perovskite Solar Cell Text Corpus for Karpathy's autoresearch A 98.9 MB text corpus of perovskite solar cell scientific literature formatted for direct use with karpathy/autoresearch — the autonomous LLM-driven hyperparameter search framework that trains a GPT from scratch and has an AI agent iteratively modify train.py to minimize val_bpb (bits per byte). 📊 Dataset Stats Metric Value Total documents 19,730 Total text 98.9 MB (~103M characters)… See the full description on the dataset page: https://huggingface.co/datasets/CollinL/perovskite-solar-cell-efficiency-autoresearch.texttext-generation10K<n<100K0 likes175 downloads5mo agoHugging Face05Sovesh /autoresearch-es Autoresearch Spanish Dataset Source: wikimedia/wikipedia (20231101.es) · 91 shards · val: shard_00090.parquet text1M<n<10M0 likes167 downloads7mo agoHugging Face06Sovesh /autoresearch-zh Autoresearch Chinese Dataset Source: wikimedia/wikipedia (20231101.zh) · 50 shards · val: shard_00049.parquet text1M<n<10M0 likes146 downloads7mo agoHugging Face07Sovesh /autoresearch-de Autoresearch German Dataset Source: wikimedia/wikipedia (20231101.de) · 108 shards · val: shard_00107.parquet text1M<n<10M0 likes103 downloads7mo agoHugging Face08sebastianboehler /autoresearch-manim Autoresearch Manim Curated Manim code-generation examples exported from the autoresearch_manim_finetune pipeline. Preview Gallery Preview Preview Preview Machine learning: attention plus residual mixing Physics: boundary layer flow near a surface Biology: neuron structure and signal direction Finance: compound growth over time Economics: production frontier tradeoff Neuroscience: action potential phases Summary Focus:… See the full description on the dataset page: https://huggingface.co/datasets/sebastianboehler/autoresearch-manim.imagetext-generationn<1K0 likes91 downloads2mo agoHugging Face09Sovesh /autoresearch-ja Autoresearch Japanese Dataset Source: wikimedia/wikipedia (20231101.ja) · 131 shards · val: shard_00130.parquet text1M<n<10M0 likes33 downloads7mo agoHugging Face10Sovesh /autoresearch-gu Autoresearch Gujarati Dataset Source: wikimedia/wikipedia (20231101.gu) · 3 shards · val: shard_00002.parquet text10K<n<100K0 likes26 downloads7mo agoHugging Face11davegraham /autoresearch-experiments Autoresearch Cross-Platform Experiments Dataset Description This dataset contains 2,637 hyperparameter optimization experiments from an autonomous LLM-driven ML research project. An LLM agent (Claude Sonnet) autonomously proposes hyperparameter modifications, trains a small language model for 5 minutes, evaluates validation bits-per-byte (val_bpb), and iterates. Experiments span 3 hardware platforms, 5 GPU models, and 7 text datasets, making this a unique resource for… See the full description on the dataset page: https://huggingface.co/datasets/davegraham/autoresearch-experiments.tabulartabular-regression1K<n<10K0 likes21 downloads6mo agoHugging Face12ProlificAI /autoresearch-hitl-annotations Autoresearch × Prolific HITL dataset Annotation dataset from the study "When does autoresearch need a human?" — a case study running Karpathy's autoresearch on a DPO task and evaluating the resulting models with 300 Prolific participants. Full interactive report covers per-pair stats, Bradley-Terry ranking, LLM-clustered comment themes, and methodology. What's in this dataset Two configs: annotations (default, annotations.parquet) — 1,507 rows. Each row is one… See the full description on the dataset page: https://huggingface.co/datasets/ProlificAI/autoresearch-hitl-annotations.tabular1K<n<10K3 likes16 downloads3mo agoHugging Face13Sovesh /autoresearch-or Autoresearch Odia Dataset Source: wikimedia/wikipedia (20231101.or) · 2 shards · val: shard_00001.parquet text10K<n<100K0 likes14 downloads7mo agoHugging Face14zjhhhh /DeepScaleR-Autoresearch-responses-ver0tabular1K<n<10K0 likes14 downloads3mo agoHugging Face15zjhhhh /DeepScaleR-Autoresearch-codebook-ver1tabularn<1K0 likes12 downloads3mo agoHugging Face16zjhhhh /DeepScaleR-Autoresearch-responses-ver1tabular1K<n<10K0 likes12 downloads3mo agoHugging Face17Sovesh /autoresearch-hi Autoresearch Hindi Dataset Source: wikimedia/wikipedia (20231101.hi) · 13 shards · val: shard_00012.parquet text100K<n<1M0 likes11 downloads7mo agoHugging Face18reasoning-degeneration-dev /adaevolve-autoresearch-smoketest-3itertabularn<1K0 likes8 downloads6mo agoHugging Face19mendax0110 /autoresearchpp autoresearch-cpp C++20 / LibTorch port of karpathy/autoresearch. An AI agent modifies train source files, builds, runs a fixed-budget experiment, and keeps changes only when val_bpb improves. Requirements CMake >= 3.25 C++20 compiler (GCC 12+, Clang 15+, MSVC 19.34+) LibTorch (CPU, CUDA, or MPS build) NVIDIA GPU optional — CPU and Apple MPS are fully supported Setup Linux + CUDA 1. Download LibTorch: wget… See the full description on the dataset page: https://huggingface.co/datasets/mendax0110/autoresearchpp.textn<1K0 likes8 downloads5mo agoHugging Face20zjhhhh /DeepScaleR-Autoresearch-codebook-ver0tabularn<1K0 likes7 downloads3mo agoHugging Face21Creekside /lambda-text-autoresearchtext10K<n<100K0 likes6 downloads6mo agoHugging Face22Yy245 /AutoResearch-paper-md AutoResearch Paper Markdown This repository contains MinerU-converted markdown files for the AutoResearch paper corpus. The markdown files are stored inside tar shards to make upload/download reliable on Hugging Face. Only the following content is uploaded by the accompanying script: markdown_shards/markdown-*.tar paper_index.json README.md Original PDF files are intentionally not uploaded to this repository. Summary Item Count / Size Markdown files… See the full description on the dataset page: https://huggingface.co/datasets/Yy245/AutoResearch-paper-md.text10K<n<100K0 likes6 downloads4mo agoHugging Face23reasoning-degeneration-dev /adaevolve-autoresearch-run1tabularn<1K0 likes3 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.