datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AI-CUDA-Engineer-Archive
The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition
We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.arc-whestbench-public-2026
Organized by:
Alignment Research Center (ARC),
AIcrowd
WhestBench 2026: ARC White-Box Estimation Challenge
WhestBench is a benchmark for white-box activation estimation: given the weights of a randomly initialized ReLU multi-layer perceptron (MLP) and a strict floating-point-operation (FLOP) budget, predict the average post-activation value of every neuron when the network is fed standard Gaussian inputs.
This is the WhestBench 2026… See the full description on the dataset page: https://huggingface.co/datasets/aicrowd/arc-whestbench-public-2026.ArchEGraph
ArchEGraph
ArchEGraph is a building-energy dataset organized for graph-based and weather-conditioned learning.
Dataset Summary
Total cases in manifest.csv: 49,326
Unique buildings: 5,481
Unique weather IDs: 64
n_steps range: 968 to 8,760
n_spaces range: 1 to 231
This repository currently stores:
manifest.csv (index of all cases)
building/ (5,481 files)
geometry/ (5,482 files)
weather/ (64 files)
energy/ (49,326 files; nested under subfolders like 00/)
split/… See the full description on the dataset page: https://huggingface.co/datasets/ArchEGraph/ArchEGraph.scheduleSee https://github.com/ust-archive/ust-archive for more information.
rl-run-archive-2026
RL run archive 2026
Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer
checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for
long-term preservation and reproducibility.
Layout mirrors the verified backup trees they were copied from:
tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive
carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rl-run-archive-2026.planck-2018-chains
Planck 2018 Cosmological-Parameter Chains
This dataset contains the Planck Public Release 3 cosmological-parameter
full grid, COM_CosmoParams_fullGrid_R3.01.zip. It contains the Markov
chains and their GetDist and CosmoMC companions for 336 combinations of
cosmological model and likelihood or external-data selection. The 1,296
chain roots become 5,184 Parquet tables containing 27,699,519 rows.
The tables retain the source's headerless, positional structure. Arrow
fields are… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/planck-2018-chains.wmap-single-year-maps
WMAP DR5 Single-Year I/Q/U Maps
The preview renders the canonical year-1 K1 TEMPERATURE field in its source
NESTED order, using a Galactic Mollweide projection and a symmetric 99.5th
percentile colour range.
This dataset contains the complete full-resolution single-year I/Q/U release
served by LAMBDA: ten WMAP differencing assemblies for each of nine observing
years. These are year-specific, per-assembly measurements. They are distinct
from wmap-band-maps-9yr, whose five maps… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/wmap-single-year-maps.gwosc-o2-strain
GWOSC O2 16 kHz gravitational-wave strain
This dataset contains the 16,384 Hz H1, L1, and V1 strain records released by
the Gravitational Wave Open Science Center for the
second observing run (O2). Each detector's 1 Hz data-quality and
hardware-injection masks are separate configurations, preserving the source
cadences.
Observing run
O2
GPS extent
1164558336–1187737600
Detectors
H1 (Hanford), L1 (Livingston), and V1 (Virgo)
Strain
16,384 Hz, float64
Masks… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gwosc-o2-strain.crimson-hexagonal-archive
The Crimson Hexagonal Archive — machine-readable representation
Query this without downloading anything. Every config is served by the Hugging Face datasets-server over plain HTTP, no auth, no client library. Use /rows — it is the reliable one. It reads the parquet directly and answers in under two seconds:
https://datasets-server.huggingface.co/rows?dataset=leesharks%2Fcrimson-hexagonal-archive&config=deposits&split=train&offset=0&length=10… See the full description on the dataset page: https://huggingface.co/datasets/leesharks/crimson-hexagonal-archive.trellis500k-sketchfab-archivesmoltbook-observatory-archive
Observatory Dataset
This dataset is an incremental export of a SQLite observatory database, published as
date-partitioned Parquet files for efficient browsing and querying on Hugging Face.
For example, you can filter data by wildcards on date:
ds = load_dataset(
"SimulaMet/moltbook-observatory-archive",
"posts",
data_files="data/posts/2026-01-2*.parquet", # 20–29
split="train"
)
Each SQLite table is exposed as a separate dataset subset. Use dropdown above the… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/moltbook-observatory-archive.gaia-dr3-source
Gaia DR3 Source
This dataset mirrors the complete ESA Gaia Data Release 3 gaia_source
bulk-download table. It contains one record for every published Gaia source and
all 152 columns served by ESA, including source identifiers, astrometry,
photometry, observing statistics, quality fields, classifications, and
astrophysical parameters.
Each of ESA's 3,386 compressed ECSV shards is retained as one dataset
configuration. The configuration name is the source filename stem verbatim.… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gaia-dr3-source.meteogate-archive
Meteogate European Weather Observations Archive
A continuously growing archive of real-time meteorological observations from the
EUMETNET Meteogate E-SOH service,
covering thousands of weather stations across Europe.
Data Structure
Each Parquet file contains observations in long format (one row per
station × variable × timestamp) with the following columns:
Column
Type
Description
timestamp
datetime
Observation time in UTC
station_id
string
WIGOS… See the full description on the dataset page: https://huggingface.co/datasets/alexdum/meteogate-archive.gwosc-o1-strain
GWOSC O1 16 kHz gravitational-wave strain
This dataset contains the 16,384 Hz H1 and L1 strain records released by the
Gravitational Wave Open Science Center for Advanced
LIGO's first observing run (O1). Each detector's 1 Hz data-quality and
hardware-injection masks are included as separate configurations so that
measurements with different cadences remain separate.
Observing run
O1, GPS 1126051217–1137254417
Detectors
H1 (Hanford) and L1 (Livingston)
Strain
16… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gwosc-o1-strain.fermi-lat-weekly-photons
Fermi-LAT weekly photons
This dataset contains Fermi Large Area Telescope all-sky weekly photon files
from mission week w009 through w153, frozen on 2026-08-30. Its 145
configurations correspond one-to-one with the weekly p305_v001 FITS files.
Each Parquet row is an EVENTS row, with the 23 FITS-named columns in their
stored order and shape.
Mission weeks run Thursday through Wednesday in UTC. The first configuration
begins with the science-phase interval on 2008-08-04; w153… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/fermi-lat-weekly-photons.archive-dolma3-6t-dedup-state
archive-dolma3-6t-dedup-state
ARCHIVE: backup of the SOC-90 dedup state (Bloom filter + filtered docs). Reproducibility-only; current dedup state is in HCAI-Lab/dolma3-6t-bloom-index and dolma3-6t-unique.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/soc-90-keep-backup… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-6t-dedup-state.trellis500k-github-archives-10glenans-isobars-archivearc-agi-3-schema-traces
ARC-AGI-3 Schema Gameplay Trajectories
This release contains 50 ARC-AGI-3 gameplay trajectories and a dependency-free
scoring utility. The trajectories are split evenly across two collections:
gpt_5_6_sol/: 25 GPT-5.6 Sol trajectories.
claude_fable_opus/: 25 trajectories from Claude Opus 4.8 and Claude Fable 5.
Each trajectory directory includes run.json, a streamed events.jsonl event
log, sanitized session data, snapshots, and the shareable text/image files
produced during… See the full description on the dataset page: https://huggingface.co/datasets/schema-harness/arc-agi-3-schema-traces.swas-fits-spectra
SWAS FITS Spectra
The SWAS Tenth Public Release pointing-based IRAF-FITS Level 1.5 product serves 6,962 containers containing four or six named spectral image members.
Use
from datasets import load_dataset
ds = load_dataset("astro-legacy-archive/swas-fits-spectra", "14498-5856_0001__SWAS_IMEXT001_B1", split="train")
row = ds[0]
values = row[ds.column_names[0]]
print(len(values), values[:3])
725 [0.00926896184682846, 0.007523952983319759, 0.001602768199518323]… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/swas-fits-spectra.trellis500k-github-archives-9cbi-archive-raw
Central Bank of Ireland Archive: original source files
6,309 original files, 6.56 GB. Every PDF, spreadsheet, Word document and
archive gathered from the Central Bank of Ireland's public website, stored by
content hash so that a search result can be turned back into the document a
human would actually read.
This is the raw tier. If you want the text, you almost certainly want
aditya487/cbi-archive-corpus
instead: 5,568 documents and 89,242 page or pseudo-page rows as Parquet… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-raw.moltbook-observatory-archive
Observatory Dataset
This dataset is an incremental export of a SQLite observatory database, published as
date-partitioned Parquet files for efficient browsing and querying on Hugging Face.
For example, you can filter data by wildcards on date:
ds = load_dataset(
"SimulaMet/moltbook-observatory-archive",
"posts",
data_files="data/posts/2026-01-2*.parquet", # 20–29
split="train"
)
Each SQLite table is exposed as a separate dataset subset. Use dropdown above the… See the full description on the dataset page: https://huggingface.co/datasets/Hectorize/moltbook-observatory-archive.arctic
Arctic Shift Reddit Archive
Every Reddit comment and submission since 2005, organized as monthly Parquet shards
What is it?
The full Reddit archive from Arctic Shift, converted to Parquet and hosted here for easy access. Covers every public subreddit from 2005-12 through 2026-02.
Right now the archive has 1.6B items (362.1M comments, 1.2B submissions) in 181.4 GB of compressed Parquet. Comments and submissions are stored as separate datasets, split into monthly shards… See the full description on the dataset page: https://huggingface.co/datasets/Dk587/arctic.cbi-released-products
Cosmic Background Imager released products
This repository contains the numerical products released for four generations of Cosmic Background Imager (CBI) analysis: the 2000 deep fields, the 2000 mosaic fields, the final 2002–2005 temperature and polarization analysis, and the final 2000–2005 total-intensity analysis. It contains the published band powers, their window functions, the same-band correlation blocks and full Fisher matrices distributed in CosmoMC .newdat files, and… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/cbi-released-products.arca-importaciones-argentina
ARCA Importaciones Argentina
Dataset público derivado de la Información Agregada de Comercio Exterior publicada por ARCA Argentina.
Cobertura
El objetivo histórico abarca todos los meses publicados por ARCA desde 02/2017 hasta 08/2026.
Dos granularidades reales de ARCA
ARCA no mantuvo el mismo formato durante todo el período. El pipeline detecta el encabezado de cada ZIP y conserva la semántica correcta:
data/items/YYYYMM.parquet: meses con… See the full description on the dataset page: https://huggingface.co/datasets/alexbozz1/arca-importaciones-argentina.w2t-llm-arc-easy-lora
W2T Llm Arc Easy Lora
This repository contains artifacts for the W2T paper:
Paper: W2T: LoRA Weights Already Know What They Can Do
Repo: Weight2Token
Summary
ARC-Easy LoRA checkpoints and prepared metadata used for performance prediction.
Source Status
Storage location: local
Verification status: confirmed
Files
See manifest.json for the exact local or remote source paths used to prepare this release.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/Xiaolong-Han/w2t-llm-arc-easy-lora.arc_easytrellis500k-github-archives-5arc-agi3-kimi-k2.7-ar25
ARC-AGI-3 ar25 — Agent Trajectories (kimi-k2.7)
Gameplay trajectories from the harness×model pair kimi-k2.7 playing the
ARC-AGI-3 game ar25, part of the
ARA-as-world-model generalization experiment. The agent builds a structured world model
(an Agent-Native Research Artifact) live during play and consults it to crack levels it
cannot solve from cold exploration.
One dataset repo per harness×model×game: sibling repos
arc-agi3-<harness>-<model>-<game> hold the same game played by… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-kimi-k2.7-ar25.
