datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AI-CUDA-Engineer-Archive
The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition
We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.scheduleSee https://github.com/ust-archive/ust-archive for more information.
rl-run-archive-2026
RL run archive 2026
Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer
checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for
long-term preservation and reproducibility.
Layout mirrors the verified backup trees they were copied from:
tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive
carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rl-run-archive-2026.planck-2018-chains
Planck 2018 Cosmological-Parameter Chains
This dataset contains the Planck Public Release 3 cosmological-parameter
full grid, COM_CosmoParams_fullGrid_R3.01.zip. It contains the Markov
chains and their GetDist and CosmoMC companions for 336 combinations of
cosmological model and likelihood or external-data selection. The 1,296
chain roots become 5,184 Parquet tables containing 27,699,519 rows.
The tables retain the source's headerless, positional structure. Arrow
fields are… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/planck-2018-chains.wmap-single-year-maps
WMAP DR5 Single-Year I/Q/U Maps
The preview renders the canonical year-1 K1 TEMPERATURE field in its source
NESTED order, using a Galactic Mollweide projection and a symmetric 99.5th
percentile colour range.
This dataset contains the complete full-resolution single-year I/Q/U release
served by LAMBDA: ten WMAP differencing assemblies for each of nine observing
years. These are year-specific, per-assembly measurements. They are distinct
from wmap-band-maps-9yr, whose five maps… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/wmap-single-year-maps.gwosc-o2-strain
GWOSC O2 16 kHz gravitational-wave strain
This dataset contains the 16,384 Hz H1, L1, and V1 strain records released by
the Gravitational Wave Open Science Center for the
second observing run (O2). Each detector's 1 Hz data-quality and
hardware-injection masks are separate configurations, preserving the source
cadences.
Observing run
O2
GPS extent
1164558336–1187737600
Detectors
H1 (Hanford), L1 (Livingston), and V1 (Virgo)
Strain
16,384 Hz, float64
Masks… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gwosc-o2-strain.crimson-hexagonal-archive
The Crimson Hexagonal Archive — machine-readable representation
Query this without downloading anything. Every config is served by the Hugging Face datasets-server over plain HTTP, no auth, no client library. Use /rows — it is the reliable one. It reads the parquet directly and answers in under two seconds:
https://datasets-server.huggingface.co/rows?dataset=leesharks%2Fcrimson-hexagonal-archive&config=deposits&split=train&offset=0&length=10… See the full description on the dataset page: https://huggingface.co/datasets/leesharks/crimson-hexagonal-archive.trellis500k-sketchfab-archivesmoltbook-observatory-archive
Observatory Dataset
This dataset is an incremental export of a SQLite observatory database, published as
date-partitioned Parquet files for efficient browsing and querying on Hugging Face.
For example, you can filter data by wildcards on date:
ds = load_dataset(
"SimulaMet/moltbook-observatory-archive",
"posts",
data_files="data/posts/2026-01-2*.parquet", # 20–29
split="train"
)
Each SQLite table is exposed as a separate dataset subset. Use dropdown above the… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/moltbook-observatory-archive.gaia-dr3-source
Gaia DR3 Source
This dataset mirrors the complete ESA Gaia Data Release 3 gaia_source
bulk-download table. It contains one record for every published Gaia source and
all 152 columns served by ESA, including source identifiers, astrometry,
photometry, observing statistics, quality fields, classifications, and
astrophysical parameters.
Each of ESA's 3,386 compressed ECSV shards is retained as one dataset
configuration. The configuration name is the source filename stem verbatim.… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gaia-dr3-source.meteogate-archive
Meteogate European Weather Observations Archive
A continuously growing archive of real-time meteorological observations from the
EUMETNET Meteogate E-SOH service,
covering thousands of weather stations across Europe.
Data Structure
Each Parquet file contains observations in long format (one row per
station × variable × timestamp) with the following columns:
Column
Type
Description
timestamp
datetime
Observation time in UTC
station_id
string
WIGOS… See the full description on the dataset page: https://huggingface.co/datasets/alexdum/meteogate-archive.gwosc-o1-strain
GWOSC O1 16 kHz gravitational-wave strain
This dataset contains the 16,384 Hz H1 and L1 strain records released by the
Gravitational Wave Open Science Center for Advanced
LIGO's first observing run (O1). Each detector's 1 Hz data-quality and
hardware-injection masks are included as separate configurations so that
measurements with different cadences remain separate.
Observing run
O1, GPS 1126051217–1137254417
Detectors
H1 (Hanford) and L1 (Livingston)
Strain
16… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gwosc-o1-strain.fermi-lat-weekly-photons
Fermi-LAT weekly photons
This dataset contains Fermi Large Area Telescope all-sky weekly photon files
from mission week w009 through w153, frozen on 2026-08-30. Its 145
configurations correspond one-to-one with the weekly p305_v001 FITS files.
Each Parquet row is an EVENTS row, with the 23 FITS-named columns in their
stored order and shape.
Mission weeks run Thursday through Wednesday in UTC. The first configuration
begins with the science-phase interval on 2008-08-04; w153… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/fermi-lat-weekly-photons.archive-dolma3-6t-dedup-state
archive-dolma3-6t-dedup-state
ARCHIVE: backup of the SOC-90 dedup state (Bloom filter + filtered docs). Reproducibility-only; current dedup state is in HCAI-Lab/dolma3-6t-bloom-index and dolma3-6t-unique.
Provenance
This dataset was renamed on 2026-05-25 as part of the HCAI-Lab HF naming convention cleanup (PR 3). See docs/HCAI_LAB_NAMING_CONVENTION.md in the project repo for the convention.
Field
Value
Previous name
HCAI-Lab/soc-90-keep-backup… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/archive-dolma3-6t-dedup-state.trellis500k-github-archives-10glenans-isobars-archiveswas-fits-spectra
SWAS FITS Spectra
The SWAS Tenth Public Release pointing-based IRAF-FITS Level 1.5 product serves 6,962 containers containing four or six named spectral image members.
Use
from datasets import load_dataset
ds = load_dataset("astro-legacy-archive/swas-fits-spectra", "14498-5856_0001__SWAS_IMEXT001_B1", split="train")
row = ds[0]
values = row[ds.column_names[0]]
print(len(values), values[:3])
725 [0.00926896184682846, 0.007523952983319759, 0.001602768199518323]… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/swas-fits-spectra.trellis500k-github-archives-9cbi-archive-raw
Central Bank of Ireland Archive: original source files
6,309 original files, 6.56 GB. Every PDF, spreadsheet, Word document and
archive gathered from the Central Bank of Ireland's public website, stored by
content hash so that a search result can be turned back into the document a
human would actually read.
This is the raw tier. If you want the text, you almost certainly want
aditya487/cbi-archive-corpus
instead: 5,568 documents and 89,242 page or pseudo-page rows as Parquet… See the full description on the dataset page: https://huggingface.co/datasets/aditya487/cbi-archive-raw.moltbook-observatory-archive
Observatory Dataset
This dataset is an incremental export of a SQLite observatory database, published as
date-partitioned Parquet files for efficient browsing and querying on Hugging Face.
For example, you can filter data by wildcards on date:
ds = load_dataset(
"SimulaMet/moltbook-observatory-archive",
"posts",
data_files="data/posts/2026-01-2*.parquet", # 20–29
split="train"
)
Each SQLite table is exposed as a separate dataset subset. Use dropdown above the… See the full description on the dataset page: https://huggingface.co/datasets/Hectorize/moltbook-observatory-archive.cbi-released-products
Cosmic Background Imager released products
This repository contains the numerical products released for four generations of Cosmic Background Imager (CBI) analysis: the 2000 deep fields, the 2000 mosaic fields, the final 2002–2005 temperature and polarization analysis, and the final 2000–2005 total-intensity analysis. It contains the published band powers, their window functions, the same-band correlation blocks and full Fisher matrices distributed in CosmoMC .newdat files, and… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/cbi-released-products.trellis500k-github-archives-5juno-microwave-maps
Juno microwave sky maps and map-space companions
This dataset contains the 2025 LAMBDA release of Juno Microwave Radiometer sky
maps and correlated-noise companions. Each of the 48 FITS bintables is one
Parquet configuration named by its source filename stem. Its column is T,
with the source table shape, float64 or int64 dtype, and row order. The six HDF5 companions add 20 dense-matrix
configurations named by the source filename stem, __, and the HDF5 leaf name.
Each leaf name… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/juno-microwave-maps.ust-rankings
UST Rankings
Daily course and instructor rating marts for UST Rankings, built from the
ust-archive datasets.
File
Contents
courses.parquet
Current Course metadata by Course Code.
course-ratings.parquet
Longitudinal course ratings by term and criterion.
instructor-ratings.parquet
Longitudinal instructor ratings by term and criterion.
course-rankings.parquet
Latest-term course ratings.
instructor-rankings.parquet
Latest-term instructor ratings.… See the full description on the dataset page: https://huggingface.co/datasets/ust-archive/ust-rankings.planck-frequency-maps-pr4
Planck PR4 NPIPE Frequency Maps
This dataset contains the nine Planck Public Release 4 frequency maps
produced by the NPIPE joint LFI/HFI processing pipeline. It serves the
full-mission, full-channel source tables exactly as released: source
column names, units, fixed-vector shapes, row order, dtypes, and value
bits are preserved. Metadata makes the source representation
self-describing without flattening, renaming, or normalising it.
The included maps are the R4.00 products at… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/planck-frequency-maps-pr4.trellis500k-github-archives-4msam-released-products
MSAM released flight products
This dataset contains the complete numeric contents of the three MSAM1 flight
archives served by LAMBDA: the June 1992, June 1994, and June 1995 observation
tables, sampled beam maps, and released 1994/1995 covariance matrices. The 38
configuration identifiers preserve the flight directory and source filename
stem.
How to use
python -m venv .venv && .venv/bin/pip install datasets huggingface_hub pyarrow
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/msam-released-products.cosb-photon-events
COS-B photon events
python -m venv .venv && .venv/bin/pip install datasets huggingface_hub pyarrow
This dataset contains the archive-natural event products for all 65 COS-B
pointings served by HEASARC. Each configuration is one source event file;
pointing 37, targeted on M31, is used in the examples below.
Cite H. A. Meyer-Hasselwander et al. (1986), Explanatory Supplement to the
COS-B Final Database, Proc. Cosmic Ray Conf., La Jolla, ESA Publication.
No explicit dataset… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/cosb-photon-events.trellis500k-github-archives-8us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites
Historical US layoffs archive: 6,799 WARN Act notices that state websites no longer list (2000-2025), recovered
Rebuilt 2026-09-22. Five state labor agencies — Connecticut, Michigan, New York,
North Carolina and Pennsylvania — retired the web pages their older WARN Act
layoff notices lived on. Their current pages start years later. This dataset is
every notice in our file that came from one of those retired pages and is not
on the agency's live page today: 6,799 notices, 6,694… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-historical-layoffs-archive-warn-act-notices-removed-from-state-websites.
