datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai2_arc
Dataset Card for "ai2_arc"
Dataset Summary
A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in
advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains
only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also
including a corpus of over 14 million science sentences… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ai2_arc.AI-CUDA-Engineer-Archive
The AI CUDA Engineer Archive 👷: Agentic CUDA Kernel Discovery, Optimization & Composition
We release The AI CUDA Engineer archive, a dataset consisting of approximately 30,000 CUDA kernels generated by The AI CUDA Engineer. It is released under the CC-By-4.0 license and can be accessed via HuggingFace and interactively visualized here. The dataset is based on the Kernel tasks provided in KernelBench and includes a torch reference implementation, torch, NCU and Clang-tidy… See the full description on the dataset page: https://huggingface.co/datasets/SakanaAI/AI-CUDA-Engineer-Archive.picbreeder-vlm-archive
Picbreeder-VLM Archive
Every image evolved by the swarm of vision-language-model "breeders" in
In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models
(GECCO 2026), together with the CPPN genomes that produced them, the agents' reasoning transcripts, the
lineage graphs, and the analysis artifacts behind the paper and blog.
The original Picbreeder (Secretan et al., 2008) let crowds of
humans collaboratively evolve images from
CPPN… See the full description on the dataset page: https://huggingface.co/datasets/picbreeder-vlm/picbreeder-vlm-archive.jfk-archives
Dataset Card for JFK Archives
This dataset is a collection of all records pertaining to the assassination of the
US president, John F. Kennedy, released until April 2025 through archives.org
by the US government.
Dataset Details
Dataset Description
The original data downloaded from archives.org
consists of 56,300 scanned documents in PDF format, released until April 2025. The files are
organized by their release year(s): 2107-2018, 2021, 2022, 2023 and 2025.… See the full description on the dataset page: https://huggingface.co/datasets/farhanhubble/jfk-archives.distilabel-intel-orca-dpo-pairs-binarizedThis is the binarized version of distilabel Orca Pairs for DPO and ORPO.
Reference: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs?row=0
arc-whestbench-public-2026
Organized by:
Alignment Research Center (ARC),
AIcrowd
WhestBench 2026: ARC White-Box Estimation Challenge
WhestBench is a benchmark for white-box activation estimation: given the weights of a randomly initialized ReLU multi-layer perceptron (MLP) and a strict floating-point-operation (FLOP) budget, predict the average post-activation value of every neuron when the network is fed standard Gaussian inputs.
This is the WhestBench 2026… See the full description on the dataset page: https://huggingface.co/datasets/aicrowd/arc-whestbench-public-2026.scheduleSee https://github.com/ust-archive/ust-archive for more information.
SCPWiki-Archive-02-March-2025-Datasetsfree-music-archive-full
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-full.planck-2018-chains
Planck 2018 Cosmological-Parameter Chains
This dataset contains the Planck Public Release 3 cosmological-parameter
full grid, COM_CosmoParams_fullGrid_R3.01.zip. It contains the Markov
chains and their GetDist and CosmoMC companions for 336 combinations of
cosmological model and likelihood or external-data selection. The 1,296
chain roots become 5,184 Parquet tables containing 27,699,519 rows.
The tables retain the source's headerless, positional structure. Arrow
fields are… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/planck-2018-chains.okapi_arc_challengemsmarco-v2.1-snowflake-arctic-embed-l
Snowflake Arctic Embed L Embeddings for MSMARCO V2.1 for TREC-RAG
This dataset contains the embeddings for the MSMARCO-V2.1 dataset which is used as the corpora for TREC RAG
All embeddings are created using Snowflake's Arctic Embed L and are intended to serve as a simple baseline for dense retrieval-based methods.
Retrieval Performance
Retrieval performance for the TREC DL21-23, MSMARCOV2-Dev and Raggy Queries can be found below with BM25 as a baseline. For both… See the full description on the dataset page: https://huggingface.co/datasets/Snowflake/msmarco-v2.1-snowflake-arctic-embed-l.wmap-single-year-maps
WMAP DR5 Single-Year I/Q/U Maps
The preview renders the canonical year-1 K1 TEMPERATURE field in its source
NESTED order, using a Galactic Mollweide projection and a symmetric 99.5th
percentile colour range.
This dataset contains the complete full-resolution single-year I/Q/U release
served by LAMBDA: ten WMAP differencing assemblies for each of nine observing
years. These are year-specific, per-assembly measurements. They are distinct
from wmap-band-maps-9yr, whose five maps… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/wmap-single-year-maps.gwosc-o2-strain
GWOSC O2 16 kHz gravitational-wave strain
This dataset contains the 16,384 Hz H1, L1, and V1 strain records released by
the Gravitational Wave Open Science Center for the
second observing run (O2). Each detector's 1 Hz data-quality and
hardware-injection masks are separate configurations, preserving the source
cadences.
Observing run
O2
GPS extent
1164558336–1187737600
Detectors
H1 (Hanford), L1 (Livingston), and V1 (Virgo)
Strain
16,384 Hz, float64
Masks… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gwosc-o2-strain.crimson-hexagonal-archive
The Crimson Hexagonal Archive — machine-readable representation
Query this without downloading anything. Every config is served by the Hugging Face datasets-server over plain HTTP, no auth, no client library. Use /rows — it is the reliable one. It reads the parquet directly and answers in under two seconds:
https://datasets-server.huggingface.co/rows?dataset=leesharks%2Fcrimson-hexagonal-archive&config=deposits&split=train&offset=0&length=10… See the full description on the dataset page: https://huggingface.co/datasets/leesharks/crimson-hexagonal-archive.gwosc-o3a-strain
GWOSC O3a 16 kHz strain
Upload in progress. Verified span shards are being added while source spans finish downloading and conversion.
This dataset contains the public Gravitational Wave Open Science Center O3a
strain release at 16,384 Hz for H1, L1, V1. Each detector is represented
independently. A contiguous source span produces three Parquet files:
Strain, DQmask, and Injmask.
Licence and acknowledgement
Creative Commons Attribution 4.0 International
Data… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gwosc-o3a-strain.ACVA-10percentArc-Corpusarc-voicesamples-generatedmoltbook-observatory-archive
Observatory Dataset
This dataset is an incremental export of a SQLite observatory database, published as
date-partitioned Parquet files for efficient browsing and querying on Hugging Face.
For example, you can filter data by wildcards on date:
ds = load_dataset(
"SimulaMet/moltbook-observatory-archive",
"posts",
data_files="data/posts/2026-01-2*.parquet", # 20–29
split="train"
)
Each SQLite table is exposed as a separate dataset subset. Use dropdown above the… See the full description on the dataset page: https://huggingface.co/datasets/SimulaMet/moltbook-observatory-archive.meteogate-archive
Meteogate European Weather Observations Archive
A continuously growing archive of real-time meteorological observations from the
EUMETNET Meteogate E-SOH service,
covering thousands of weather stations across Europe.
Data Structure
Each Parquet file contains observations in long format (one row per
station × variable × timestamp) with the following columns:
Column
Type
Description
timestamp
datetime
Observation time in UTC
station_id
string
WIGOS… See the full description on the dataset page: https://huggingface.co/datasets/alexdum/meteogate-archive.free-music-archive-medium
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-medium.arc-speeches-refinedarc-agigwosc-o1-strain
GWOSC O1 16 kHz gravitational-wave strain
This dataset contains the 16,384 Hz H1 and L1 strain records released by the
Gravitational Wave Open Science Center for Advanced
LIGO's first observing run (O1). Each detector's 1 Hz data-quality and
hardware-injection masks are included as separate configurations so that
measurements with different cadences remain separate.
Observing run
O1, GPS 1126051217–1137254417
Detectors
H1 (Hanford) and L1 (Livingston)
Strain
16… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gwosc-o1-strain.fermi-lat-weekly-photons
Fermi-LAT weekly photons
This dataset contains Fermi Large Area Telescope all-sky weekly photon files
from mission week w009 through w153, frozen on 2026-08-30. Its 145
configurations correspond one-to-one with the weekly p305_v001 FITS files.
Each Parquet row is an EVENTS row, with the 23 FITS-named columns in their
stored order and shape.
Mission weeks run Thursday through Wednesday in UTC. The first configuration
begins with the science-phase interval on 2008-08-04; w153… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/fermi-lat-weekly-photons.arc-multilingualgoogle-code-archive
Google Code Archive Dataset
Dataset Description
This dataset was compiled from the Google Code Archive, a preserved snapshot of projects hosted on Google Code, Google's open-source project hosting service that operated from 2006 to 2016. Google Code was one of the major code hosting platforms of its era, hosting hundreds of thousands of open-source projects before its shutdown. The archive provides a unique historical record of open-source development during a formative… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/google-code-archive.free-music-archive-small
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-small.cbi-released-products
Cosmic Background Imager released products
This repository contains the numerical products released for four generations of Cosmic Background Imager (CBI) analysis: the 2000 deep fields, the 2000 mosaic fields, the final 2002–2005 temperature and polarization analysis, and the final 2000–2005 total-intensity analysis. It contains the published band powers, their window functions, the same-band correlation blocks and full Fisher matrices distributed in CosmoMC .newdat files, and… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/cbi-released-products.
