datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yonder
Yonder: A 4.65M-Frame Drone-Perspective Dataset for Indoor Navigation
The cross-simulator generalization gap.
Yonder is the largest publicly available drone-perspective dataset for indoor
navigation, plus a closed-loop benchmark designed to expose a failure mode invisible
to standard offline metrics: perception trained on one simulator does not transfer
cleanly to a different simulator, even when both target the same task.
This dataset accompanies the NeurIPS 2026 Datasets &… See the full description on the dataset page: https://huggingface.co/datasets/astralhf/yonder.astrovision-data
About
AstroVision is a first-of-a-kind, large-scale dataset of real small body images from both legacy and ongoing deep space missions, which currently features 115,970 densely annotated, real images of sixteen small bodies from eight missions. AstroVision was developed to facilitate the study of computer vision and deep learning for autonomous navigation in the vicinity of a small body, with speicial emphasis on training and evaluation of deep learning-based keypoint detection… See the full description on the dataset page: https://huggingface.co/datasets/travisdriver/astrovision-data.ocaml-opam-ppxlib-json-astversion https://git-lfs.github.com/spec/v1
oid sha256:b797e216eb8d6720d794ca5d07a01019cac2d79df456fbcc69b6407497cc267c
size 473
Ezaris-Training-Sets
Ezaris-Training-Sets
Complete training data, checkpoints, code and provenance for the Ezaris-1B program (formerly LUNA-1B) — a 1.2B-parameter Llama-style model trained from scratch by ASTERIZER.
Layout
Path
Contents
pretrain_raw_datasets/
Raw pretraining sources (SmolLM-corpus, fineweb-edu, OpenWebMath, Finemath, Wiki, code)
pretrain_cleaned_datasets/
Cleaned / deduplicated / tokenized pretrain sets + phase-2 exact & near-dedup outputs… See the full description on the dataset page: https://huggingface.co/datasets/ASTERIZER/Ezaris-Training-Sets.minty-astro-ph
MINT-1T ArXiv Astro-ph
An astronomy-focused subset of mlfoundations/MINT-1T-ArXiv, filtered to include only papers from the astro-ph arXiv category (including cross-listed papers).
Overview
Papers
~845k
Total size
~804 GB
Format
WebDataset tar shards
Shards
287 (astro-ph-00000.tar to astro-ph-00286.tar)
Shard size
~3 GB each
Source
MINT-1T (Awadalla et al., 2024)
Data Format
Each tar shard contains paired files per paper:… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/minty-astro-ph.AstroM3Dataset
AstroM3Dataset
Description
AstroM3Dataset is a time-series astronomy dataset containing photometry, spectra, and metadata features for variable stars.
The dataset was constructed by cross-matching publicly available astronomical datasets,
primarily from the ASAS-SN (Shappee et al. 2014) variable star catalog (Jayasinghe et al. 2019)
and LAMOST spectroscopic survey (Cui et al. 2012), along with data from
WISE (Wright et al. 2010), GALEX (Morrissey et al. 2007), 2MASS… See the full description on the dataset page: https://huggingface.co/datasets/AstroMLCore/AstroM3Dataset.scarlet-test-datasec-10k-markdown-uncompressed
📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents)
Dataset Summary
This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025).
The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.planck-2018-chains
Planck 2018 Cosmological-Parameter Chains
This dataset contains the Planck Public Release 3 cosmological-parameter
full grid, COM_CosmoParams_fullGrid_R3.01.zip. It contains the Markov
chains and their GetDist and CosmoMC companions for 336 combinations of
cosmological model and likelihood or external-data selection. The 1,296
chain roots become 5,184 Parquet tables containing 27,699,519 rows.
The tables retain the source's headerless, positional structure. Arrow
fields are… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/planck-2018-chains.generated-csvsSBI-16-3D
SBI-16-3D Dataset
SBI-16-3D is a dataset which is part of the AstroCompress project. It contains data assembled from the James Webb Space Telescope (JWST). Note that the underlying data is released under CC-4.0.
Usage
You first need to install the datasets and astropy packages:
pip install datasets astropy
There are two datasets: tiny and full, each with train and test splits. The tiny dataset has 2 4D images in the train and 1 in the test. The full dataset contains all… See the full description on the dataset page: https://huggingface.co/datasets/AstroCompress/SBI-16-3D.wmap-single-year-maps
WMAP DR5 Single-Year I/Q/U Maps
The preview renders the canonical year-1 K1 TEMPERATURE field in its source
NESTED order, using a Galactic Mollweide projection and a symmetric 99.5th
percentile colour range.
This dataset contains the complete full-resolution single-year I/Q/U release
served by LAMBDA: ten WMAP differencing assemblies for each of nine observing
years. These are year-specific, per-assembly measurements. They are distinct
from wmap-band-maps-9yr, whose five maps… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/wmap-single-year-maps.gwosc-o2-strain
GWOSC O2 16 kHz gravitational-wave strain
This dataset contains the 16,384 Hz H1, L1, and V1 strain records released by
the Gravitational Wave Open Science Center for the
second observing run (O2). Each detector's 1 Hz data-quality and
hardware-injection masks are separate configurations, preserving the source
cadences.
Observing run
O2
GPS extent
1164558336–1187737600
Detectors
H1 (Hanford), L1 (Livingston), and V1 (Virgo)
Strain
16,384 Hz, float64
Masks… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gwosc-o2-strain.BigAudioDataset
AstraMindAI/BigAudioDataset
Dataset Description
AstraMindAI/BigAudioDataset is a large-scale, multilingual dataset designed for a wide range of audio and speech processing tasks. It comprises a diverse collection of audio clips, including both spoken voice and music, making it a valuable resource for training and evaluating models for automatic speech recognition (ASR), text-to-speech (TTS), audio classification, and more.
The voice data is aggregated from well-known… See the full description on the dataset page: https://huggingface.co/datasets/AstraMindAI/BigAudioDataset.gwosc-o3a-strain
GWOSC O3a 16 kHz strain
Upload in progress. Verified span shards are being added while source spans finish downloading and conversion.
This dataset contains the public Gravitational Wave Open Science Center O3a
strain release at 16,384 Hz for H1, L1, V1. Each detector is represented
independently. A contiguous source span produces three Parquet files:
Strain, DQmask, and Injmask.
Licence and acknowledgement
Creative Commons Attribution 4.0 International
Data… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gwosc-o3a-strain.gaia-dr3-source
Gaia DR3 Source
This dataset mirrors the complete ESA Gaia Data Release 3 gaia_source
bulk-download table. It contains one record for every published Gaia source and
all 152 columns served by ESA, including source identifiers, astrometry,
photometry, observing statistics, quality fields, classifications, and
astrophysical parameters.
Each of ESA's 3,386 compressed ECSV shards is retained as one dataset
configuration. The configuration name is the source filename stem verbatim.… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gaia-dr3-source.clouds-decoded-rain-check
Sentinel-2 Cloud Property Retrievals over the UK and India
Rendered per-scene quicklook layers of cloud optical and microphysical properties
retrieved from Sentinel-2 (MSI) Level-1C imagery over eight tiles (four in the UK
and four in India). These images power the interactive report at
asterisk-labs-clouds-decoded-rain-check.static.hf.space. Cloud properties are estimated with the clouds-decoded retrieval code, developed as part of the clouds decoded project funded by ARIA.… See the full description on the dataset page: https://huggingface.co/datasets/asterisk-labs/clouds-decoded-rain-check.CCM_datagwosc-o1-strain
GWOSC O1 16 kHz gravitational-wave strain
This dataset contains the 16,384 Hz H1 and L1 strain records released by the
Gravitational Wave Open Science Center for Advanced
LIGO's first observing run (O1). Each detector's 1 Hz data-quality and
hardware-injection masks are included as separate configurations so that
measurements with different cadences remain separate.
Observing run
O1, GPS 1126051217–1137254417
Detectors
H1 (Hanford) and L1 (Livingston)
Strain
16… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/gwosc-o1-strain.fermi-lat-weekly-photons
Fermi-LAT weekly photons
This dataset contains Fermi Large Area Telescope all-sky weekly photon files
from mission week w009 through w153, frozen on 2026-08-30. Its 145
configurations correspond one-to-one with the weekly p305_v001 FITS files.
Each Parquet row is an EVENTS row, with the 23 FITS-named columns in their
stored order and shape.
Mission weeks run Thursday through Wednesday in UTC. The first configuration
begins with the science-phase interval on 2008-08-04; w153… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/fermi-lat-weekly-photons.Music-POSTPROCESS-509ab05ecacheudio-POSTPROCESS-3fd79cfbPluto-Nano-1.0-Pretrain-v2
ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2)
Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI).
v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.asta-benchMusic-POSTPROCESS-32eadf7eLAVIBData for LAVIB: A Large-scale Video Interpolation Benchmark (arxiv link: arxiv.org/abs/2406.09754)
SBI-16-2DSBI-16-2D is a dataset which is part of the AstroCompress project.
It contains imaging data assembled from the Hubble Space Telescope (HST).AstroDimbfcl-v1-non-live-ast-hermes
