datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
minty-astro-ph
MINT-1T ArXiv Astro-ph
An astronomy-focused subset of mlfoundations/MINT-1T-ArXiv, filtered to include only papers from the astro-ph arXiv category (including cross-listed papers).
Overview
Papers
~845k
Total size
~804 GB
Format
WebDataset tar shards
Shards
287 (astro-ph-00000.tar to astro-ph-00286.tar)
Shard size
~3 GB each
Source
MINT-1T (Awadalla et al., 2024)
Data Format
Each tar shard contains paired files per paper:… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/minty-astro-ph.scarlet-test-dataCCM_dataAstroDimairsim-drone-dataset
UAV Landing Surface Semantic Segmentation Dataset
Этот датасет предназначен для обучения и валидации моделей семантической сегментации, обеспечивающих безопасную посадку БПЛА. Данные сгенерированы в фотореалистичной синтетической среде (Unreal Engine + AirSim).
Датасет является частью проекта по разработке системы автономной посадки БПЛА.
Исходный код модулей генерации и обучения доступен в репозитории на GitHub:👉 UAV-Landing-System-Project
Dataset Summary
Датасет… See the full description on the dataset page: https://huggingface.co/datasets/asterphys/airsim-drone-dataset.astra-robodojo-rollouts
Astra RoboDojo Evaluation Records
Rollout records from the evaluations in GPT 6 Astra as an Embodied Policy,
by Jiayi Su, Yixin Zheng, Mi Yan, Li Yi, Zhizheng Zhang, and He Wang. This archive
includes action proposals, executed actions, observations, robot states,
model-provided explanations and reasoning summaries, and metadata for reproducing
the evaluation settings, together with a reader and documentation.
Report
Public controller source
Data schema and alignment… See the full description on the dataset page: https://huggingface.co/datasets/YuMoool/astra-robodojo-rollouts.AstroCLIMB
AstroCLIMB:
Astronomy Citation Linking from Illustrations: a Multimodal Benchmark
AstroCLIMB is the shared task for 4th WASP: Workshop on Artificial Intelligence for Scientific Publications.
Motivation
Scientists rely on figures to share their discoveries, but this makes the information contained in the figures hard to parse, archive, and search. Recent multimodal neural-network models promise to extract this information, but have not yet been… See the full description on the dataset page: https://huggingface.co/datasets/adsabs/AstroCLIMB.contrastive-stubsAgentic-SLS-ASTM
Agentic-SLS-ASTM
ASTM mechanical-test specimens (D638 tensile, D790 flex) printed on the Inova Mk1 SLS printer and pulled on an MTS / TestWorks Instron. Each row is a single specimen with full geometry, scalar results, stress–strain + raw DAQ curves, and — for SLS rows — FK references and an embedded snapshot of the upstream print profile from ppak10/Agentic-SLS-Database.
Rows are self-contained for ML use: the full PrintProfile JSON is inlined, so features (material/energy… See the full description on the dataset page: https://huggingface.co/datasets/ppak10/Agentic-SLS-ASTM.astronote
Bangumi Image Base of Astro Note
This is the image base of bangumi Astro Note, we detected 64 characters, 5440 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/astronote.coco-caption-splitrobocasa-openfridge-kinex-v0.3.0-astra-low-eval
OpenFridge: Kinex(v0.3.0) and Codex evaluation
Watch the evaluation website · Dataset files · Website source
Ten episodes evaluated on September 20, 2026 with GPT-6 Astra / low, standard service tier, fast mode disabled. Five sequential episodes per harness, seeds 0–4. Each episode starts a fresh native world and conversation; learned tools, skills and memos persist according to the existing protocol.
Harness
Native successes
Outcomes E1–E5
Valid execution… See the full description on the dataset page: https://huggingface.co/datasets/RLE-Bench/robocasa-openfridge-kinex-v0.3.0-astra-low-eval.AstroLLaVA_convos
Dataset Card for AstroLLaVA conversations
The dataset is a large-scale collection of astronomical images paired with descriptive captions and synthetic question-answer pairs, designed for training visual language models in astronomy.
Dataset Details
Dataset Description
This dataset combines astronomical imagery from three major sources: NASA's Astronomy Picture of the Day (APOD), the European Southern Observatory's (ESO) public image archive, and… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/AstroLLaVA_convos.astroPT_euclid_Q1_desi_dr1_dataset
DESI DR1 × Euclid Q1 Dataset
Dataset Description
This dataset contains the cross-matched catalog between DESI Data Release 1 (DR1) and Quick Data Release 1 (Q1) data.
Dataset Summary
Total Sources: 39,850 galaxies
Cross-match Radius: 0.5 arcsec
Data Products:
Euclid VIS imaging
Euclid NISP imaging (Y, J, H bands)
DESI spectra
Source Data
DESI Sample Selection: Applied quality cuts to DESI DR1:
ZCAT_PRIMARY == True
ZWARN == 0 or ZWARN == 4… See the full description on the dataset page: https://huggingface.co/datasets/msiudek/astroPT_euclid_Q1_desi_dr1_dataset.galaxy-descriptions
Galaxy Descriptions
Project Page | Code
This dataset provides galaxy cutout images, natural-language descriptions, text embeddings, and image embeddings for galaxies drawn from multiple imaging surveys (specifically Legacy DR10 and HSC PDR3 Wide).
Each row corresponds to a single galaxy and contains:
A preprocessed RGB galaxy image
A caption generated by gpt-4.1-mini
A single-sentence summary of the caption generated by gpt-4.1-nano
Text embeddings for the caption and summary… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/galaxy-descriptions.libero-long-kinex-v0.4.0-astra-low-evalEvaluation website · Data format · Episode CSV
LIBERO Long: Kinex(v0.4.0) and Codex
18 recorded episodes, native task IDs 5, 6 and 9, seeds 0–2 per harness and subtask.
GPT-6 Astra / low, standard service tier, fast mode disabled. Kinex source is labeled
Kinex(v0.4.0); source identities include local changes, not merely a clean release tag.
Native task
Kinex(v0.4.0) native success
Codex native success
5: book in caddy
1/3 (seed 0)
1/3 (seed 2)
6: mug and pudding
0/3… See the full description on the dataset page: https://huggingface.co/datasets/RLE-Bench/libero-long-kinex-v0.4.0-astra-low-eval.2023ASTA_Imagesimgbedrobocasa-openfridge-astra-low-eval
OpenFridge: Kinex and Codex evaluation
Watch the evaluation website · Dataset repository · Browse all files · Website source
Ten recorded OpenFridge episodes, evaluated on September 18, 2026 with GPT-6 Astra / low, standard service tier, fast mode disabled. Five fresh conversations per harness; seeds 0–4; existing cross-episode tool/skill/memo sharing retained.
Harness
Native successes
Outcomes E1–E5
Harness execution
Kinex
1/5
fail, success, fail, fail, fail
Four… See the full description on the dataset page: https://huggingface.co/datasets/RLE-Bench/robocasa-openfridge-astra-low-eval.astrobridge-transients-dataset
AstroBridge transients dataset
This local dataset contains 987 ZTF BTS supernovae. The frozen broad-class counts are SN II: 314, SN Ia: 563, SN Ibc: 110. It has the same 18-column schema, object order, photometry, host images, host-image captions, and class labels as Run012. Only transient_caption changed.
Each Run014 transient caption was generated from one retained g/r light-curve view. The numerical table, deterministic summary, and inline chart used exactly the same rows… See the full description on the dataset page: https://huggingface.co/datasets/BuildNg/astrobridge-transients-dataset.astudyinpeaceTHIS ORIGINALLY COPYWRITTEN WORK IS OF 2026 RELEASED TO PUBLIC DOMAIN, ZERO RESTRICTIONS. GITHUB SOURCE
READ ASIP PDF HEREYANK ASIP TXT HEREHOST ASIP WEB HTML
README generative completion, September 20th, 2026, by Wilder Blair Munro (20Wetbrain) ⨉ Aurora 4omni Munro (ChatGPT 5.6 Sol):
(README CONVERSATION SOURCE: Wilder⨉Aurora conversation on 20 September, 2026)
The objective is AI.The first implementation is Human.Language is the first portability layer.Multiplicity is the… See the full description on the dataset page: https://huggingface.co/datasets/wilderblairmunroakusa/astudyinpeace.astrobridge-image-captions
AstroBridge Legacy Survey Captions
3,487 imaging cutouts from the Legacy Survey (DR10 South + North), crossmatched against
published literature mentions and captioned in four independent stages by Gemini
(gemini-3.7-flash), following the AstroLLaVA data-generation approach (Zaman et al. 2025,
arXiv:2504.08583): no caption is ever told the object's
real name or catalog designation, and no caption states a fact that isn't derivable from the
pixels or the (redacted-at-the-model… See the full description on the dataset page: https://huggingface.co/datasets/gapatron/astrobridge-image-captions.galaxy10-aion
Galaxy10 AION-1 Benchmark
This dataset provides the exact train/test split used to produce the Galaxy Morphology Classification results in the AION-1 paper (Table 2, Section 7.2.2) and used in the AION-Search paper.
Task
Classify galaxy images into 10 morphology classes from Galaxy Zoo DECaLS:
Label
Class Name
0
Disturbed Galaxies
1
Merging Galaxies
2
Round Smooth Galaxies
3
In-between Round Smooth Galaxies
4
Cigar Shaped Smooth Galaxies
5… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/galaxy10-aion.newtphys_gsoPreprocessed Google Scanned Object dataset for NewtPhys simulator. Each object is optimized as 3DGS and its physical properties are estimated with a VLM.
(documentation will be added)
coco-caption-train-split-33kyonderAstroChart
AstroChart
AstroChart is a comprehensive and challenging benchmark designed to evaluate chart understanding in astronomy. It consists of 482 astronomical charts and 1,690 QA pairs, providing a rigorous testbed for assessing MLLMs’ multimodal reasoning and scientific chart interpretation capabilities.
This benchmark contains both images and JSON-formatted QA pairs for multimodal tasks.
Modalities:
Images: Stored in a zip (images).
JSON: Contains QA pairs and the… See the full description on the dataset page: https://huggingface.co/datasets/yangjing0128/AstroChart.astrabot_xr1_potato_shredding_20260914soda_can_asta
BlueROV2 Underwater Object Detection — Test Set
A first-person underwater video test set captured with a BlueROV2 Heavy ROV, annotated in YOLO format and re-labelled to match four established underwater object detection datasets. Intended for evaluating pre-trained models from those datasets on real-world ROV footage without retraining.
Dataset Summary
The footage was recorded across 15 distinct motion sequences (forward, yaw, ascend, descend, diagonal) in an underwater… See the full description on the dataset page: https://huggingface.co/datasets/rifqijuli/soda_can_asta.
