datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AstroVLBench
AstroVLBench: A Benchmark for Vision-Language Models in Observational Astronomy
AstroVLBench is a comprehensive benchmark for evaluating vision-language models (VLMs) on authentic astronomical observation tasks. It comprises over 4,100 expert-verified evaluation instances spanning five observational modalities central to extragalactic and time-domain research: optical imaging, radio interferometric imaging, multi-wavelength photometry, time-domain light curves, and optical… See the full description on the dataset page: https://huggingface.co/datasets/anonymous4ai/AstroVLBench.astrobridge-yse-test-dataset-v2
AstroBridge YSE external test dataset v2
This dataset contains 266 object-disjoint, spectroscopically labeled YSE DR1 transients. The broad-class counts are SN II: 71, SN Ia: 180, SN Ibc: 15.
V2 shortens the YSE forced-photometry time coverage to resemble the alert-photometry coverage of the AstroBridge BTS training dataset. For each object, it retains the smallest inclusive time interval that contains every positive measurement with flux/uncertainty at least 5 and at least five… See the full description on the dataset page: https://huggingface.co/datasets/BuildNg/astrobridge-yse-test-dataset-v2.AstroAlertBench
AstroAlertBench
Dataset on the Hugging Face Hub: AnonymousUser16384/AstroAlertBench.
Vision–language benchmark on publicly available ZTF alerts brokered by ALeRCE. Each example pairs tabular metadata (ZTF-style candidate fields) with a single RGB stamp montage (science, reference, difference) as one PNG.
Dataset structure
Path
Description
data/manifest_benchmark_final.csv
Primary benchmark metadata — 1500 rows (300 per class: SN, AGN, VS, asteroid, bogus)… See the full description on the dataset page: https://huggingface.co/datasets/AnonymousUser16384/AstroAlertBench.mirabest-radio-astronomy-unofficial
MiraBest Radio Astronomy Dataset (Unofficial)
⚠️ IMPORTANT: This is an unofficial repository containing a processed version of the MiraBest dataset formatted for stable diffusion fine-tuning. This repository is not affiliated with the original authors.
Unofficial processing of the MiraBest radio astronomy dataset with original classification labels and natural language captions for diffusion fine-tuning. Original dataset by Porter & Scaife (2023).
Original Dataset
The… See the full description on the dataset page: https://huggingface.co/datasets/kwazzi-jack/mirabest-radio-astronomy-unofficial.astrobridge-yse-test-dataset
AstroBridge YSE external test dataset
This dataset contains 266 spectroscopically labeled YSE DR1 transients that do not overlap the final AstroBridge BTS training dataset by normalized TNS identity or a two-arcsecond transient-coordinate match. The broad-class counts are SN II: 71, SN Ia: 180, SN Ibc: 15.
Each row has at least five valid ZTF g and five valid ZTF r observations, ordered by time. The light-curve flux arrays are in nJy. The ATCAT arrays retain SNANA FLUXCAL at… See the full description on the dataset page: https://huggingface.co/datasets/BuildNg/astrobridge-yse-test-dataset.sentinel-lfm-mining-patches
sentinel-lfm — illegal-mining single-frame patches
128px RGB patches cropped from the Roboflow illegal-mining dataset, labelled
mine (1) / no-mine (0). Split by source image (no leakage) into
train/val/test. Provided as PNGs + vlm_sft-format JSONL (one image + prompt
-> JSON answer) so it drops straight into VLM fine-tuning.
split
pos
neg
total
train
1410
555
1965
val
303
66
369
test
303
116
419
RGB only (no multispectral). Each JSONL row is a single-turn VLM… See the full description on the dataset page: https://huggingface.co/datasets/ASTRALK/sentinel-lfm-mining-patches.AstroVLBench
AstroVLBench: A Benchmark for Vision-Language Models in Observational Astronomy
AstroVLBench is a comprehensive benchmark for evaluating vision-language models (VLMs) on authentic astronomical observation tasks. It comprises over 4,100 expert-verified evaluation instances spanning five observational modalities central to extragalactic and time-domain research: optical imaging, radio interferometric imaging, multi-wavelength photometry, time-domain light curves, and optical… See the full description on the dataset page: https://huggingface.co/datasets/XiaomanZhang/AstroVLBench.astro-iq
astro-iq — Astrophotography Sub-frame Quality Dataset
A machine-learning dataset for no-reference quality scoring of raw
astrophotography sub-exposures. Each sample is a 512×512 greyscale PNG
thumbnail (center-cropped from the original 6248×4176 sensor frame) paired
with photometric features and a continuous quality label in [0, 1].
Dataset summary
Property
Value
Total frames
8,378
Labelled frames
7,828
Sensors
ZWO ASI2600MC Pro (OSC), ZWO… See the full description on the dataset page: https://huggingface.co/datasets/ddompe/astro-iq.
