datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmu_sdss_sdss
mmu_sdss_sdss HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_sdss_sdss.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_sdss_sdss.GaSNet-II-SDSS-dataset
The SDSS spectra used in the paper: https://arxiv.org/abs/2311.04146
The code is available on Github: https://github.com/Fucheng-Zhong/GaSNet-II
If the dataset or code helps in your research, please cite paper 2311.04146
split_sdss_hsc_embeddingssdss---
description: 'Spectra dataset based on SDSS-IV.
'
homepage: https://www.sdss.org/
version: 1.0.0
citation: "% % ACKNOWLEDGEMENTS\n% % From: https://www.sdss4.org/collaboration/citing-sdss/\n\
% \n% Funding for the Sloan Digital Sky Survey IV has been provided by the Alfred \ P. Sloan Foundation, the U.S. Department of Energy Office of Science, and the \ Participating Institutions. SDSS acknowledges support and resources from the Center \ for High-Performance Computing at the… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/sdss.mmu-norm-sdss
SDSS spectra — L1 (release v1)
This L1 repository contains 271,966 spectra matched across the release, in 11 shards (about 15 GB).
Schema
spectrum struct per row: lambda (vacuum, heliocentric, observer frame, Å; variable length ≤ ~4,700), flux (×10⁻¹⁷ erg s⁻¹ cm⁻² Å⁻¹), ivar, lsf_sigma, mask (native; nonzero = bad — verified), valid. Plus object_id, positions, metadata.
Padding: the upstream arrays used lambda = −1 with zero inverse variance for padding. Those… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-sdss.mmu-l2-sdss
L2 model-ready views (release v1)
The mmu-l2-* repositories provide processed, model-ready versions of the matching mmu-norm-* L1 data, while L1 keeps the normalized source measurements. L2 applies documented processing steps for training and evaluation, with the details for reversing each transformation stored in the row or in provenance.json.
Repo
View
Reversal
mmu-l2-tess
per-sector relative flux f/median−1, time from first valid cadence
flux =… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-l2-sdss.mmu_sdss_sdss---
description: 'HATS version of MultimodalUniverse/sdss: Spectra dataset based on SDSS-IV.
'
homepage: https://www.sdss.org/
version: 1.0.0
citation: "% % ACKNOWLEDGEMENTS\n% % From: https://www.sdss4.org/collaboration/citing-sdss/\n\
% \n% Funding for the Sloan Digital Sky Survey IV has been provided by the Alfred \ P. Sloan Foundation, the U.S. Department of Energy Office of Science, and the \ Participating Institutions. SDSS acknowledges support and resources from the Center \ for… See the full description on the dataset page: https://huggingface.co/datasets/LSDB/mmu_sdss_sdss.IllustrisTNG_SKIRT_SDSS
IllustrisTNG SKIRT SDSS
Preprocessed synthetic galaxy images derived from the
IllustrisTNG cosmological simulations.
Raw multi-band FITS images were produced with the
SKIRT Monte Carlo radiative transfer code in SDSS
photometric bands and subsequently processed into 128 × 128 RGB images
ready for machine-learning applications.
The dataset is designed as training data for the
Spherinator /
HiPSter framework, but it is
general-purpose and suitable for any galaxy-morphology task.… See the full description on the dataset page: https://huggingface.co/datasets/HITS-AIN/IllustrisTNG_SKIRT_SDSS.hsc_sdss_cross_matchedSDSS-GalaxyZoo2-Preprocessed
Galaxy Zoo 2 - Preprocessed Dataset
Preprocessed galaxy images from SDSS with Galaxy Zoo 2 morphological classifications, ready for deep learning.
📁 Dataset Structure
├── labels.csv # All ~239K galaxy labels (101 MB)
├── labels_sampled.csv # 25K stratified sample (11 MB)
├── preprocessed_69x69.zip # Full dataset at 69×69 resolution (2.8 GB)
├── preprocessed_224x224.zip # Full dataset at 224×224 resolution (3.14 GB)
├──… See the full description on the dataset page: https://huggingface.co/datasets/marwa000000000/SDSS-GalaxyZoo2-Preprocessed.sdss_gaia_crossmatched
Crossmatched samples from the Multimodal Universe for SDSS/Gaia
Mother paper here: https://arxiv.org/abs/2412.02527
sdss-asteroid-taxonomy
SDSS-based Asteroid Taxonomy
Part of the Orbital Mechanics Datasets collection on Hugging Face.
Compositional taxonomy for 107,466 SDSS photometric observations of 63,468
asteroids, classified using the scheme of Carvano et al. (2010). Each observation includes
SDSS u'g'r'i'z' log-reflectances, a taxonomic class assignment, and a probability score.
Orbital elements from the asteroid catalog are merged in for asteroids with known orbits.
Dataset description
The Sloan… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/sdss-asteroid-taxonomy.desi_sdss_crossmatched
Crossmatched samples from the Multimodal Universe for DESI/SDSS
Mother paper here: https://arxiv.org/abs/2412.02527
sdss-nsa-metasdss_dr17_spectra_for_classification
SDSS DR17 Spectral Classification Parquet Benchmark
This fixed supervised benchmark classifies SDSS DR17 optical spectra as GALAXY, QSO, or STAR. Rows contain native FITS arrays, common-grid spectral inputs, identifiers, derived spectral-quality fields, and metadata matched from specObj-dr17.fits.
Files
train.parquet
validation.parquet
test.parquet
dataset_metadata.json
Split
Number of spectra
Train
22,581
Validation
4,839
Test
4,839
The… See the full description on the dataset page: https://huggingface.co/datasets/BrunoBarreto/sdss_dr17_spectra_for_classification.sdss_hsc_crossmatched
Crossmatched samples from the Multimodal Universe for SDSS/HSC
Mother paper here: https://arxiv.org/abs/2412.02527
sdssmmu-sdss-hscThis dataset is an adaptation and just for test purposes to implement MMU streaming.
Here a link to the original data sources:
hsc: https://huggingface.co/datasets/MultimodalUniverse/hsc
sdss: https://huggingface.co/datasets/MultimodalUniverse/sdss
HSC
description: 'Image dataset based on HSC SSP PRD3.
' homepage: https://hsc-release.mtk.nao.ac.jp/doc/ version: 1.0.0 citation: "% CITATION\n@article{Aihara_2017,\n title={The Hyper Suprime-Cam SSP
\ Survey: Overview and survey… See the full description on the dataset page: https://huggingface.co/datasets/TobiasPitters/mmu-sdss-hsc.deepfacelabmefile:///home/mzhh/%E6%A1%8C%E9%9D%A2/DFL_Me2.55/_internal/core
file:///home/mzhh/%E6%A1%8C%E9%9D%A2/DFL_Me2.55/_internal/DFLIMG
file:///home/mzhh/%E6%A1%8C%E9%9D%A2/DFL_Me2.55/_internal/doc
file:///home/mzhh/%E6%A1%8C%E9%9D%A2/DFL_Me2.55/_internal/facelib
file:///home/mzhh/%E6%A1%8C%E9%9D%A2/DFL_Me2.55/_internal/flaskr
file:///home/mzhh/%E6%A1%8C%E9%9D%A2/DFL_Me2.55/_internal/localization
file:///home/mzhh/%E6%A1%8C%E9%9D%A2/DFL_Me2.55/_internal/mainscripts… See the full description on the dataset page: https://huggingface.co/datasets/sdssfdf/deepfacelabme.sdss_dr12_stars_regression
SDSS DR12 Stellar Spectra Parquet Benchmark
This dataset provides a fixed supervised benchmark split for stellar atmospheric parameter estimation from SDSS DR12 optical spectra.
It contains three Parquet files:
train.parquet
validation.parquet
test.parquet
The split sizes are:
Split
Number of spectra
Train
29,422
Validation
4,846
Test
14,443
Each row corresponds to one SDSS stellar spectrum and includes raw spectral arrays, processed spectral features… See the full description on the dataset page: https://huggingface.co/datasets/BrunoBarreto/sdss_dr12_stars_regression.sds_section1_extraction2
Dataset Card for sds_section1_extraction2
This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Using this dataset with Argilla
To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code:
import argilla as rg
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Eshwar14/sds_section1_extraction2.mmu-sdss-with-coordinatessdss_hsc_embeddings
SDSS ↔ HSC Embeddings (The Platonic Universe)
Precomputed cross-survey embeddings for matched sources in SDSS (spectra) and HSC (images).Each row is one object with multiple HSC image-embedding vectors and one SDSS spectral-embedding vector.HSC columns include families like AstroPT, ConvNeXt, DINOv2, I-JEPA, and ViT (suffix _hsc); SDSS spectra use specformer_base_sdss.
Load in Python
from datasets import load_dataset
import numpy as np
ds =… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/sdss_hsc_embeddings.pretraindatah2f_sdss_pasquet_subsetmmu-sdss-partitionedAnnoy-PyEdu-Rs-Licensesgalaxy-classification-sdssSDS_Sim_Data_500_Rowssdss_hsc_crossmatched
Crossmatched samples from the Multimodal Universe for SDSS/HSC
Mother paper here: https://arxiv.org/abs/2412.02527
