datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmu_hsc_pdr3_dud_22.5
mmu_hsc_pdr3_dud_22.5 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_hsc_pdr3_dud_22.5.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/mmu_hsc_pdr3_dud_22.5.TempleOS-Source-CodeHSCodeComp
HSCodeComp: A Realistic and Expert-Level Benchmark for Deep Search Agents in Hierarchical Rule Application
Paper | Code | Dataset on Hugging Face
⭐ MarcoPolo Team ⭐
Alibaba Group
🗂️ Data
📌 Overview
HSCodeComp is the first realistic, expert-level e-commerce benchmark designed to evaluate deep search agents on their ability to perform Level-3 knowledge—hierarchical rule application—a critical yet overlooked capability in current agent evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ATH-MaaS/HSCodeComp.legacysurvey_hsc_crossmatched
Crossmatched samples from the Multimodal Universe for Legacy Survey/HSC
Mother paper here: https://arxiv.org/abs/2412.02527
mmu_hsc_pdr3_wide_21
mmu_hsc_pdr3_wide_21 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_hsc_pdr3_wide_21.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB… See the full description on the dataset page: https://huggingface.co/datasets/hugging-science/mmu_hsc_pdr3_wide_21.split_legacysurvey_hsc_embeddingsTrain, calibration and test sets across models for Legacy Survey ↔ HSC embeddings (source: UniverseTBD/legacysurvey_hsc_embeddings).
mmu-norm-hsc
HSC PDR3 Deep/UltraDeep image cutouts — L1 (release v1)
This L1 repository contains 55,346 objects matched across the release, in 79 shards (about 57 GB). Each object has 160×160-pixel cutouts at 0.168″ per pixel in five bands (hsc-g/r/i/z/y), with per-pixel inverse variance and a mask.
Schema
image struct: band (5), flux (ADU at AB zeropoint 27), ivar, mask (true = CLEAN — verified empirically), psf_fwhm, scale. Plus cmodel magnitudes/errors, extendedness… See the full description on the dataset page: https://huggingface.co/datasets/kshitijd/mmu-norm-hsc.mmu-hsc-with-coordinatessplit_legacysurvey_hsc_embeddingsTrain, calibration and test sets across models for Legacy Survey ↔ HSC embeddings (source: UniverseTBD/legacysurvey_hsc_embeddings).
stocks-HSCL-1D-candlessplit_sdss_hsc_embeddingshsc---
description: 'Image dataset based on HSC SSP PRD3.
'
homepage: https://hsc-release.mtk.nao.ac.jp/doc/
version: 1.0.0
citation: "% CITATION\n@article{Aihara_2017,\n title={The Hyper Suprime-Cam SSP \ Survey: Overview and survey design},\n volume={70},\n ISSN={2053-051X},\n \ url={http://dx.doi.org/10.1093/pasj/psx066},\n DOI={10.1093/pasj/psx066},\n \ number={SP1},\n journal={Publications of the Astronomical Society of Japan},\n \ publisher={Oxford University Press… See the full description on the dataset page: https://huggingface.co/datasets/MultimodalUniverse/hsc.legacysurvey_hsc_embeddings
Legacy Survey ↔ HSC Embeddings (The Platonic Universe)
Precomputed cross-survey embeddings for matched sources in Legacy Survey and HSC.Each row is one object with multiple backbone embeddings for both surveys (paired by suffixes _legacysurvey and _hsc).
Examples of columns (see Viewer for full list):
AstroPT: astropt_15m_hsc, astropt_15m_legacysurvey, astropt_95m_*, astropt_850m_*
ConvNeXt: convnext_nano_*, convnext_tiny_*, convnext_base_*, convnext_large_*
DINOv2: dino_small_*… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/legacysurvey_hsc_embeddings.split_jwst_hsc_embeddingshsc-jwst-images-high-snrhsc-jwst-imagesmmu_hsc_pdr3_dud_22.5_minihscode
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/ohjimin/hscode.desi_hsc_crossmatched
Crossmatched samples from the Multimodal Universe for DESI/HSC
Mother paper here: https://arxiv.org/abs/2412.02527
HSCodeComp
HSCodeComp: A Realistic and Expert-Level Benchmark for Deep Search Agents in Hierarchical Rule Application
Paper | Code | Dataset on Hugging Face
⭐ MarcoPolo Team ⭐
Alibaba Group
🗂️ Data
📌 Overview
HSCodeComp is the first realistic, expert-level e-commerce benchmark designed to evaluate deep search agents on their ability to perform Level-3 knowledge—hierarchical rule application—a critical yet overlooked capability in current agent evaluation… See the full description on the dataset page: https://huggingface.co/datasets/chenjianhui0428/HSCodeComp.HSCodeComp
HSCodeComp: A Realistic and Expert-Level Benchmark for Deep Search Agents in Hierarchical Rule Application
Paper | Code | Dataset on Hugging Face
⭐ MarcoPolo Team ⭐
Alibaba Group
🗂️ Data
📌 Overview
HSCodeComp is the first realistic, expert-level e-commerce benchmark designed to evaluate deep search agents on their ability to perform Level-3 knowledge—hierarchical rule application—a critical yet overlooked capability in current agent evaluation… See the full description on the dataset page: https://huggingface.co/datasets/saadz506/HSCodeComp.hsc_pred_v_trueproduct_hscodehsc_sdss_cross_matchedhsc-wuxia-full
HSC Wuxia Full Dataset
This dataset is the complete master parallel corpus of Chinese-English sentence pairs developed for the Bachelor's Thesis (Trabajo de Fin de Grado - TFG) titled:
"Enfoques de traducción automática con modelos de lenguaje en obras wuxia"
(Degree in Data Science and Engineering, Universidade da Coruña)
It contains the full set of clean, aligned parallel sentence pairs extracted from the structural and semantic alignment of Wuxia novels.
Corpus… See the full description on the dataset page: https://huggingface.co/datasets/hsilvosa/hsc-wuxia-full.legacysurvey_hsc_crossmatched
Crossmatched samples from the Multimodal Universe for Legacy Survey/HSC
Mother paper here: https://arxiv.org/abs/2412.02527
hsc-pdr3-wide-20-embeddings
HSC PDR3 Wide r<20 Embeddings
AION-Search and AION embeddings for HSC PDR3 Wide galaxies with r_mag < 20 mag
License & data source
The embeddings and packaging in this repository are released under the MIT License.
The underlying catalog data are derived from the Hyper Suprime-Cam Subaru Strategic Program (HSC-SSP) and remain subject to the original HSC-SSP data-use policy and required acknowledgements.
Embeddings Citation
@misc{koblischke2025semantic… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/hsc-pdr3-wide-20-embeddings.HSC-GalaxiesML-VAE-embeddingshsc_flow_embeddingshsc-anomaly-expert
HSC Anomaly Expert
This dataset contains 1,000 galaxy images independently scored by 16 expert
astronomers. The ten fixed few-shot examples form the train split; the
remaining 990 images form the test split used for evaluation, containing
45 expert-consensus anomalies and 945 other images.
The images come from public data release 2 of the Subaru Hyper Suprime-Cam
(HSC) survey; this is the HSC in the dataset name. The source images,
Zooniverse subject metadata, volunteer… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/hsc-anomaly-expert.
