datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
galaxy-mentions-hats
Galaxy Mentions HATS
A HATS catalog of 43,546 resolved literature mentions from
astronolan/galaxy-mentions,
prepared for efficient spatial crossmatching with the
Multimodal Universe HATS catalogs.
Each row is a mention in a paper, not a deduplicated astronomical object. Multiple rows may
therefore describe the same galaxy. The optional wiki_entity_id groups mentions using the current
1 arcsecond connected-components build, while mention_id remains the unique row identifier.… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/galaxy-mentions-hats.galaxies-with-hats
Galaxies with HATS
This is a HATS (HEALPix Adaptive Tiling Scheme)
version of Smith42/galaxies
(revision v2.0): ~8.5 million 256×256 pixel PNG galaxy cutouts from the
DESI Legacy Survey DR8, centred on
the galaxy source, spatially partitioned on the sky, and bundled with all ~160
metadata columns (Galaxy Zoo DESI morphologies, NSA photometry, redshifts, and
OSSY / ALFALFA / JHU-MPA cross-matches) in every row.
Where the original dataset is ordered for ML training, this… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/galaxies-with-hats.galaxy_descriptions_hats
Galaxy Descriptions (HATS)
Original dataset: Nolan Koblischke's astronolan/galaxy-descriptions
This is the same dataset, repartitioned into the spatially indexed HATS format. The scientific content is unchanged.
This catalog contains the original 275,613 galaxy rows, including images, captions, summaries, text embeddings, AION image embeddings, coordinates, survey names, and object identifiers. Conversion added the HATS _healpix_29 spatial index and organized the rows into… See the full description on the dataset page: https://huggingface.co/datasets/Smith42/galaxy_descriptions_hats.hatsunebijouTeVCat_HATS
TeVCat in HATS format
A snapshot of TeVCat, the online catalog of TeV gamma-ray sources, converted to the HATS format so it can be queried and crossmatched with LSDB alongside other HATS catalogs, such as those from the Multimodal Universe.
Snapshot
2026-09-17
Sources
363
Size
~0.4 MB
Build code and validation
github.com/youyou-astro/TeVCat_HATS
TeVCat is maintained by its own team; this dataset only redistributes a formatted copy. See Citation and terms… See the full description on the dataset page: https://huggingface.co/datasets/youyouli-astro/TeVCat_HATS.resumes
Dataset Card for Advanced Resume Parser & Job Matcher Resumes
This dataset contains a merged collection of real and synthetic resume data in JSON format. The resumes have been normalized to a common schema to facilitate the development of NLP models for candidate-job matching in the technical recruitment domain.
Dataset Details
Dataset Description
This dataset is a combined collection of real resumes and synthetically generated CVs.
Curated by: datasetmaster… See the full description on the dataset page: https://huggingface.co/datasets/Hatshe/resumes.HATS-fr
🗃️ HATS Dataset
HATS (Human Assessed Transcription Side-by-Side) is a data set for French 🇫🇷 which consists of 1,000 triplets (reference, hypothesis A, hypothesis B) and 7,150 human choice annotated by 143 subjects 🫂 Their objective was to select, given a textual reference, which of two erroneous hypotheses is the best.
Curated by: Thibault Bañeras-Roux, Richard Dufour, Jane Wottawa, Mickael Rouvier, Teva Merlin
Funded by: Agence Nationale de la Recherche - DIETS… See the full description on the dataset page: https://huggingface.co/datasets/thibault-baneras-roux/HATS-fr.items_fullitems_raw_liteitems_raw_full
