datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
catalog
Mesh-LLM Catalog
This dataset is the Hugging Face-backed catalog for Mesh-LLM.
The runtime catalog entries live under entries/**/*.json. The Dataset Viewer
uses catalog_rows.jsonl, a flat generated table with one row per model variant.
The catalog deliberately excludes raw blob URLs. Entries should resolve to
Hugging Face repositories and canonical Mesh refs.
ine-catalog
INE
Este repositorio contiene todas las tablas¹ del Instituto Nacional de Estadística exportadas a ficheros Parquet.
Puedes encontrar cualquiera de las tablas o sus metadatos en la carpeta tablas.
Cada tabla está identificado un una ID. Puedes encontrar la ID de la tabla tanto en el INE (es el número que aparece en la URL) or en el archivo tablas.jsonl de este repositorio que puedes explorar en el Data Viewer.
Por ejemplo, la tabla de Índices nacionales de clases se corresponde… See the full description on the dataset page: https://huggingface.co/datasets/datania/ine-catalog.CATalog
Dataset Summary
CATalog is a diverse, open-source Catalan corpus for language modelling. It consists of text documents from 26 different sources, including web crawling, news, forums, digital libraries and public institutions, totaling in 17.45 billion words.
Supported Tasks and Leaderboards
Fill-Mask
Text Generation
other:Language-Modelling: The dataset is suitable for training a model in Language Modelling, predicting the next word in a given context. Success is… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CATalog.usgs-global-earthquake-catalog
USGS Global Earthquake Catalog
Provides historical data on global seismic events, sourced directly from the U.S. Geological Survey (USGS) Earthquake Hazards Program via its FDSN Event Web Service.
Each record represents a single seismic event (primarily earthquakes) and contains detailed information, including:
Event Time & Location: Precise timestamp, geographic coordinates (latitude, longitude), and depth of the event.
Magnitude: The magnitude of the event (mag) and the method… See the full description on the dataset page: https://huggingface.co/datasets/mnemoraorg/usgs-global-earthquake-catalog.european-open-data-catalogue
European Open Data Catalogue
This repository publishes independently versioned metadata and licensed source snapshots:
A discovery catalogue with 15565 dataset entries from
ISTAT, Eurostat, OECD, ILO, DoveVannoINostriSoldi (DVNS) and Cruscotto Italia.
3 independently pinned availability indexes with
911,795 joint combinations across 35 datasets, built from complete
source responses within the explicitly declared scope.
Licensed Cruscotto source snapshots, stored separately from… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-open-data-catalogue.product-catalogue
The Catalogue: Product Taxonomy Classification Benchmark
A large-scale, multimodal benchmark dataset for product taxonomy classification, featuring real e-commerce products with images, descriptions, and hierarchical category labels.
Dataset Description
The Catalogue is a benchmark dataset designed to evaluate AI models on the task of classifying products into a standardized taxonomy. Each sample includes a product image, title, description, brand, and the… See the full description on the dataset page: https://huggingface.co/datasets/Shopify/product-catalogue.sat-catalogos
SAT CFDI Catálogos
Catálogos oficiales del SAT (Servicio de Administración Tributaria) para CFDI.
Incluye los catálogos del Anexo 20 (factura electrónica y retenciones)
y los complementos de carta porte, nómina, comercio exterior y recepción de pagos.
Uso
from datasets import load_dataset, get_dataset_config_names
# Cargar un catálogo específico
ds = load_dataset("mayrop/sat-catalogos", "anexo_20_factura_electronica__4_0__c_uso_cfdi")
df = ds["train"].to_pandas()… See the full description on the dataset page: https://huggingface.co/datasets/mayrop/sat-catalogos.michael-hafftka-catalog-raisonne
Michael Hafftka – Catalog Raisonné Dataset (1970s–2025)
A rare, large-scale dataset of ~3,800 paintings by a single artist, spanning over five decades.
This dataset provides a longitudinal view of artistic development, combining images with structured metadata (title, year, medium, dimensions, collection) and including works held in major museum collections.
Example from the dataset (1995)
Why this dataset is unique
Single-artist consistency: Unlike most art… See the full description on the dataset page: https://huggingface.co/datasets/Hafftka/michael-hafftka-catalog-raisonne.gwas-catalogmalware-families-catalog
Mirrors and canonical source
This dataset is published identically across multiple platforms. The canonical source is the official SystemHelpDesk MSP site; all mirrors link back to it.
Canonical (SystemHelpDesk MSP): https://malware-families-catalog.systemhelpdesk.com/
Mirror (GitHub Pages): https://jordanricky1604-ship-it.github.io/malware-families-catalog/
GitHub repository: https://github.com/jordanricky1604-ship-it/malware-families-catalog
Hugging Face dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Jordan123234/malware-families-catalog.data_gouv_datasets_catalog-full-documents
🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée
Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data.
Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes :
titre et description,
organisation productrice,
licence,
couverture spatiale et temporelle,
fréquence de mise à jour,
formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.data-gouv-datasets-catalog
📢 Sondage 2026 : Utilisation des datasets publiques de MediaTech
Vous utilisez ce dataset ou d’autres datasets de notre collection MediaTech ? Votre avis compte !
Aidez-nous à améliorer nos datasets publiques en répondant à ce sondage rapide (5 min) : 👉 https://grist.numerique.gouv.fr/o/albert/forms/gF4hLaq9VvUog6c5aVDuMw/11
Merci pour votre contribution ! 🙌
🇫🇷 Data.gouv.fr Datasets Catalog
This dataset contains a processed and embedded version of the… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/data-gouv-datasets-catalog.shades_nationalityPossibly a placeholder dataset for the original here: https://huggingface.co/datasets/bigscience-catalogue-data/bias-shades
Data Statement for SHADES
How to use this document:
Fill in each section according to the instructions. Give as much detail as you can, but there's no need to extrapolate. The goal is to help people understand your data when they approach it. This could be someone looking at it in ten years, or it could be you yourself looking back at the data in two years.… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-catalogue-data/shades_nationality.catalogSee https://github.com/ust-archive/ust-archive for more information.
lunar-eclipse-catalog
Five Millennium Catalog of Lunar Eclipses
Credit: NASA/Apollo 8
Part of a dataset collection on Hugging Face.
Dataset description
Complete catalog of lunar eclipses spanning five millennia (-1999 to +3000), computed by Fred Espenak as part of NASA's Five Millennium Canon of Lunar Eclipses.
A lunar eclipse occurs when the Moon passes through Earth's shadow. The Moon can enter the faint penumbral shadow (producing a subtle darkening), the darker umbral… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/lunar-eclipse-catalog.bpl-card-catalog
Boston Public Library Rare Books Card Catalog Dataset
Dataset Description
This dataset contains approximately 410,000 digitized catalog cards from the Boston Public Library's Rare Books Department card catalog. The cards represent the main entry catalog (author/title cross-referenced) covering printed materials from various historical periods.
Why This Dataset?
Historical card catalogs are rich sources of:
Bibliographic information structured in… See the full description on the dataset page: https://huggingface.co/datasets/biglam/bpl-card-catalog.vision-catalog-entity-color-v1vargov-design-catalog
Vargov® Design Catalog — 605 lighting and decorative compositions in 8 languages
A machine-readable catalog of the full body of work of Vargov® Design, an author-driven
studio of lighting and decorative compositions founded by designer Anton Vargov (Moscow).
Every record is one composition: its identifier, category, canonical URLs, image links,
awards, links to its 3D model, and editorial copy written by the studio in eight
languages — Russian, English, German, Italian, French… See the full description on the dataset page: https://huggingface.co/datasets/vargov-design/vargov-design-catalog.living-catalog
Living catalog
This is the directory of record. Every public Council of AI dataset, Space, model, API, and RAG pointer is a row. Rebuild by running publish_living_catalog.py (overnight keeper calls it).
Viewer Space: https://huggingface.co/spaces/csoai/living-catalog
Living board: GET https://councilof.ai/api/gspc (counts are derived there, never typed here as scores)
Verify: https://councilof.ai/gspc-verify (free forever)
What this is not
Not a certification… See the full description on the dataset page: https://huggingface.co/datasets/csoai/living-catalog.galaxy-chirality-catalog
DESI Legacy Galaxy Chirality Catalog
This dataset accompanies the current Paper IV manuscript, An Observed-Label Chirality-Dipole Null in 949,584 High-Confidence DESI Spirals and an 8.5-Million-Galaxy Catalog.
The primary high-confidence observed-label statistic is consistent with zero under fixed-occupancy label randomization (z=0.7053169638, one-sided empirical-rank p=0.2246775322). This is not a calibrated true-spin, physical-amplitude, or primordial-parity bound.… See the full description on the dataset page: https://huggingface.co/datasets/bamfai/galaxy-chirality-catalog.model-catalogfurniture-catalog
Furniture Catalog
Image catalog and training splits from the bachelor's thesis "Visual Furnishings Compatibility Learning and Retrieval Using Machine Learning" (Ukrainian Catholic University, 2026).
5 171 individual furniture images across two room types, 1 781 room scene images, plus triplet training data (golden / train / val splits) used to train the compatibility model.
Repo layout
{room}/
{category}/ *.jpg — individual furniture images… See the full description on the dataset page: https://huggingface.co/datasets/Darebal/furniture-catalog.automobile_catalogue_jp_beirThis is a copy of https://huggingface.co/datasets/jinaai/automobile_catalogue_jp reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/automobile_catalogue_jp_beir.beverages_catalogue_ru_beirThis is a copy of https://huggingface.co/datasets/jinaai/beverages_catalogue_ru reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/beverages_catalogue_ru_beir.oceania-gov-open-data-catalog
Oceania Government Open Data — Combined Catalogue (hourly snapshot)
Combined regional catalogue of Oceania (Australia + New Zealand) public-service open
data harvested from both data.gov.au and data.govt.nz portals, including state,
territory and local-council publishers.
License declaration
License: other (see below). Records in this catalogue inherit the licence of their
source dataset. Where the source declares a standard open licence the record is tagged
with… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/oceania-gov-open-data-catalog.bright-star-catalog
Bright Star Catalogue (BSC5)
Part of the Astronomy Datasets collection on Hugging Face.
The Bright Star Catalogue (BSC5, 5th Revised Edition) containing 9,110 naked-eye stars
brighter than visual magnitude ~6.5 with UBVRI photometry, MK spectral types, proper motions,
radial and rotational velocities, and multiplicity information.
Dataset description
The Bright Star Catalogue (Hoffleit & Warren, 1991) is THE standard reference for naked-eye
stars. Originally compiled… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/bright-star-catalog.games-catalog
Council of AI — games catalog
Catalog door. Games load into Council Space. No contest engine on this card.
Live: https://councilof.ai/gspc-arena
Council OS: https://councilof.ai/os
Council Space: https://councilof.ai/gspc-arena
Measurement, not certification. Empty slots are not for sale. No scores on this card.
Jail is a measured floor, not a 16th pane.
The live board is the authority
GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub… See the full description on the dataset page: https://huggingface.co/datasets/csoai/games-catalog.phenotype-catalog
Ethnic Erotic Phenotype Catalog
A structured complement to Wikipedia for ethnographic data — 1,700+ ethnic groups indexed with normalized linguistic, geographic, cultural, and phenotype metadata, plus 23K+ notable-people references and 5K+ vision-grounded per-image phenotype observations.
Curated from the live catalog at ethnicerotic.com and published as an open dataset for anthropological reference, AI training, and ethnographic research.
What's in v6
Two columns… See the full description on the dataset page: https://huggingface.co/datasets/EthnicErotic/phenotype-catalog.bpl-card-catalog-lance-fullnarit-ghosts-halo-catalogs
NARIT GHOSTS Halo Catalogs
Dataset Description
This dataset contains reduced stellar catalogs, combined FITS images, and candidate substructure catalogs from the GHOSTS (Galaxy Halos, Outer disks, Substructure, Thick disks, and Star clusters) Survey observed by the Hubble Space Telescope (HST).
It serves as the primary data lake for the automated astronomical pipeline designed to detect faint stellar substructures (like Ultra-Faint Dwarfs and stellar streams) in… See the full description on the dataset page: https://huggingface.co/datasets/appleboiy/narit-ghosts-halo-catalogs.
