datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
catalog
Mesh-LLM Catalog
This dataset is the Hugging Face-backed catalog for Mesh-LLM.
The runtime catalog entries live under entries/**/*.json. The Dataset Viewer
uses catalog_rows.jsonl, a flat generated table with one row per model variant.
The catalog deliberately excludes raw blob URLs. Entries should resolve to
Hugging Face repositories and canonical Mesh refs.
xr-motion-dataset-catalogue
XR Motion Dataset Catalogue
Overview
The XR Motion Dataset Catalogue, accompanying our paper "Navigating the Kinematic Maze: A Comprehensive Guide to XR Motion Dataset Standards," standardizes and simplifies access to Extended Reality (XR) motion datasets. The catalogue represents our initiative to streamline the usage of kinematic data in XR research by aligning various datasets to a consistent format and structure.
Dataset Specifications
All datasets in this… See the full description on the dataset page: https://huggingface.co/datasets/cschell/xr-motion-dataset-catalogue.ine-catalog
INE
Este repositorio contiene todas las tablas¹ del Instituto Nacional de Estadística exportadas a ficheros Parquet.
Puedes encontrar cualquiera de las tablas o sus metadatos en la carpeta tablas.
Cada tabla está identificado un una ID. Puedes encontrar la ID de la tabla tanto en el INE (es el número que aparece en la URL) or en el archivo tablas.jsonl de este repositorio que puedes explorar en el Data Viewer.
Por ejemplo, la tabla de Índices nacionales de clases se corresponde al… See the full description on the dataset page: https://huggingface.co/datasets/datania/ine-catalog.CATalog
Dataset Summary
CATalog is a diverse, open-source Catalan corpus for language modelling. It consists of text documents from 26 different sources, including web crawling, news, forums, digital libraries and public institutions, totaling in 17.45 billion words.
Supported Tasks and Leaderboards
Fill-Mask
Text Generation
other:Language-Modelling: The dataset is suitable for training a model in Language Modelling, predicting the next word in a given context. Success is… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/CATalog.usgs-global-earthquake-catalog
USGS Global Earthquake Catalog
Provides historical data on global seismic events, sourced directly from the U.S. Geological Survey (USGS) Earthquake Hazards Program via its FDSN Event Web Service.
Each record represents a single seismic event (primarily earthquakes) and contains detailed information, including:
Event Time & Location: Precise timestamp, geographic coordinates (latitude, longitude), and depth of the event.
Magnitude: The magnitude of the event (mag) and the method… See the full description on the dataset page: https://huggingface.co/datasets/mnemoraorg/usgs-global-earthquake-catalog.Products-Catalogeuropean-open-data-catalogue
European Open Data Catalogue
This repository publishes independently versioned metadata and licensed source snapshots:
A discovery catalogue with 15565 dataset entries from
ISTAT, Eurostat, OECD, ILO, DoveVannoINostriSoldi (DVNS) and Cruscotto Italia.
3 independently pinned availability indexes with
911,795 joint combinations across 35 datasets, built from complete
source responses within the explicitly declared scope.
Licensed Cruscotto source snapshots, stored separately from… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/european-open-data-catalogue.product-catalogue
The Catalogue: Product Taxonomy Classification Benchmark
A large-scale, multimodal benchmark dataset for product taxonomy classification, featuring real e-commerce products with images, descriptions, and hierarchical category labels.
Dataset Description
The Catalogue is a benchmark dataset designed to evaluate AI models on the task of classifying products into a standardized taxonomy. Each sample includes a product image, title, description, brand, and the… See the full description on the dataset page: https://huggingface.co/datasets/Shopify/product-catalogue.mercadona-catalog
🍍 Mercadona Catalog
Catálogo completo de productos y precios de la tienda online de Mercadona, exportado semanalmente desde su API pública (no oficial).
Estructura
Archivo
Contenido
categories.json
Árbol de categorías (secciones y subcategorías)
categories/<id>.json
Detalle por categoría con productos asociados
product_ids.json
Índice con todos los IDs de producto
products/<id>.json
Detalle completo por producto (precio, descripción, imágenes… See the full description on the dataset page: https://huggingface.co/datasets/datania/mercadona-catalog.botrail-catalog
botrail-catalog
Robot manipulator / end-effector model catalog for botrail.
Browse it visually — thumbnails, datasheets, 3D with joint sliders:
spaces/botrail/catalog
Entry point: index.json
Each package ships normalized URDF / USD (usdc) / GLB / collision meshes with
per-file license records and a validation report (levels V0-V5).
distribution: recipe_only packages publish metadata + thumbnail only
(upstream licenses do not permit redistributing the sources). One that a
newer… See the full description on the dataset page: https://huggingface.co/datasets/botrail/botrail-catalog.sat-catalogos
SAT CFDI Catálogos
Catálogos oficiales del SAT (Servicio de Administración Tributaria) para CFDI.
Incluye los catálogos del Anexo 20 (factura electrónica y retenciones)
y los complementos de carta porte, nómina, comercio exterior y recepción de pagos.
Uso
from datasets import load_dataset, get_dataset_config_names
# Cargar un catálogo específico
ds = load_dataset("mayrop/sat-catalogos", "anexo_20_factura_electronica__4_0__c_uso_cfdi")
df = ds["train"].to_pandas()… See the full description on the dataset page: https://huggingface.co/datasets/mayrop/sat-catalogos.michael-hafftka-catalog-raisonne
Michael Hafftka – Catalog Raisonné Dataset (1970s–2025)
A rare, large-scale dataset of ~3,800 paintings by a single artist, spanning over five decades.
This dataset provides a longitudinal view of artistic development, combining images with structured metadata (title, year, medium, dimensions, collection) and including works held in major museum collections.
Example from the dataset (1995)
Why this dataset is unique
Single-artist consistency: Unlike most art… See the full description on the dataset page: https://huggingface.co/datasets/Hafftka/michael-hafftka-catalog-raisonne.gwas-catalogplaywright_with_chunk_google_map_huggingface_379_d4a19c_catalog
Derived: Product Catalog Insights
Aggregated product catalog statistics enriched with ticket-derived signals.
Attribution & Redistribution
This dataset is derived from the following upstream source(s):
Upstream Source
Repository
License
github_fetch_huggingface_terminal_9046_nxauzm_upstream_support_tickets
TianfuXinqu/github_fetch_huggingface_terminal_9046_nxauzm_upstream_support_tickets
apache-2.0
Redistribution is permitted under the apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/TianfuXinqu/playwright_with_chunk_google_map_huggingface_379_d4a19c_catalog.malware-families-catalog
Mirrors and canonical source
This dataset is published identically across multiple platforms. The canonical source is the official SystemHelpDesk MSP site; all mirrors link back to it.
Canonical (SystemHelpDesk MSP): https://malware-families-catalog.systemhelpdesk.com/
Mirror (GitHub Pages): https://jordanricky1604-ship-it.github.io/malware-families-catalog/
GitHub repository: https://github.com/jordanricky1604-ship-it/malware-families-catalog
Hugging Face dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Jordan123234/malware-families-catalog.data_gouv_datasets_catalog-full-documents
🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée
Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data.
Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes :
titre et description,
organisation productrice,
licence,
couverture spatiale et temporelle,
fréquence de mise à jour,
formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.data-gouv-datasets-catalog
📢 Sondage 2026 : Utilisation des datasets publiques de MediaTech
Vous utilisez ce dataset ou d’autres datasets de notre collection MediaTech ? Votre avis compte !
Aidez-nous à améliorer nos datasets publiques en répondant à ce sondage rapide (5 min) : 👉 https://grist.numerique.gouv.fr/o/albert/forms/gF4hLaq9VvUog6c5aVDuMw/11
Merci pour votre contribution ! 🙌
🇫🇷 Data.gouv.fr Datasets Catalog
This dataset contains a processed and embedded version of the… See the full description on the dataset page: https://huggingface.co/datasets/AgentPublic/data-gouv-datasets-catalog.bigbounce-anomaly-catalog
BigBounce — Multi-Survey Autoencoder Anomaly Catalog (Paper 3)
Corrected release description. This dataset backs Golden (2026), Paper 3 —
A Multi-Survey Autoencoder Anomaly-Candidate Catalog. It is released
CC-BY-4.0. The corrected submission release is frozen at the immutable tag
p3-v3.1.161. The frozen inventory is heterogeneous and
is not a complete six-survey, independently rerunnable per-object product.
Consumers should verify downloaded files against
RELEASE_MANIFEST.json… See the full description on the dataset page: https://huggingface.co/datasets/bamfai/bigbounce-anomaly-catalog.catalogSee https://github.com/ust-archive/ust-archive for more information.
visual-product-cataloguemeilisearch-catalog-dumplunar-eclipse-catalog
Five Millennium Catalog of Lunar Eclipses
Credit: NASA/Apollo 8
Part of a dataset collection on Hugging Face.
Dataset description
Complete catalog of lunar eclipses spanning five millennia (-1999 to +3000), computed by Fred Espenak as part of NASA's Five Millennium Canon of Lunar Eclipses.
A lunar eclipse occurs when the Moon passes through Earth's shadow. The Moon can enter the faint penumbral shadow (producing a subtle darkening), the darker umbral… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/lunar-eclipse-catalog.shades_nationalityPossibly a placeholder dataset for the original here: https://huggingface.co/datasets/bigscience-catalogue-data/bias-shades
Data Statement for SHADES
How to use this document:
Fill in each section according to the instructions. Give as much detail as you can, but there's no need to extrapolate. The goal is to help people understand your data when they approach it. This could be someone looking at it in ten years, or it could be you yourself looking back at the data in two years.… See the full description on the dataset page: https://huggingface.co/datasets/bigscience-catalogue-data/shades_nationality.animezone-catalogoverlap-catalog-ego100k
Overlap catalog: Egocentric-100K
Perceptual fingerprints of builddotai/Egocentric-100K,
computed with overlap at 4 fps on
decoded pixels - so re-encoding, container swaps and metadata stripping do not
hide a match.
Import this into a local overlap index and every dataset you are offered is
screened against the public corpus:
overlap import block00 block01 ... # or just the blocks you want
overlap compare vendor-offer.ovlm # "81% of this is public Egocentric-100K"… See the full description on the dataset page: https://huggingface.co/datasets/WorldArchive/overlap-catalog-ego100k.vision-catalog-entity-color-v1living-catalog
Living catalog
This is the directory of record. Every public Council of AI dataset, Space, model, API, and RAG pointer is a row. Rebuild by running publish_living_catalog.py (overnight keeper calls it).
Viewer Space: https://huggingface.co/spaces/csoai/living-catalog
Living board: GET https://councilof.ai/api/gspc (counts are derived there, never typed here as scores)
Verify: https://councilof.ai/gspc-verify (free forever)
What this is not
Not a certification… See the full description on the dataset page: https://huggingface.co/datasets/csoai/living-catalog.galaxy-chirality-catalog
DESI Legacy Galaxy Chirality Catalog
This dataset accompanies the current Paper IV manuscript, An Observed-Label Chirality-Dipole Null in 949,584 High-Confidence DESI Spirals and an 8.5-Million-Galaxy Catalog.
The primary high-confidence observed-label statistic is consistent with zero under fixed-occupancy label randomization (z=0.7053169638, one-sided empirical-rank p=0.2246775322). This is not a calibrated true-spin, physical-amplitude, or primordial-parity bound.… See the full description on the dataset page: https://huggingface.co/datasets/bamfai/galaxy-chirality-catalog.model-catalogfurniture-catalog
Furniture Catalog
Image catalog and training splits from the bachelor's thesis "Visual Furnishings Compatibility Learning and Retrieval Using Machine Learning" (Ukrainian Catholic University, 2026).
5 171 individual furniture images across two room types, 1 781 room scene images, plus triplet training data (golden / train / val splits) used to train the compatibility model.
Repo layout
{room}/
{category}/ *.jpg — individual furniture images… See the full description on the dataset page: https://huggingface.co/datasets/Darebal/furniture-catalog.
