datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
product-database
Open Food Facts Database
What is 🍊 Open Food Facts?
A food products database
Open Food Facts is a database of food products with ingredients, allergens, nutrition facts and all the tidbits of information we can find on product labels.
Made by everyone
Open Food Facts is a non-profit association of volunteers. 25.000+ contributors like you have added 1.7 million + products from 150 countries using our Android or iPhone app or their… See the full description on the dataset page: https://huggingface.co/datasets/openfoodfacts/product-database.grob-products-updated
bep40/grob-products-updated
Generated by ML Intern
This dataset repository was generated by ML Intern, an agent for machine learning research and development on the Hugging Face Hub.
Try ML Intern: https://smolagents-ml-intern.hf.space
Source code: https://github.com/huggingface/ml-intern
Usage
from datasets import load_dataset
dataset = load_dataset('bep40/grob-products-updated')
cbi-released-products
Cosmic Background Imager released products
This repository contains the numerical products released for four generations of Cosmic Background Imager (CBI) analysis: the 2000 deep fields, the 2000 mosaic fields, the final 2002–2005 temperature and polarization analysis, and the final 2000–2005 total-intensity analysis. It contains the published band powers, their window functions, the same-band correlation blocks and full Fisher matrices distributed in CosmoMC .newdat files, and… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/cbi-released-products.data-product-benchmark
DPDisc Dataset
Paper | Code
Dataset Description
This dataset provides a benchmark for automatic data product creation. The task is framed as follows: given a natural language data product request and a corpus of text and tables, the objective is to identify the relevant tables and text documents that should be included in the resulting data product which would useful to the given data product request. The benchmark brings together three variants: HybridQA, TAT-QA, and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/data-product-benchmark.AMAZON-Products-2023
Dataset Card for Amazon Products 2023
Dataset Summary
This dataset contains product metadata from Amazon, filtered to include only products that became available in 2023. The dataset is intended for use in semantic search applications and includes a variety of product categories.
Number of Rows: 117,243
Number of Columns: 15
Data Source
The data is sourced from Amazon Reviews 2023.
It includes product information across multiple categories, with… See the full description on the dataset page: https://huggingface.co/datasets/milistu/AMAZON-Products-2023.msam-released-products
MSAM released flight products
This dataset contains the complete numeric contents of the three MSAM1 flight
archives served by LAMBDA: the June 1992, June 1994, and June 1995 observation
tables, sampled beam maps, and released 1994/1995 covariance matrices. The 38
configuration identifiers preserve the flight directory and source filename
stem.
How to use
python -m venv .venv && .venv/bin/pip install datasets huggingface_hub pyarrow
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/msam-released-products.vsa-released-products
VSA February 2004 released products
This dataset contains LAMBDA's February 2004 VSA spectrum, x factors,
16 window functions, and Hessian product. LAMBDA serves the Hessian as 256
lines with one numeric value per line; its documentation identifies those
values as the 16×16 M_ij, but does not state how flat positions map to
matrix indices. Unlabelled source columns remain Column 1, Column 2, and
so on.
Use
python -m venv .venv && .venv/bin/pip install datasets… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/vsa-released-products.amazon_productsMoRoOp
MoRoOp: A Dataset of Autonomous Mobile Robot Operations in Production Logistics
MoRoOp contains 72 hours of operational data from an autonomous mobile robot (AMR). The data was recorded at the KTH IPU Lab as part of the Self-Organizing Production Logistics (SOPL) project. The current release (1.0.0) covers five consecutive days, 2026-07-27 through 2026-07-31, organized into morning and evening shifts.
Data organization
The data/ directory contains one folder for… See the full description on the dataset page: https://huggingface.co/datasets/Self-Organizing-Production-Logistics/MoRoOp.products-2017Many e-shops have started to mark-up product data within their HTML pages using the schema.org vocabulary. The Web Data Commons project regularly extracts such data from the Common Crawl, a large public web crawl. The Web Data Commons Training and Test Sets for Large-Scale Product Matching contain product offers from different e-shops in the form of binary product pairs (with corresponding label "match" or "no match")
In order to support the evaluation of machine learning-based matching methods, the data is split into training, validation and test set. We provide training and validation sets in four different sizes for four product categories. The labels of the test sets were manually checked while those of the training sets were derived using shared product identifiers from the Web via weak supervision.
The data stems from the WDC Product Data Corpus for Large-Scale Product Matching - Version 2.0 which consists of 26 million product offers originating from 79 thousand websites.wb-products
Dataset Card for Wildberries products
Dataset Summary
This dataset was scraped from product pages on the Russian marketplace Wildberries. It includes all information from the product card and metadata from the API, excluding image URLs. The dataset was collected by processing approximately 160 million products out of a potential 230 million, starting from the first product. Data collection had to be stopped due to serious rate limits that prevented further progress. The… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/wb-products.amazon_product_reviews_video_games#Title
class-released-products
CLASS released measurement products
The preview renders the released Q_POLARISATION field from
class_dr1_40GHz_skymap_n128 in its source RING order.
This dataset contains LAMBDA's three CLASS DR1 40 GHz maps, three 90 GHz EE
spectra, and 40 GHz circular-polarization limits. The release's
transfer functions,
beam,
bandpass,
simulations,
masks,
synchrotron-beta,
reobserved and
combined
auxiliary maps, and software are excluded. Configuration names are source
filename stems.… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/class-released-products.amazon-product-datasetAMAZON-Products-2023-Arabic
Dataset Card for Amazon Products 2023 Arabic
Dataset Summary
This dataset contains product metadata from Amazon, filtered to include only products that became available in 2023. The dataset is intended for use in semantic search applications and includes a variety of product categories.
Number of Rows: 117,243
Number of Columns: 17
Data Source
The data is sourced from Amazon Reviews 2023.
It includes product information across multiple categories, with… See the full description on the dataset page: https://huggingface.co/datasets/milistu/AMAZON-Products-2023-Arabic.hm_ecommerce_products
H&M Personalized Fashion Recommendations - Enhanced Dataset
Dataset Description
This dataset is a processed and enhanced version of the H&M Personalized Fashion Recommendations Kaggle competition dataset. The original dataset has been cleaned and augmented with pre-computed embeddings and accessible image URLs to facilitate fashion recommendation research and multimodal retrieval applications.
Dataset Summary
The H&M dataset contains rich product metadata… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/hm_ecommerce_products.so100_guess_who_baseThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 95,
"total_frames": 51045,
"total_tasks": 1,
"total_videos": 95,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:95"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/produc-xuan/so100_guess_who_base.product-taxonomy-bench
Dataset Summary
product-taxonomy-bench is an anonymised benchmark dataset for predicting Shopify Product Taxonomy categories from Shopify product tags.
This dataset does not include raw product titles, raw tags, or product URLs. Tags are anonymised as tagNNNNNN.
Start Here
Read this dataset card for the snapshot layout and field definitions.
Open the benchmark notebook at notebooks/product_taxonomy_bench.ipynb. It defaults to the fixed paper snapshot at revision… See the full description on the dataset page: https://huggingface.co/datasets/gregb/product-taxonomy-bench.amazon-product-searchPlease refer the original repository https://github.com/amazon-science/esci-data.
africa-synth-energy-oilgas-gas-production-nigeria
Africa Synth Energy Oilgas Gas Production Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-oilgas-gas-production-nigeria.africa-owid-gas-production-by-country
Gas Production By Country | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-gas-production-by-country.cube-move-prod-4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 15880,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/benlevin/cube-move-prod-4.africa-owid-natural-gas-production-by-region-terawatt-hours-twh
Natural Gas Production By Region Terawatt Hours Twh | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: parquet - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-natural-gas-production-by-region-terawatt-hours-twh.groot1_smileface_Prod_try2
groot1_smileface_Prod_try2
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
rebot-prod1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "seeed_b601_dm_follower",
"total_episodes": 43,
"total_frames": 38530,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:43"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/tinjyuu/rebot-prod1.record-pickandplace-prodThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 31494,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/kakimoto/record-pickandplace-prod.africa-synth-textile-garment-production-all
African Textile & Garment Production | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: csv - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-textile-garment-production-all.remote-worker-productivity
Key Features:
Primary Research Focus:
Age vs Productivity correlation
Years of Experience impact on remote work effectiveness
WFH Days per Week optimal balance analysis
Multiple productivity metrics (not just one score)
Dataset Highlights:
1,500 rows - Perfect size for analysis
30+ columns - Rich feature set
Realistic correlations - Built-in meaningful relationships
Clean data - No missing values, proper data types
Multiple target variables - 5 different… See the full description on the dataset page: https://huggingface.co/datasets/nprak26/remote-worker-productivity.quiet-released-products
QUIET released measurement products
This dataset mirrors a selected eight-file fragment of the QUIET release served
by LAMBDA: the four raw maps under its “QUIET Galactic Fields 1st+2nd Season
Observations” heading and the four first- and second-season angular-power
tables. It is not the complete QUIET release. LAMBDA separately releases
raw-map covariance matrices, QUIET+WMAP coadds and their covariance matrices,
masks, galaxy-map data, beams, and bandpasses.
Each… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/quiet-released-products.recalldb-product-recalls-sample
RecallDB — U.S. Product Recall Database (Sample)
Full dataset: recalldb.dataengineered.io · $49 one-time snapshot → Buy on Stripe · the same sample on Kaggle
127,783 official recalls · 292,790 recalled products · CPSC · FDA · FSIS · NHTSA · USCG · 100% source-linked
RecallDB is a normalized, provenance-tracked dataset of official U.S. federal product recalls. It joins five official source families into one relational model: CPSC consumer products, NHTSA vehicles, FDA/openFDA… See the full description on the dataset page: https://huggingface.co/datasets/Ichlibitiche/recalldb-product-recalls-sample.
