datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
retail-products-philippinesRetailAction
Dataset Card for RetailAction
This is a FiftyOne dataset with 21000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/RetailAction")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/RetailAction.Bitext-retail-ecommerce-llm-chatbot-training-dataset
Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.retail_gaze
Dataset Card for retail_gaze
This is a FiftyOne dataset with 42 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/retail_gaze")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/retail_gaze.Retail-3I
Dataset Card for Retail-3I
Retail-3I is a retail-domain dataset built on top of tau2-bench for evaluating and training tool-calling LLM agents under three realistic user-intent conditions:
(1) Ambiguous intent: the user request is underspecified.
(2) Changing intent: the user revises or extends the goal after seeing intermediate results.
(3) Infeasible intent: the user request conflicts with tool limits, inventory/policy constraints, or missing capabilities.
There's also a… See the full description on the dataset page: https://huggingface.co/datasets/Ziyiii0-0/Retail-3I.online-retailonline_retailRetailRocket-Recommender-Dataretail-product-checkout
Retail Product Checkout (RPC) Dataset
Overview
This repository provides a Hugging Face–compatible distribution of the
Retail Product Checkout (RPC) dataset, originally introduced by
Wei et al. for research on automatic checkout and fine-grained retail
product recognition.
This dataset was not created or collected by the maintainer of this
repository. All credit for data collection, annotation, and dataset
design belongs to the original authors.
Original… See the full description on the dataset page: https://huggingface.co/datasets/benjamintli/retail-product-checkout.retail-forecast-optimize-benchmark
Retail Forecast-Optimize Benchmark
Benchmark dataset for predict-then-optimize retail inventory decisions.
Contents
File pattern
Description
{scenario}_s{seed}.json
Full policy comparison per scenario instance
samples/sample_*.json
Retail scenario definitions
eval_results.json
Aggregated leaderboard metrics
chronos-2-live-forecasts/
Live Chronos-2 outputs from HF Jobs (when available)
Scenarios
Baseline Operations
Promotion… See the full description on the dataset page: https://huggingface.co/datasets/aniketraj0224/retail-forecast-optimize-benchmark.Online-Retailonline-retailBitext-retail-banking-llm-chatbot-training-dataset
Bitext - Retail Banking Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail Banking] sector can be easily achieved using our two-step approach to LLM Fine-Tuning.… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-banking-llm-chatbot-training-dataset.Blenderkit_retail_bench_mini
Blenderkit Retail Bench Mini
Mini subset of the Blenderkit retail 3D-text benchmark (131 assets, ~4.8 GB).
Layout
Path
Description
Source GLB assets
Conditioning renders
Mesh dumps
PBR dumps
Dual-grid views (1024)
PBR voxels
Shape latents
PBR / texture latents
Per-asset metadata
Aggregate stats
Quick start
Or with mirror:
RetailAction
RetailAction Dataset
Paper: RetailAction: Dataset for Multi-View Spatio-Temporal Localization of Human-Object Interactions in RetailAccepted at: ICCV 2025 – Retail Vision WorkshopAuthors: Davide Mazzini, Alberto Raimondi, Bruno Abbate, Daniel Fischetti, David M. WoollardOrganization: Standard AI
Overview
RetailAction is a large-scale dataset designed for multi-view spatio-temporal localization of human–object interactions in real-world retail environments.
Unlike… See the full description on the dataset page: https://huggingface.co/datasets/standard-cognition/RetailAction.mSOP-765k
mSOP-765k: A Benchmark For Multi-Modal Structured Output Predictions
The mSOP-765k dataset serves as a benchmark for Multi-Modal Structured Output Predictions.
The dataset contains approximately 765k data, comprising both images and textual data.
Data
Image Data:
The images are cropped from scanned advertisement leaflets.
The image data is divided into train and test splits.
The image dataset is available in two versions: one with images resized so that the longer edge… See the full description on the dataset page: https://huggingface.co/datasets/retail-product-promotion/mSOP-765k.RetailAction
Dataset Card for RetailAction
RetailAction is a large-scale dataset for multi-view spatio-temporal localization of human-object interactions in real-world retail environments. It contains 21,000 dual-camera samples (42,000 synchronized top-view video clips, ~41 hours total) captured across 10 U.S. convenience stores. Each sample carries point-based annotations marking the precise location where a customer touches a product, along with temporal boundaries and action class… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/RetailAction.Online-Retail-II-UCIus-retail-store-closings-layoffs-warn-act-notices-daily
US retail store closings and layoffs — the actual WARN Act filings, rebuilt every day
Last rebuilt: 2026-09-22. 4,321 layoff and closure notices filed by
department stores, supermarkets and grocers, big-box and specialty chains, apparel and footwear retailers, pharmacies with a retail name, and outlet and dollar-store operators with US state labor departments — 427,305 workers,
1,167 employers, 46 states, 1988–2026.
1,874 of the notices (43.4%) were recorded by the state as a… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-retail-store-closings-layoffs-warn-act-notices-daily.smart-retail-shelf-auditing-v1abhishekrp1517_online-retail-transactions-dataset
Online Retail transactions Dataset
Analyzing Online Retail Transactions: Understanding Customer Behavior and Trends
Dataset Info
Source: Kaggle
Original Size: 29.01 MB
Kaggle Downloads: 6,396
Files: 2
Files
Online Retail.csv
Online Retail.xlsx
Mirrored from Kaggle
Mercedes_Benz_Retail_Receivables_LLC_1463814
cik
form
accessionNumber
fileNumber
filmNumber
reportDate
url
1463814
ABS-EE
0000950131-18-001228
333-212311
18946066
2018-06-30
https://sec.gov/Archives/edgar/data/1463814/000095013118001228
1463814
ABS-EE
0000950131-18-001229
333-212311
18946077
2018-06-30
https://sec.gov/Archives/edgar/data/1463814/000095013118001229
1463814
ABS-EE
0000950131-18-001670
333-212311
181029607
2018-07-31
https://sec.gov/Archives/edgar/data/1463814/000095013118001670
1463814
ABS-EE… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/Mercedes_Benz_Retail_Receivables_LLC_1463814.RetailBanking-Conversations
Dataset Description
RetailBanking-Conversations is a synthetic dataset designed to train and evaluate language models in the retail banking domain, it has been created using the open source library wizardSdata that eable the creation of synthetic datasets in any field.
The dataset contains 320 realistic conversations, across 160 unique financial profiles and 10 key retail banking topics, between financial advisors and clients, covering 10 main categories of banking products and… See the full description on the dataset page: https://huggingface.co/datasets/danystar/RetailBanking-Conversations.retail-merchbench
MerchBench
MerchBench is an open-source benchmark for retail AI model routing. It helps retailers evaluate which model or workflow tier is economically sufficient for different classes of merchandise-planning decisions.
The core idea is simple: retail AI should not be governed by model leaderboards alone. It should be governed by decision risk, reversibility, economic impact, deterministic controls, and human-review policy.
If you are searching for retail AI evaluation, LLM… See the full description on the dataset page: https://huggingface.co/datasets/Novice-ninja/retail-merchbench.retail_restock_11_feb_no_verbThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "Unitree_G1_Dex3",
"total_episodes": 124,
"total_frames": 13344,
"total_tasks": 3,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:124"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/tma-hiverobots/retail_restock_11_feb_no_verb.retail
retail
An executable Environment for evaluating and training tool-using agents. Kullback built it from recorded traces of a working agent, and Leibler publishes it.
The package holds the rebuilt world: a database, one function per tool that behaves the way the real tool was observed to behave, the compiled policy, and the Starting state each Task begins from. It also holds the Task list with the instruction a candidate gets, and a code-only Verifier per Task that grades the… See the full description on the dataset page: https://huggingface.co/datasets/leibler/retail.retailhero-uplifthttps://ods.ai/competitions/x5-retailhero-uplift-modeling
retail-bank-servicing-alignment-sft
Retail Bank Servicing Alignment SFT
The training corpus for the Granite retail-bank servicing agent. It is the
released tool-use SFT corpus merged with a servicing-alignment continuation
curriculum that teaches multi-turn behaviours the base corpus does not: what to
do when the customer says "that one", when a policy question interrupts a
transfer, when the agent's own previous turn was wrong, and when the honest
answer is that the agent cannot see what it was asked about.
Every… See the full description on the dataset page: https://huggingface.co/datasets/spkc83/retail-bank-servicing-alignment-sft.nigerian-banking-retail-transactions
Nigerian Banking Retail Transactions | Africa (Electric Sheep Africa metadata inventory)
Size category: 1M<n<10M - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/nigerian-banking-retail-transactions.m5-retail-demand-forecasting-benchmarks
M5 Retail Demand Forecasting & Inventory Risk Benchmarks
This dataset contains the heavily processed artifacts, extracted time-series features, baseline benchmarks, and model artifacts for the M5 Retail Demand Forecasting dataset.
It includes:
Over 1GB of highly engineered temporal, pricing, and calendar features.
Volatility and shortfall risk metrics for 42,840 time series.
XGBoost, Prophet, and SARIMA predictions (point + 95% intervals).
Isolation Forest anomaly detection… See the full description on the dataset page: https://huggingface.co/datasets/snchakri/m5-retail-demand-forecasting-benchmarks.
