datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bangla-english-and-code-mixed-ecommerce-review-dataset
BanglishRev: A Large-Scale Bangla-English and Code-mixed Dataset of Product Reviews in E-Commerce
Description
The BanglishRev dataset is the largest e-commerce product review dataset to date for reviews written in Bengali, English, a mixture of both and Banglish, Bengali words written with English alphabets. The dataset comprises of 1.74 million written reviews from 3.2 million ratings information collected from a total of 128k products being sold in online… See the full description on the dataset page: https://huggingface.co/datasets/BanglishRev/bangla-english-and-code-mixed-ecommerce-review-dataset.marqo-general-ecommerce-evalai-generated-ecommerce-images
AI-Generated E-Commerce Images
Overview
This dataset contains 6,031 AI-generated images depicting common e-commerce
after-sales scenarios across 12 categories. The dataset also ships with
companion annotations (annotations_ai.jsonl) describing each image in
chat-completion format.
Dataset Structure
Category
Description
damaged_electronics
Damaged electronics (keyboards, laptops, headphones, etc.)
damaged_phone_screen
Smartphones with cracked or… See the full description on the dataset page: https://huggingface.co/datasets/JoyCN/ai-generated-ecommerce-images.ecommerce_products_clipEcommerce_Datasetai-generated-ecommerce-images
AI-Generated E-Commerce Images
Overview
This dataset contains 6,031 AI-generated images depicting common e-commerce
after-sales scenarios across 12 categories. The dataset also ships with
companion annotations (annotations_ai.jsonl) describing each image in
chat-completion format.
Dataset Structure
Category
Description
damaged_electronics
Damaged electronics (keyboards, laptops, headphones, etc.)
damaged_phone_screen
Smartphones with cracked or… See the full description on the dataset page: https://huggingface.co/datasets/prajwalkothwal/ai-generated-ecommerce-images.MM-Bench-E-CommerceThis is the HuggingFace repository of the paper named MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding in WSDM 2026 (oral).
In this paper, we argue that generative Multimodal Large Language Models (MLLMs) hold significant potential for improving product representation learning.
We propose the first generative MLLM-based model named MOON for product representation learning.
Furthermore, we contruct and publish a large-scale real-world… See the full description on the dataset page: https://huggingface.co/datasets/Daoze/MM-Bench-E-Commerce.temu-ecommerce-pricing-workflow-sample
Temu E-commerce Pricing & SKU Variant Dataset
85 products · 382 SKUs · $2.14–$219.72 price range · avg 29.5% discount where measurable
A real, production-quality sample of Temu product listings and SKU-level pricing data, captured by Octoparse Managed Data Service via a managed anti-bot pipeline. Every row is real market data — no synthetic expansion, no mock prices.
Built for teams working on competitor price monitoring, dynamic pricing models, product matching, and e-commerce AI… See the full description on the dataset page: https://huggingface.co/datasets/Octoparse/temu-ecommerce-pricing-workflow-sample.E-Commerce-Product-Recommendationindian-fashion-ecommerceecommerce-taxonomye-commerce-sample-imagesecommerce_customer_behavior_datasetecommerce-taxonomybakllava-multiturn-ecommerceproducts_ecommerce_embeddings
Dataset Card for "products_ecommerce_embeddings"
The dataset is based on https://github.com/querqy/chorus/tree/main/data-encoder
More Information needed
labelled-ecommerce-taxonomyebay_ecommerce_sportfood_part1labelled-ecommerce-taxonomyecommerce-taxonomyecommerce-taxonomylabelled-ecommerce-taxonomyml-datasets-image-rakuten-ecommerceEcommerce_QAEcommercee-commerce-image-captioning
Dataset Card for Turkish E-Commerce Product Captions
Dataset Details
Dataset Description
This dataset consists of Turkish captions for product images collected from publicly accessible pages on Trendyol. The captions were generated using a captioning model and curated manually by the dataset author for academic and research purposes. The dataset includes approximately 200K samples and is intended for image captioning tasks in Turkish.
Curated by: Mustafa… See the full description on the dataset page: https://huggingface.co/datasets/forza61/e-commerce-image-captioning.ecommerce-search-artifactslabelled-ecommerce-taxonomylabelled-ecommerce-taxonomytimescale-ecommerce
