datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
E-commerce
E-commerce Dataset (AutoGEO)
This is a commercial-domain dataset released with AutoGEO for Generative Engine Optimization (GEO) research.
📄 Paper: "What Generative Search Engines Like and How to Optimize Web Content Cooperatively"👥 Authors: Yujiang Wu*, Shanshan Zhong*, Yubin Kim, Chenyan Xiong (*Equal contribution)🚀 Code: AutoGEO on GitHub
Dataset Configurations
main: Primary train/test data for GEO training and evaluation (~1.6k train / ~400 test)… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/E-commerce.ecommerce-search-extraction
Ionio E-commerce Search Query Extraction
Built with: simula — schema-driven synthetic data generation with auditable taxonomy lineage.
An English synthetic dataset for training and evaluating systems that translate natural-language
shopping requests into narrow, atomic, database-queryable JSON. It contains 10,985 accepted
examples from a 13,000-attempt generation run. No accepted rows were trimmed from this release.
Each example pairs a realistic typed or spoken shopper query… See the full description on the dataset page: https://huggingface.co/datasets/Ionio-ai/ecommerce-search-extraction.Sora-Ecommerce-Guide
Sora Ecommerce Guide Dataset
This dataset contains comprehensive documentation, user guides, admin operating procedures, and system flow architectures for the Sora Ecommerce platform, structured in flat instruction/input/output format matching standard fine-tuning benchmarks.
Splits
train: 9 samples
test: 2 samples
Features
instruction: System/task instruction context.
input: The prompt, question, or user query.
output: Complete step-by-step… See the full description on the dataset page: https://huggingface.co/datasets/HeinKoZin/Sora-Ecommerce-Guide.ecommerce-chat-tool-calling
E-commerce Chat Tool-Calling Dataset (generic, schema-following)
Synthetic training data that teaches a small model (e.g.
google/functiongemma-270m-it) to map natural-language shopping requests
onto whatever tool schema is declared in the prompt — not onto one
hard-coded API.
A visitor chats with the store's corner chatbot: "I need an inexpensive
top-loading washing machine, preferably from a German manufacturer" → the
model must emit the declared tool call with the right query… See the full description on the dataset page: https://huggingface.co/datasets/Qrzysztof/ecommerce-chat-tool-calling.sea-ecommerce-customer-support-sample
SEA Multilingual E-commerce Customer Support Sample
This public sample contains 1,000 synthetic, AI-generated customer-support
conversations for Southeast Asian e-commerce scenarios.
Languages
English
Chinese
Malay
Indonesian
Formats
CSV
JSONL
Intended Use
Use this sample for inspection, evaluation, prototyping, multilingual testing,
and intent-classification experiments.
Important Limitations
This is synthetic… See the full description on the dataset page: https://huggingface.co/datasets/nwchang/sea-ecommerce-customer-support-sample.ecommerce-query-rewriting
#e-commerce-query-rewriting-dataset
Hub: mudasir13cs/ecommerce-query-rewriting
A dataset of 10,000 examples pairing ambiguous, context-dependent user queries with their fully resolved, context-aware rewrites for e-commerce product search. Built for fine-tuning LLMs to resolve pronouns, ellipsis, ordinals, and other conversational shortcuts using prior search context — the kind of resolution real shopping assistants need to handle turns like "show me that one" or "the cheaper… See the full description on the dataset page: https://huggingface.co/datasets/mudasir13cs/ecommerce-query-rewriting.Devanagari-Ecommerce-Dataset
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Prepared by:
Aakash… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Devanagari-Ecommerce-Dataset.Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset.TR-ECommerce-CustomerSupport-Instructions
TR E-Commerce Customer Support Instructions 🇹🇷
A high-quality Turkish E-Commerce Customer Support dataset designed for fine-tuning large language models (LLMs) on instruction-following customer support tasks.
Dataset Summary
Feature
Value
Language
Turkish (tr)
Domain
E-Commerce Customer Support
Format
Conversation (Conversational/Chat)
Chain-of-Thought (CoT)
✅ Natural paragraph reasoning (thinking field)
Total Categories
20
Total Rows
186… See the full description on the dataset page: https://huggingface.co/datasets/Mer1Alii/TR-ECommerce-CustomerSupport-Instructions.ecommerce-faq-llama2-chatskincare-ecommerce-FAQ
Dataset Card for Dataset Name
FAQ template for skincare E-Commerce
Dataset Details
Dataset Description
his dataset contain question and answer on the FAQ between a buyer and a seller in a skincare online shop
dataset contain 2 columns of [Buyer] and [AI Assistant]
Curated by: UrFavB0i
Language(s) (NLP): English-only
License: Apache 2.0
Uses
this dataset can be used to fine tune Large Language Model like Falcon, LLama model etc
Direct Use… See the full description on the dataset page: https://huggingface.co/datasets/UrFavB0i/skincare-ecommerce-FAQ.ecommerce-product-descriptions-v1
E-Commerce Product Descriptions Dataset
This dataset contains 150 high-quality synthetic product descriptions across multiple e-commerce categories including earbuds, smartwatches, and laptops.
Dataset Structure
Each entry follows the Alpaca-style instruction format:
: Task description
cmd: Failure calling service input: Failed transaction (2147483646): Context/brand features
: Generated product description
: Product category
: Uniqueness ratio (0.7 - 0.99)… See the full description on the dataset page: https://huggingface.co/datasets/Rahul5262/ecommerce-product-descriptions-v1.sea-ecommerce-customer-support-sample
SEA Multilingual E-commerce Customer Support Sample
This public sample contains 1,000 synthetic, AI-generated customer-support
conversations for Southeast Asian e-commerce scenarios.
Languages
English
Chinese
Malay
Indonesian
Formats
CSV
JSONL
Intended Use
Use this sample for inspection, evaluation, prototyping, multilingual testing,
and intent-classification experiments.
Important Limitations
This is synthetic… See the full description on the dataset page: https://huggingface.co/datasets/dinesshxryu/sea-ecommerce-customer-support-sample.ecommerce-production-incident-postmortems
Autonomous E-Commerce Production Incident Post-Mortems (NexusOS v2.0)
Dataset Summary
This repository contains high-fidelity, meticulously graded post-mortem diagnostic samples mapping real-world infrastructure failures in e-commerce ecosystems (WordPress, WooCommerce, and core transaction layers) directly to expert root-cause analyses and isolated code execution patches.
The data is structured into optimized binary Parquet formats, making it immediately ready for… See the full description on the dataset page: https://huggingface.co/datasets/Amman-shah/ecommerce-production-incident-postmortems.
