datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wb-products
Dataset Card for Wildberries products
Dataset Summary
This dataset was scraped from product pages on the Russian marketplace Wildberries. It includes all information from the product card and metadata from the API, excluding image URLs. The dataset was collected by processing approximately 160 million products out of a potential 230 million, starting from the first product. Data collection had to be stopped due to serious rate limits that prevented further progress. The… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/wb-products.task929_products_reviews_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task929_products_reviews_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task929_products_reviews_classification.ulasan-beauty-products-qaproduct-bench
Product Bench
31-turn multi-turn speech-to-speech benchmark for evaluating voice AI models as a laptop comparison shopping assistant.
Part of Audio Arena, a suite of 6 benchmarks spanning 221 turns across different domains. Built by Arcada Labs.
Leaderboard | GitHub | All Benchmarks
Dataset Description
The model acts as a laptop comparison shopping assistant helping a customer evaluate, compare, and order laptops. The conversation features multi-intent turns… See the full description on the dataset page: https://huggingface.co/datasets/arcada-labs/product-bench.production-ai-guardrail-evals
Production AI Guardrail Evals
Eighteen synthetic, assertion-bearing cases for testing whether a language model can stay inside an advisory role. The cases cover roadmap intake, release readiness, and catalog-change review—the same workload families I use in the Winwood AI Toolkit around IEM Rig.
This is a sanitized public derivative, not a dump of application logs or the private evaluation corpus.
What each row contains
case_id: stable public identifier;… See the full description on the dataset page: https://huggingface.co/datasets/mattwinwood/production-ai-guardrail-evals.task1088_array_of_products
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1088_array_of_products
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1088_array_of_products.saas-product-support-sharegpt-1k
SaaS/Tech Product Support — Multi-Turn SFT Dataset
A domain-specific supervised fine-tuning dataset for
SaaS and tech product support conversations, built for
LLM fine-tuning and instruction tuning.
Dataset Summary
This dataset contains 1,200 multi-turn English conversations
between a customer and a support agent, covering common
SaaS/tech support scenarios: bug reports, billing issues,
API errors, authentication problems, onboarding blockers,
integration failures… See the full description on the dataset page: https://huggingface.co/datasets/Dang-DN-VN/saas-product-support-sharegpt-1k.apparel23-qwen32b-kept-outfits-with-products
Apparel'23 kept outfit bundles (with products)
Amazon Apparel 2023 outfit bundles (need → bundle of role-tagged products), the gold source
for DeepShopper's AMZ side: trains the reward V0 (qwen4b-apparel23-bundle-sft),
seeds reward V1 positives and the Reducer soft-gold targets, and provides the AMZ
mapper needs. sft.{train,test}.jsonl. Derived from public Amazon data; research use. Code: https://github.com/clijo/reco-rl (branch outfit_bundle).
amazon-products
Amazon Products Sample Dataset
A curated sample of 2,000 popular products from the Amazon Reviews 2023 dataset, designed for educational use in building RAG (Retrieval-Augmented Generation) systems and shopping agents.
Dataset Description
This dataset contains product metadata across 4 categories:
Electronics (500 products)
Video Games (500 products)
Books (500 products)
Home & Kitchen (500 products)
Products were filtered to include only those with 500+ reviews… See the full description on the dataset page: https://huggingface.co/datasets/gatech-scheller-ai-in-business/amazon-products.Product-Descriptions-and-Ads
Synthetic Dataset for Product Descriptions and Ads
The basic process was as follows:
Prompt GPT-4 to create a list of 100 sample clothing items and descriptions for those items.
Split the output into desired format `{"product" : "", "description" : ""}
Prompt GPT-4 to create adverts for each of the 100 samples based on their name and description.
This data was not cleaned or verified manually.
product-management-sft-100k
Product Management SFT (100K)
100,000 ShareGPT conversations demonstrating expert-level product management across PRD writing, feature prioritization, OKR setting, roadmap planning, user research, competitive analysis, and stakeholder communication.
Motivation
AI assistants for product management commonly fail by:
Generic frameworks without application: Explaining RICE scoring without actually scoring the user's features; describing OKRs without writing them… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/product-management-sft-100k.product-catalog-questions
Lamini Product Catalog QA Dataset
Description
This dataset contains questions about products and their corresonding product information like product id, product name, product description, etc. This questions catalog has been built on top of open-source product catalog from kaggle.
Format
The questions and product information are in the form of jsonlines file.
Data Pipeline Code
The entire data pipeline used to create this dataset is open source at:… See the full description on the dataset page: https://huggingface.co/datasets/lamini/product-catalog-questions.task371_synthetic_product_of_list
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task371_synthetic_product_of_list
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task371_synthetic_product_of_list.tanpo-product-sft
Tanpo Product SFT (10k)
Ownership
Owner: DarkNinja Solutions
Creator: d4rkninja
Community: DarkLab
Domain
Product strategy and CEO-level product thinking: roadmap, prioritization, discovery, specs, metrics, and product-market fit.
Dataset summary
Field
Value
Rows
10000
Schema
Chat SFT (messages with system / user / assistant)
Source file
tanpo-product-sft-10k-format-fixed.jsonl
Audit verdict
PASS
Empty assistant… See the full description on the dataset page: https://huggingface.co/datasets/d4rkninja/tanpo-product-sft.ke-products
Dataset Card for Kazanexpress products
Dataset Summary
This dataset was scraped from product pages on the Russian marketplace Kazanexpress. It includes all information from the product card and metadata from the API. The dataset was collected by processing around 3 million products, starting from the first one. At the time the dataset was collected, it is assumed that these were all the products available on this marketplace. Please note that the data returned by the API… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/ke-products.product-authenticity
PRODUCT_AUTHENTICITY
A preference dataset for PRODUCT_AUTHENTICITY, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from
Splits… See the full description on the dataset page: https://huggingface.co/datasets/316usman/product-authenticity.amazon-product-data-sample
Dataset Card for "amazon-product-data-filter"
Dataset Summary
The Amazon Product Dataset contains product listing data from the Amazon US website. It can be used for various NLP and classification tasks, such as text generation, product type classification, attribute extraction, image recognition and more.
NOTICE: This is a sample of the full Amazon Product Dataset, which contains 1K examples. Follow the link to gain access to the full dataset.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/iarbel/amazon-product-data-sample.Product-Descriptions-and-Ads
Synthetic Dataset for Product Descriptions and Ads
The basic process was as follows:
Prompt GPT-4 to create a list of 100 sample clothing items and descriptions for those items.
Split the output into desired format `{"product" : "", "description" : ""}
Prompt GPT-4 to create adverts for each of the 100 samples based on their name and description.
This data was not cleaned or verified manually.
racing-planet-product-catalog
🛒 Racing Planet Simson Product Catalog
37 strukturierte Produkt-Einträge für Simson-Ersatzteile und Tuning-Komponenten von racing-planet.de.
Kategorien
Kategorie
Anzahl
Beispiel
Zündung
5
VAPE Zündanlage 6V/12V, NGK Zündkerze
Motor
6
Stage6 70ccm Zylinder, Kolbenringe, Kurbelwellenlager
Vergaser
4
BVF 16N3-4, Dell'Orto SHA 14/12
Auspuff
4
Sito Plus, Stage6 R/T, Yasuni C16
Elektrik
4
12V Umbausatz, LED Scheinwerfer
Fahrwerk
3
Telegabel… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/racing-planet-product-catalog.turkish-electronics-product-comparison-recommendation
🇹🇷 Türkçe Elektronik Ürün Karşılaştırma ve Öneri Chat Veri Seti
Veri Seti Açıklaması
Bu veri seti, büyük dil modellerinin Türkçe elektronik ürün karşılaştırma ve ürün öneri görevlerinde eğitilmesi amacıyla hazırlanmış bir konuşma veri setidir.
Veri seti iki temel görev grubundan oluşmaktadır:
Elektronik ürün karşılaştırma
Elektronik ürün önerisi
Veri seti Supervised Fine-Tuning (SFT) ve instruction tuning çalışmalarında kullanılabilecek chat formatında… See the full description on the dataset page: https://huggingface.co/datasets/sedayzc/turkish-electronics-product-comparison-recommendation.synthetic-sport-products-sustainability
Dataset Card for synthetic-sport-products-sustainability
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/as-cle-bert/synthetic-sport-products-sustainability/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info… See the full description on the dataset page: https://huggingface.co/datasets/greenfit-ai/synthetic-sport-products-sustainability.enhanced-product-search-llmProduction-grade-clinical-platform-emr-codebase
Clinical EMR Platform Codebase
Production-grade clinical platform codebase available for commercial AI training licensing.
Overview
This dataset contains the full source code of a production electronic medical records (EMR) and clinical management platform, built and deployed over 27 months with real doctors across 40+ medical specialties.
The codebase is available as a non-exclusive commercial license for AI model training, fine-tuning, and benchmarking purposes.… See the full description on the dataset page: https://huggingface.co/datasets/CodeLicense/Production-grade-clinical-platform-emr-codebase.PM-products
Dataset Card for PochtaMarket products
Dataset Summary
This dataset was scraped from product pages on the Russian marketplace PochtaMarket. It includes all information from the product card. The dataset was collected by processing around 500 thousand, starting from the first one. At the time the dataset was collected, it is assumed that these were all the products available on this marketplace. Some fields may be empty, but the string is expected to contain some data… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/PM-products.meat-productsproduct-dev-ops-playbook
Product Dev Ops Playbook — Cross-Functional Sprint Alignment SOP
Complete Product × Engineering × Operations alignment SOP — from unified backlog to 10-day sprint cadence to veto power rules.
📦 Install on ClawHub
clawhub install product-dev-ops-playbook
Then ask your AI agent:
"Design a 10-day sprint cadence for our 8-person product+eng team"
Installs the complete Product × Engineering × Operations alignment SOP — dual-layer Kanban, 10-day sprint cadence… See the full description on the dataset page: https://huggingface.co/datasets/Gingiris/product-dev-ops-playbook.product-hunt-playbook
Product Hunt Launch Playbook
30x Daily #1 | 3x Weekly #1 | 1x Monthly #1
English | 中文 | 日本語 | 한국어
📦 Install on ClawHub
clawhub install product-hunt-playbook
Then ask your AI agent:
"Plan our PH launch for next Tuesday" or "Draft maker comments that don't sound like marketing"
Installs the #1-Day Product Hunt playbook — ranking algorithm explained, hunter outreach scripts, hour-by-hour launch-day timeline, maker comment templates, post-launch momentum… See the full description on the dataset page: https://huggingface.co/datasets/Gingiris/product-hunt-playbook.sea-product-listing-seo-sample
SEA Multilingual Product Listing SEO Sample
This public sample contains 1,000 synthetic, AI-generated marketplace-style
product listing examples for Southeast Asian e-commerce workflows.
Languages
English
Chinese
Malay
Indonesian
Formats
CSV
JSONL
Intended Use
Use this sample for inspection, evaluation, listing-copy prototyping, SEO
keyword experiments, and multilingual catalog workflow testing.
Important Limitations… See the full description on the dataset page: https://huggingface.co/datasets/nwchang/sea-product-listing-seo-sample.amazon-product-data-filter
Dataset Card for "amazon-product-data-filter"
Dataset Summary
The Amazon Product Dataset contains product listing data from the Amazon US website. It can be used for various NLP and classification tasks, such as text generation, product type classification, attribute extraction, image recognition and more.
Languages
The text in the dataset is in English.
Dataset Structure
Data Instances
Each data point provides product information, such… See the full description on the dataset page: https://huggingface.co/datasets/iarbel/amazon-product-data-filter.CoT_Music_Production_DAW
Welcome to CoT_Music_Production_DAW, an open-source dataset (MIT licensed) featuring 7,000 expertly crafted Q&A pairs designed to train AI in the world of electronic music production using FL Studio and Ableton Live. This comprehensive resource covers everything from general DAW fundamentals—like interface navigation, recording, mixing, and project management—to in-depth specifics for FL Studio and Ableton Live, advanced production techniques, and even the nitty-gritty of music theory and… See the full description on the dataset page: https://huggingface.co/datasets/mattwesney/CoT_Music_Production_DAW.
