datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mo-customer-support-tweets-945k
Customer Support on Twitter Dataset 945k
Dataset Description
Context
This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions.
Content
Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/MohammadOthman/mo-customer-support-tweets-945k.customer-support-th-26.9k
customer-support-th-26.9k
Thai customer-support instruction dataset (~26.9k examples). Thai-localized version of the Bitext customer-support dataset — instruction templates, intent/category labels, and response templates in Thai.
Format
Field
Description
instruction
Customer question template in Thai (may contain {{placeholders}})
response
Support response template in Thai
category
Coarse category (e.g. ORDER)
intent
Fine-grained intent (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/customer-support-th-26.9k.customer-support
Description
Topic: Customer Support Interactions
Domains: E-commerce, Telecommunications, Software Services
Number of Entries: 1,000
Dataset Type: Raw Dataset
Model Used: Meta Llama4 Maverick 17B Instruct V1
Language: English
customer-support-chatml
Customer Support ChatML Dataset
This dataset is a curated and preprocessed version of the
Bitext Customer Support Dataset.
Dataset Description
The dataset has been converted to ChatML format for fine-tuning conversational AI models.
Format
Each example contains:
text: The complete conversation in ChatML format
messages: JSON string of the conversation as a list of messages
instruction: The original user query
response: The original assistant response… See the full description on the dataset page: https://huggingface.co/datasets/Shivam271089/customer-support-chatml.synthetic-customer-support-sft-smoke
Synthetic customer-support SFT dataset
Synthetically generated with HuggingFaceTB/SmolLM2-135M-Instruct from seeded scenario prompts (18 products x 16 issue types x 5 customer tones).
Rows: 4 kept after validation (4 completions failed parsing and were dropped)
Format: messages column (system/user/assistant), ready for TRL SFTTrainer
Metadata: category (issue type), tone (customer tone)
Seed: 42
Generated on 2026-09-02. Model-generated content: review before production use.
sea-ecommerce-customer-support-sample
SEA Multilingual E-commerce Customer Support Sample
This public sample contains 1,000 synthetic, AI-generated customer-support
conversations for Southeast Asian e-commerce scenarios.
Languages
English
Chinese
Malay
Indonesian
Formats
CSV
JSONL
Intended Use
Use this sample for inspection, evaluation, prototyping, multilingual testing,
and intent-classification experiments.
Important Limitations
This is synthetic… See the full description on the dataset page: https://huggingface.co/datasets/nwchang/sea-ecommerce-customer-support-sample.customer_support_auto_completionkazakh-customer-support-qa
Dataset Card for kazakh-customer-support-qa
Maintained by: Mäñgi ÜTM (mangi-llm)
Dataset Summary
kazakh-customer-support-qa is a small, hand-curated question–answer dataset in the Kazakh language, built to represent realistic customer-support conversations across several industries (banking, telecom, retail/service centers, sales, and general support). Each record pairs a short customer question with a concise, policy-safe answer, and many answers include… See the full description on the dataset page: https://huggingface.co/datasets/mangi-llm/kazakh-customer-support-qa.customer-support-dpo-100k
Customer Support DPO 100K
A synthetic Direct Preference Optimization (DPO) dataset of 100,000 customer support interactions with chosen (high-quality) and rejected (poor-quality) response pairs. Designed to train AI models to provide genuinely helpful, specific, and empathetic customer support.
Dataset Description
This dataset covers 23 real-world customer support scenarios across B2B and B2C contexts. Each record includes a customer message, a high-quality chosen… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/customer-support-dpo-100k.uzbek-customer-support-dialogs
Uzbek Customer Support Dialogs 🇺🇿
A high-quality dataset of 990 customer support conversations in Uzbek (Latin script), designed for training and fine-tuning conversational AI models.
This is one of the first large-scale customer support datasets in Uzbek, created to address the gap of low-resource NLP for Central Asian languages.
📋 Dataset Description
990 conversational dialogs in natural Uzbek (Latin script)
11 customer support categories: Order, Shipping, Cancel… See the full description on the dataset page: https://huggingface.co/datasets/AsrorAsr/uzbek-customer-support-dialogs.synth-customer-support-expanded-R
Expanded E-Commerce & Subscription Customer Support Dataset
Dataset Summary
This high-quality synthetic dataset contains 438 realistic customer support interactions focused on e-commerce, shipping, delivery, and subscription management. It was created to provide edge-case scenarios and varied support policies (like Hazmat battery returns, subscription cancellations, tracking loops, and incorrect SKU deliveries).
This dataset is ideal for Supervised Fine-Tuning (SFT) or… See the full description on the dataset page: https://huggingface.co/datasets/KazKozDev/synth-customer-support-expanded-R.TR-ECommerce-CustomerSupport-Instructions
TR E-Commerce Customer Support Instructions 🇹🇷
A high-quality Turkish E-Commerce Customer Support dataset designed for fine-tuning large language models (LLMs) on instruction-following customer support tasks.
Dataset Summary
Feature
Value
Language
Turkish (tr)
Domain
E-Commerce Customer Support
Format
Conversation (Conversational/Chat)
Chain-of-Thought (CoT)
✅ Natural paragraph reasoning (thinking field)
Total Categories
20
Total Rows
186… See the full description on the dataset page: https://huggingface.co/datasets/Mer1Alii/TR-ECommerce-CustomerSupport-Instructions.agentic-customer-support-india
Agentic Customer Support Dataset (India E-Commerce)
An annotated customer-support query dataset for the Indian e-commerce domain, built to
train and evaluate the Agentic Customer Support System (CPG 289, Thapar Institute of
Engineering and Technology) — a multi-agent, hybrid-retrieval support architecture with a
LoRA fine-tuned Llama 3 8B language model.
Dataset Summary
150 customer support queries, each labelled with intent (6 categories), sentiment
(3 classes)… See the full description on the dataset page: https://huggingface.co/datasets/Aditvir/agentic-customer-support-india.customer-support-synth-quickdemo
expanded_E-commerce_and_SaaS_customer_support_(orders,_shipping,_billing,_subscriptions,_app_troubleshooting)
Synthetic customer-support dataset generated with synth-dataset-kit.
What Is Included
train.jsonl: training split in JSONL chat format
eval_summary.json: minimal evaluation and generation summary
expanded_E-commerce_and_SaaS_customer_support_(orders,_shipping,_billing,_subscriptions,_app_troubleshooting)_quality_report.html: visual quality report… See the full description on the dataset page: https://huggingface.co/datasets/Enkiboy/customer-support-synth-quickdemo.mo-customer-support-tweets-945k
Customer Support on Twitter Dataset 945k
Dataset Description
Context
This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions.
Content
Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/shaikayub/mo-customer-support-tweets-945k.telecom-customer-support-synthetic-replicas
Customer Support Differentially Private Synthetic Conversations Dataset
This dataset contains pairs of customer support conversations: original conversations and their synthetic counterparts generated with differential privacy (DP) guarantees (ε=8.0, δ=1e-5). The conversations cover technical support topics related to performance and speed concerns in mobile devices.
Dataset Description
Overview
The Customer Support Differentially Private Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Ming-secludy/telecom-customer-support-synthetic-replicas.mo-customer-support-tweets-945k
Customer Support on Twitter Dataset 945k
Dataset Description
Context
This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions.
Content
Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/Prady06/mo-customer-support-tweets-945k.sea-ecommerce-customer-support-sample
SEA Multilingual E-commerce Customer Support Sample
This public sample contains 1,000 synthetic, AI-generated customer-support
conversations for Southeast Asian e-commerce scenarios.
Languages
English
Chinese
Malay
Indonesian
Formats
CSV
JSONL
Intended Use
Use this sample for inspection, evaluation, prototyping, multilingual testing,
and intent-classification experiments.
Important Limitations
This is synthetic… See the full description on the dataset page: https://huggingface.co/datasets/dinesshxryu/sea-ecommerce-customer-support-sample.
