datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Customer_support_faqs_datasetDataset Name: Customer Support FAQs Dataset
Description:
This dataset contains a collection of 200 frequently asked questions (FAQs) and their corresponding answers, designed to assist in customer support scenarios. The questions cover a wide range of common customer inquiries related to account management, payment methods, order tracking, shipping, returns, and more. This dataset is intended for use in developing and training AI models for customer support chatbots, automated response systems… See the full description on the dataset page: https://huggingface.co/datasets/MakTek/Customer_support_faqs_dataset.mo-customer-support-tweets-945k
Customer Support on Twitter Dataset 945k
Dataset Description
Context
This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions.
Content
Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/MohammadOthman/mo-customer-support-tweets-945k.ecommerce-customer-support-conversationsE-Commerce Customer Support Conversations
Dataset Summary:
This dataset contains customer support queries and responses from an e-commerce context.
It is designed for training and fine-tuning AI models for automated customer service, chatbots, and natural language processing (NLP) applications.
Use Cases:
Fine-tuning conversational AI models (e.g., GPT, BERT)
Training chatbots for e-commerce support
Improving customer service automation
Sentiment and intent analysis
Dataset Format:
The… See the full description on the dataset page: https://huggingface.co/datasets/Venkatrajan247/ecommerce-customer-support-conversations.Bitext-customer-support-llm-chatbot-training-dataset-spanish
Spanish Customer Support LLM Chatbot Training Dataset
Spanish-language adaptation of the Bitext Customer Support LLM Chatbot Training Dataset.
This dataset is intended for training and evaluating Spanish-language customer-support chatbots and instruction-following large language models.
Dataset Details
Dataset Description
This dataset is a Spanish translation and adaptation of the original Bitext Customer Support LLM Chatbot Training Dataset.
The… See the full description on the dataset page: https://huggingface.co/datasets/Faramir/Bitext-customer-support-llm-chatbot-training-dataset-spanish.CUSTOMER_SUPPORT_DATASET_JSONL_V1VNOVA AI — Customer Support Dataset (100 Synthetic Scenarios)
High-quality synthetic customer support conversations designed for training and evaluating AI support agents.
About the Dataset
This dataset contains 100 fully synthetic customer–agent interactions across multiple industries, created to help developers train:
1-Customer support chatbots
2-Automated ticket resolution systems
3-LLM fine-tuning for support tasks
4-RAG workflows
5-Complaint classification models
6-Sentiment & intent… See the full description on the dataset page: https://huggingface.co/datasets/vnovaai/CUSTOMER_SUPPORT_DATASET_JSONL_V1.amazon-customer-support
Amazon Customer Support
Derived from the TWCS corpus (Kaggle: thoughtvector/customer-support-on-twitter), this dataset
contains 200 labelled customer-support interactions for evaluation / fine-tuning purposes.
Schema
Each line is a JSON object with four fields:
Field
Type
Description
query
string
The customer's raw message (input)
action
string
Agent action taken — resolve or escalated_to_human
intent
string
Classified intent — complaint, question… See the full description on the dataset page: https://huggingface.co/datasets/iam-tsr/amazon-customer-support.Bitext-customer-support-1-columncustomer-support-dpo-100k
Customer Support DPO 100K
A synthetic Direct Preference Optimization (DPO) dataset of 100,000 customer support interactions with chosen (high-quality) and rejected (poor-quality) response pairs. Designed to train AI models to provide genuinely helpful, specific, and empathetic customer support.
Dataset Description
This dataset covers 23 real-world customer support scenarios across B2B and B2C contexts. Each record includes a customer message, a high-quality chosen… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/customer-support-dpo-100k.customer_support_cmd_gen_basicCustomer_support_faqs_datasetDataset Name: Customer Support FAQs Dataset
Description:
This dataset contains a collection of 200 frequently asked questions (FAQs) and their corresponding answers, designed to assist in customer support scenarios. The questions cover a wide range of common customer inquiries related to account management, payment methods, order tracking, shipping, returns, and more. This dataset is intended for use in developing and training AI models for customer support chatbots, automated response systems… See the full description on the dataset page: https://huggingface.co/datasets/Yatzo/Customer_support_faqs_dataset.sea-ecommerce-customer-support-sample
SEA Multilingual E-commerce Customer Support Sample
This public sample contains 1,000 synthetic, AI-generated customer-support
conversations for Southeast Asian e-commerce scenarios.
Languages
English
Chinese
Malay
Indonesian
Formats
CSV
JSONL
Intended Use
Use this sample for inspection, evaluation, prototyping, multilingual testing,
and intent-classification experiments.
Important Limitations
This is synthetic… See the full description on the dataset page: https://huggingface.co/datasets/nwchang/sea-ecommerce-customer-support-sample.kazakh-customer-support-qa
Dataset Card for kazakh-customer-support-qa
Maintained by: Mäñgi ÜTM (mangi-llm)
Dataset Summary
kazakh-customer-support-qa is a small, hand-curated question–answer dataset in the Kazakh language, built to represent realistic customer-support conversations across several industries (banking, telecom, retail/service centers, sales, and general support). Each record pairs a short customer question with a concise, policy-safe answer, and many answers include… See the full description on the dataset page: https://huggingface.co/datasets/mangi-llm/kazakh-customer-support-qa.agentic-customer-support-india
Agentic Customer Support Dataset (India E-Commerce)
An annotated customer-support query dataset for the Indian e-commerce domain, built to
train and evaluate the Agentic Customer Support System (CPG 289, Thapar Institute of
Engineering and Technology) — a multi-agent, hybrid-retrieval support architecture with a
LoRA fine-tuned Llama 3 8B language model.
Dataset Summary
150 customer support queries, each labelled with intent (6 categories), sentiment
(3 classes)… See the full description on the dataset page: https://huggingface.co/datasets/Aditvir/agentic-customer-support-india.customer-support-registrysynth-customer-support-expanded-R
Expanded E-Commerce & Subscription Customer Support Dataset
Dataset Summary
This high-quality synthetic dataset contains 438 realistic customer support interactions focused on e-commerce, shipping, delivery, and subscription management. It was created to provide edge-case scenarios and varied support policies (like Hazmat battery returns, subscription cancellations, tracking loops, and incorrect SKU deliveries).
This dataset is ideal for Supervised Fine-Tuning (SFT) or… See the full description on the dataset page: https://huggingface.co/datasets/KazKozDev/synth-customer-support-expanded-R.customer_supportmo-customer-support-tweets-945k
Customer Support on Twitter Dataset 945k
Dataset Description
Context
This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions.
Content
Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/shaikayub/mo-customer-support-tweets-945k.bank_customer_supportaz_customer_support-v1.0
💬 az_customer_support-v1.0
Description:6,000 multilingual (Azerbaijani–English) instruction–response examples for common customer support scenarios.Covers diverse intents such as checking invoice status, canceling or modifying orders, tracking shipments, and more.Each entry includes the user's query (in English and Azerbaijani), the categorized intent, and a well-structured support response tailored to the request.
Use Cases:
Instruction-tuning for multilingual customer service… See the full description on the dataset page: https://huggingface.co/datasets/az-llm/az_customer_support-v1.0.ecommerce-customer-support
E-commerce Customer Support Dataset 🛒
This dataset contains high-quality e-commerce customer support queries, intents, and responses designed for fine-tuning AI models and customer service chatbots.
📌 Dataset Overview
Domain: E-commerce / Online Retail
Task: Customer Support Automation & Intent Classification
Format: JSON / JSONL
🚀 Features
Cleaned and structured data
Covers order tracking, returns, payment queries, and general support
Ready… See the full description on the dataset page: https://huggingface.co/datasets/muhammadumar009/ecommerce-customer-support.customer-support-cleaned
Customer Support Cleaned
Dataset Summary
customer-support-cleaned is a small, curated multilingual customer-support
conversation dataset derived from a raw Excel workbook (file1.xlsx). Each
record contains a customer message (user_message), the support agent's reply
(agent_response), the conversation language (normalized to ISO 639-1), and a
sentiment label (positive, neutral, or negative).
The dataset is intended for tasks such as:
Multilingual intent/sentiment… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/customer-support-cleaned.Customer-Support-Chatcustomer-support-and-servicecustomerSupport
customerSupport
Dataset Description
This is a synthetic dataset generated using the yaLLMa3 pipeline for triplets tasks in Arabic.
Dataset Summary
Domain: customer support
Data Type: triplets
Language: Arabic (ar)
Total Rows: 30
Generated: 10/12/2025
Generation Statistics
Topics Processed: 6
Success Rate: 100.0%
Generation Time: 47.63s
Rows Per Topic: 5
Dataset Structure
Data Fields
anchor: string
positive: string
negative:… See the full description on the dataset page: https://huggingface.co/datasets/omar-emad/customerSupport.mo-customer-support-tweets-945k
Customer Support on Twitter Dataset 945k
Dataset Description
Context
This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions.
Content
Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/Prady06/mo-customer-support-tweets-945k.telecom-customer-support-synthetic-replicas
Customer Support Differentially Private Synthetic Conversations Dataset
This dataset contains pairs of customer support conversations: original conversations and their synthetic counterparts generated with differential privacy (DP) guarantees (ε=8.0, δ=1e-5). The conversations cover technical support topics related to performance and speed concerns in mobile devices.
Dataset Description
Overview
The Customer Support Differentially Private Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Ming-secludy/telecom-customer-support-synthetic-replicas.customer-support-synth-quickdemo
expanded_E-commerce_and_SaaS_customer_support_(orders,_shipping,_billing,_subscriptions,_app_troubleshooting)
Synthetic customer-support dataset generated with synth-dataset-kit.
What Is Included
train.jsonl: training split in JSONL chat format
eval_summary.json: minimal evaluation and generation summary
expanded_E-commerce_and_SaaS_customer_support_(orders,_shipping,_billing,_subscriptions,_app_troubleshooting)_quality_report.html: visual quality report… See the full description on the dataset page: https://huggingface.co/datasets/Enkiboy/customer-support-synth-quickdemo.sea-ecommerce-customer-support-sample
SEA Multilingual E-commerce Customer Support Sample
This public sample contains 1,000 synthetic, AI-generated customer-support
conversations for Southeast Asian e-commerce scenarios.
Languages
English
Chinese
Malay
Indonesian
Formats
CSV
JSONL
Intended Use
Use this sample for inspection, evaluation, prototyping, multilingual testing,
and intent-classification experiments.
Important Limitations
This is synthetic… See the full description on the dataset page: https://huggingface.co/datasets/dinesshxryu/sea-ecommerce-customer-support-sample.Customer_support_faqs_datasetDataset Name: Customer Support FAQs Dataset
Description:
This dataset contains a collection of 200 frequently asked questions (FAQs) and their corresponding answers, designed to assist in customer support scenarios. The questions cover a wide range of common customer inquiries related to account management, payment methods, order tracking, shipping, returns, and more. This dataset is intended for use in developing and training AI models for customer support chatbots, automated response systems… See the full description on the dataset page: https://huggingface.co/datasets/Walaaaaaaaa/Customer_support_faqs_dataset.Customer_support_faqs_datasetDataset Name: Customer Support FAQs Dataset
Description:
This dataset contains a collection of 200 frequently asked questions (FAQs) and their corresponding answers, designed to assist in customer support scenarios. The questions cover a wide range of common customer inquiries related to account management, payment methods, order tracking, shipping, returns, and more. This dataset is intended for use in developing and training AI models for customer support chatbots, automated response systems… See the full description on the dataset page: https://huggingface.co/datasets/molotsikano/Customer_support_faqs_dataset.
