datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.customer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits… See the full description on the dataset page: https://huggingface.co/datasets/Tobi-Bueck/customer-support-tickets.telco-customer-churn
Dataset Card for Telco Customer Churn
This dataset contains information about customers of a fictional telecommunications company, including demographic information, services subscribed to, location details, and churn behavior. This merged dataset combines the information from the original Telco Customer Churn dataset with additional details.
Dataset Details
Dataset Description
This merged Telco Customer Churn dataset provides a comprehensive view of customer… See the full description on the dataset page: https://huggingface.co/datasets/aai510-group1/telco-customer-churn.Customer_support_faqs_datasetDataset Name: Customer Support FAQs Dataset
Description:
This dataset contains a collection of 200 frequently asked questions (FAQs) and their corresponding answers, designed to assist in customer support scenarios. The questions cover a wide range of common customer inquiries related to account management, payment methods, order tracking, shipping, returns, and more. This dataset is intended for use in developing and training AI models for customer support chatbots, automated response systems… See the full description on the dataset page: https://huggingface.co/datasets/MakTek/Customer_support_faqs_dataset.Customer-Support-Responsescustomer-support-on-twitter-conversationCustomer_Support_on_Twitterwayfair_customer_reviews
Wayfair Customer Reviews Dataset
This dataset contains customer reviews collected from wayfair.com.It accompanies the paper End-to-End Aspect-Guided Review Summarization at Scale, accepted to the EMNLP 2025 Industry Track.
Overview
The dataset supports tasks such as:
Aspect extraction
Product-level summarization, e.g., aggregating reviews by the product_id field.
It can be used on its own or combined with the companion Wayfair Product Summaries dataset.Both datasets… See the full description on the dataset page: https://huggingface.co/datasets/IeBoytsov/wayfair_customer_reviews.telco-customer-churn3mo-customer-support-tweets-945k
Customer Support on Twitter Dataset 945k
Dataset Description
Context
This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions.
Content
Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/MohammadOthman/mo-customer-support-tweets-945k.telco_customer_churnwellness-tourism-customerswayfair_customer_reviews
Wayfair Customer Reviews Dataset
This dataset contains customer reviews collected from wayfair.com.It accompanies the paper End-to-End Aspect-Guided Review Summarization at Scale, accepted to the EMNLP 2025 Industry Track.
Overview
The dataset supports tasks such as:
Aspect extraction
Product-level summarization, e.g., aggregating reviews by the product_id field.
It can be used on its own or combined with the companion Wayfair Product Summaries dataset.Both datasets… See the full description on the dataset page: https://huggingface.co/datasets/al5nfsharyh/wayfair_customer_reviews.E-Commerce_Customer_Support_Conversations
Dataset Card for "E-Commerce_Customer_Support_Conversations"
The dataset is synthetically generated with OpenAI ChatGPT model (gpt-3.5-turbo).
More Information needed
customer-support-client-agent-conversations
Customer Support Client-Agent Conversations Dataset
A synthetic context-summarized multi-turn customer-service question-answering dataset for banking domain conversations, designed for training and evaluating small language models on dialogue continuity and contextual understanding tasks.
Dataset Description
This dataset contains 183,337 context-summarized multi-turn customer-service conversations spanning various banking scenarios including account management… See the full description on the dataset page: https://huggingface.co/datasets/Lakshan2003/customer-support-client-agent-conversations.e-commerce-customer-support-qa
Dataset Card for Dataset Name
from: NebulaByte/E-Commerce_Customer_Support_Conversations
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/rjac/e-commerce-customer-support-qa.SPADE-customer-service-dialogue
SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection
Paper | Code
SPADE contains a repository of customer service line synthetic user dialogues with goals, augmented from MultiWOZ 2.1 using GPT-3.5 and Llama 70B.
The datasets are intended for training and evaluating machine generated text detectors in dialogue settings.
There are 15 English datasets generated using 5 different augmentation methods and 2 large language models… See the full description on the dataset page: https://huggingface.co/datasets/AngieYYF/SPADE-customer-service-dialogue.brazilian-customer-service-conversations
Brazilian Customer Service Conversations
Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR).
De um like me apoie em manter esse dataset!
Descricao
Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de:
Chatbots de atendimento
Classificacao de intencao (intent classification)
Analise de sentimento em conversas
Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.telco-customer-churncustomer_support_ticketscustomers-complaints
Dataset Card for "financial-customer-complaints-v5"
More Information needed
customerHARLEY_DAVIDSON_CUSTOMER_FUNDING_CORP_1114926
HARLEY-DAVIDSON CUSTOMER FUNDING CORP.
SEC ABS-EE asset-level filings for CIK 1114926 (HARLEY-DAVIDSON CUSTOMER FUNDING CORP.).
Filings: 383
Parquet files: 5
Total size: 8.0 MB
Reporting period start: 2019-05-31
Reporting period end: 2026-02-28
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/HARLEY_DAVIDSON_CUSTOMER_FUNDING_CORP_1114926.ecommerce-customer-support-conversationsE-Commerce Customer Support Conversations
Dataset Summary:
This dataset contains customer support queries and responses from an e-commerce context.
It is designed for training and fine-tuning AI models for automated customer service, chatbots, and natural language processing (NLP) applications.
Use Cases:
Fine-tuning conversational AI models (e.g., GPT, BERT)
Training chatbots for e-commerce support
Improving customer service automation
Sentiment and intent analysis
Dataset Format:
The… See the full description on the dataset page: https://huggingface.co/datasets/Venkatrajan247/ecommerce-customer-support-conversations.tts-eval-customer-support-202608
TTS Hard Cases — Customer Support
A compact, text-only evaluation set of hard cases for text-to-speech models, focused on the
customer support use case. Published by Gradium.
Most TTS benchmarks measure naturalness on ordinary prose, where modern models are already close to
saturated. Production voice agents fail somewhere else: on the literal payload of a support call —
the order number, the email address, the spelled-out surname, the date, the amount refunded. A voice
that… See the full description on the dataset page: https://huggingface.co/datasets/gradium/tts-eval-customer-support-202608.autotrain-data-quality-customer-reviews
AutoTrain Dataset for project: quality-customer-reviews
Dataset Descritpion
This dataset has been automatically processed by AutoTrain for project quality-customer-reviews.
Languages
The BCP-47 code for the dataset's language is en.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": " Love this truck, I think it is light years better than the competition. I have driven or… See the full description on the dataset page: https://huggingface.co/datasets/Gflorent/autotrain-data-quality-customer-reviews.customer_service_client_agent_conversations_40k_multi_task
Dataset Card for "customer_service_client_agent_conversations_40k_multi_task"
More Information needed
customer_assistant
Dataset Card for customer_assistant
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
This dataset contains:
A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/argilla/customer_assistant.customer-service
Customer Service Conversations Dataset
This dataset contains 100 realistic customer service conversations between customers and support agents. Each dialogue is 11 turns long and covers a variety of common issues such as late deliveries, billing errors, account problems, and more. It is ideal for training and evaluating AI assistants, chatbots, and customer support models.
Dataset Structure
Each conversation is stored as a JSON object with the following fields:… See the full description on the dataset page: https://huggingface.co/datasets/chenlei123/customer-service.customer-support-twitter-dataset
