datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SPADE-customer-service-dialogue
SPADE: Structured Prompting Augmentation for Dialogue Enhancement in Machine-Generated Text Detection
Paper | Code
SPADE contains a repository of customer service line synthetic user dialogues with goals, augmented from MultiWOZ 2.1 using GPT-3.5 and Llama 70B.
The datasets are intended for training and evaluating machine generated text detectors in dialogue settings.
There are 15 English datasets generated using 5 different augmentation methods and 2 large language models… See the full description on the dataset page: https://huggingface.co/datasets/AngieYYF/SPADE-customer-service-dialogue.brazilian-customer-service-conversations
Brazilian Customer Service Conversations
Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR).
De um like me apoie em manter esse dataset!
Descricao
Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de:
Chatbots de atendimento
Classificacao de intencao (intent classification)
Analise de sentimento em conversas
Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.customer-service-sft-50k
Customer Service SFT (50K)
50,000 ShareGPT-format customer service conversations across 8 industries and 18 issue types. Each conversation includes a system prompt establishing the agent's role, authority limits, and policy constraints — training models to operate within defined boundaries while resolving issues empathetically and effectively.
Motivation
Customer service is one of the highest-volume LLM deployment contexts. Models need to balance:
Empathy with… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/customer-service-sft-50k.algerian-darija-customer-service-sample
Algerian Darija customer messages — stratified sample
500 spontaneous Algerian Darija messages, written by real customers, drawn from a
first-party corpus of 869,166 customer messages. Every message here is unique
after normalization, de-identified, and typed by a human — nothing elicited, translated, scraped or
generated.
Algerian Darija (ISO 639-3 arq) is spoken by around 45 million people and is one of the worst-covered
varieties in current language models. For scale: PADIC… See the full description on the dataset page: https://huggingface.co/datasets/dzcorpora/algerian-darija-customer-service-sample.Frames-synthetic-customer-service-dialogue
Frames Synthetic Customer Service Dialogues
This contains a repository of customer service line synthetic user dialogues with goals, augmented from Frames using Qwen2.5-32B.
The datasets are intended for training and evaluating machine generated text detectors in dialogue settings.
Dataset Structure
The datasets are of parquet file format and contain the following columns:
Column
Description
dia_no
Unique ID for each dialogue. Dialogues with the same ID… See the full description on the dataset page: https://huggingface.co/datasets/AngieYYF/Frames-synthetic-customer-service-dialogue.synthetic-ecommerce-customer-service-dialogues-100k
Synthetic E-commerce Support Dialogues
Single-turn English support examples with privacy-aware responses.
Dataset summary
Rows: 100,000
Columns: 11
Train/validation/test: 80,000 / 10,000 / 10,000
Synthetic: yes, every row
Generation seed: 550031
Intended uses
Model prototyping, pipeline testing, schema experiments, and educational demonstrations.
Limitations
This dataset is synthetic and must not be represented as observed platform… See the full description on the dataset page: https://huggingface.co/datasets/synthdataq9x260918/synthetic-ecommerce-customer-service-dialogues-100k.CallAgentAI-Hinglish-Customer-Service
CallAgent AI: Hinglish Business Conversations Dataset
This dataset contains synthetic, high-quality "Hinglish" (Hindi + English code-switching) customer service interactions. It was generated by CallAgent AI (callagentai.in) — India's leading AI voice receptionist platform designed specifically for Indian SMBs.
Why this dataset exists
Global voice AI models often fail to capture the unique nuances of Indian business calls, which heavily rely on fluid language… See the full description on the dataset page: https://huggingface.co/datasets/Ghanashyaam/CallAgentAI-Hinglish-Customer-Service.customer_service_chatbotCustomer-service-tickets-qwen-qa
Customer Support Tickets QA (English) — Qwen SFT Dataset
This dataset is formatted for supervised fine-tuning (SFT) of Qwen-style chat models on customer support email tasks. source dataset: Tobi-Bueck/customer-support-tickets
It is designed for training models to read a customer ticket, understand its context, and generate an appropriate support response. Depending on the prompt design, the same data can also support auxiliary tasks such as queue prediction, priority prediction… See the full description on the dataset page: https://huggingface.co/datasets/W-L/Customer-service-tickets-qwen-qa.agentic_synthetic_customer_service_conversations
LLM-filtered Customer Service Conversations Dataset
Overview
This dataset contains simulated conversations generated by our agentic simulation system.
The conversations are filtered by a LLM to ensure they are of high quality.
Each record is stored in JSON Lines (JSONL) format and includes:
Input Settings: Metadata such as selected bank, customer, agent profiles, and task details.
Messages: The full conversation messages.
Summary: A German summary of the conversation.… See the full description on the dataset page: https://huggingface.co/datasets/marccgrau/agentic_synthetic_customer_service_conversations.spotless-customer-service-training
Spotless Bin Co Customer Service Training Data
Training data for a customer service AI model for Spotless Bin Co, a residential trash can cleaning service.
Dataset Description
This dataset contains 8,776 conversational examples across 5 categories:
Category
Count
Description
FAQs
1,951
Frequently asked questions
Service
1,925
Service explanation dialogues
Objections
1,925
Objection handling examples
Booking
1,975
Booking flow conversations
Brand
1,000… See the full description on the dataset page: https://huggingface.co/datasets/rileyseaburg/spotless-customer-service-training.llm_filtered_customer_service_conversations
LLM-filtered Customer Service Conversations Dataset
Overview
This dataset contains simulated conversations generated by our agentic simulation system.
The conversations are filtered by a LLM to ensure they are of high quality.
Each record is stored in JSON Lines (JSONL) format and includes:
Input Settings: Metadata such as selected bank, customer, agent profiles, and task details.
Messages: The full conversation messages.
Summary: A German summary of the conversation.… See the full description on the dataset page: https://huggingface.co/datasets/marccgrau/llm_filtered_customer_service_conversations.bank_customer_service
Bank Customer Service (UBA) Instruction Dataset
Overview
This dataset contains 254 instruction-style customer service interactions modeled on real support scenarios for United Bank for Africa (UBA), a major Nigerian bank. Each example pairs a customer question or complaint (e.g. transaction disputes, transfer limits, NUBAN, mobile/internet banking, stamp duty charges) with a concise, factual agent response. It was built to support fine-tuning of conversational AI… See the full description on the dataset page: https://huggingface.co/datasets/Ephraimmm/bank_customer_service.customer_service_dataset
GTBank Customer Service – Synthetic Dataset
Overview
This dataset contains 500 synthetic user_query / assistant_reply pairs modeled on common GTBank (Guaranty Trust Bank) retail banking customer support interactions, such as password resets, debit card activation, failed transactions, and KYC/transfer limits. It is entirely synthetic — no real customer data, transcripts, or GTBank systems were used — and is intended for prototyping and testing customer-support NLP… See the full description on the dataset page: https://huggingface.co/datasets/Ephraimmm/customer_service_dataset.Multilingual-Nepali-Customer-Care-Services-Datasetllm_filtered_customer_service_conversations_cleaned
LLM-filtered Customer Service Conversations Dataset (cleaned)
Overview
This dataset contains simulated conversations generated by our agentic simulation system.
The conversations are filtered by a LLM to ensure they are of high quality.
Each record is stored in JSON Lines (JSONL) format and includes:
Input Settings: Metadata such as selected bank, customer, agent profiles, and task details.
Messages: The full conversation messages.
Summary: A German summary of the… See the full description on the dataset page: https://huggingface.co/datasets/marccgrau/llm_filtered_customer_service_conversations_cleaned.Customer-Care-Services-Dataset-in-Nepali
