datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
brazilian-customer-service-conversations
Brazilian Customer Service Conversations
Dataset de conversas de atendimento ao cliente em portugues brasileiro (PT-BR).
De um like me apoie em manter esse dataset!
Descricao
Conversas sinteticas de alta qualidade simulando interacoes reais entre clientes e atendentes em diversos setores da economia brasileira. Util para treinar e avaliar modelos de:
Chatbots de atendimento
Classificacao de intencao (intent classification)
Analise de sentimento em conversas
Geracao de… See the full description on the dataset page: https://huggingface.co/datasets/RichardSakaguchiMS/brazilian-customer-service-conversations.customer-service
Customer Service Conversations Dataset
This dataset contains 100 realistic customer service conversations between customers and support agents. Each dialogue is 11 turns long and covers a variety of common issues such as late deliveries, billing errors, account problems, and more. It is ideal for training and evaluating AI assistants, chatbots, and customer support models.
Dataset Structure
Each conversation is stored as a JSON object with the following fields:… See the full description on the dataset page: https://huggingface.co/datasets/chenlei123/customer-service.customer-service-sft-50k
Customer Service SFT (50K)
50,000 ShareGPT-format customer service conversations across 8 industries and 18 issue types. Each conversation includes a system prompt establishing the agent's role, authority limits, and policy constraints — training models to operate within defined boundaries while resolving issues empathetically and effectively.
Motivation
Customer service is one of the highest-volume LLM deployment contexts. Models need to balance:
Empathy with… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/customer-service-sft-50k.algerian-darija-customer-service-sample
Algerian Darija customer messages — stratified sample
500 spontaneous Algerian Darija messages, written by real customers, drawn from a
first-party corpus of 869,166 customer messages. Every message here is unique
after normalization, de-identified, and typed by a human — nothing elicited, translated, scraped or
generated.
Algerian Darija (ISO 639-3 arq) is spoken by around 45 million people and is one of the worst-covered
varieties in current language models. For scale: PADIC… See the full description on the dataset page: https://huggingface.co/datasets/dzcorpora/algerian-darija-customer-service-sample.CallAgentAI-Hinglish-Customer-Service
CallAgent AI: Hinglish Business Conversations Dataset
This dataset contains synthetic, high-quality "Hinglish" (Hindi + English code-switching) customer service interactions. It was generated by CallAgent AI (callagentai.in) — India's leading AI voice receptionist platform designed specifically for Indian SMBs.
Why this dataset exists
Global voice AI models often fail to capture the unique nuances of Indian business calls, which heavily rely on fluid language… See the full description on the dataset page: https://huggingface.co/datasets/Ghanashyaam/CallAgentAI-Hinglish-Customer-Service.customer-service
Customer Service Conversations Dataset
This dataset contains 100 realistic customer service conversations between customers and support agents. Each dialogue is 11 turns long and covers a variety of common issues such as late deliveries, billing errors, account problems, and more. It is ideal for training and evaluating AI assistants, chatbots, and customer support models.
Dataset Structure
Each conversation is stored as a JSON object with the following fields:
id:… See the full description on the dataset page: https://huggingface.co/datasets/ai-training-datasets/customer-service.CustomerService
数据集说明
组成
类型
文件夹名称
来源
数量
说明
电信问答
telecom_Q&A
百度知道QA
87366
经过脱敏、数据清洗、人工筛选等处理
行业相关知识数据
industry_data
教科书、国际标准等
5218
通过大模型从文档得到的QA数据,部分原文档保存在source_data中
通用指令数据集
general_instruction
firefly
18123
挑选了阅读、情感理解、补全、逻辑推理等主题的通用指令
混合数据集
blended_data
-
-
按照数据集建设进程,混合后组件的训练、测试数据,可直接使用
混合数据 - V1
组成
来源
比例
条数
说明
百度知道
64%
32282
经过脱敏、数据清洗、人工筛选等处理
firefly
36%
18123
挑选了阅读、情感理解、补全、逻辑推理等主题的通用指令
标准问答
-
18
通过联通网上营业厅在线客服整理
合计
100%
50423
-… See the full description on the dataset page: https://huggingface.co/datasets/THU-StarLab/CustomerService.Customer-service-tickets-qwen-qa
Customer Support Tickets QA (English) — Qwen SFT Dataset
This dataset is formatted for supervised fine-tuning (SFT) of Qwen-style chat models on customer support email tasks. source dataset: Tobi-Bueck/customer-support-tickets
It is designed for training models to read a customer ticket, understand its context, and generate an appropriate support response. Depending on the prompt design, the same data can also support auxiliary tasks such as queue prediction, priority prediction… See the full description on the dataset page: https://huggingface.co/datasets/W-L/Customer-service-tickets-qwen-qa.rooyai-customer-service-datasetspotless-customer-service-training
Spotless Bin Co Customer Service Training Data
Training data for a customer service AI model for Spotless Bin Co, a residential trash can cleaning service.
Dataset Description
This dataset contains 8,776 conversational examples across 5 categories:
Category
Count
Description
FAQs
1,951
Frequently asked questions
Service
1,925
Service explanation dialogues
Objections
1,925
Objection handling examples
Booking
1,975
Booking flow conversations
Brand
1,000… See the full description on the dataset page: https://huggingface.co/datasets/rileyseaburg/spotless-customer-service-training.evaluation-of-customer-servicecustomer-service-nerllm_filtered_customer_service_conversations
LLM-filtered Customer Service Conversations Dataset
Overview
This dataset contains simulated conversations generated by our agentic simulation system.
The conversations are filtered by a LLM to ensure they are of high quality.
Each record is stored in JSON Lines (JSONL) format and includes:
Input Settings: Metadata such as selected bank, customer, agent profiles, and task details.
Messages: The full conversation messages.
Summary: A German summary of the conversation.… See the full description on the dataset page: https://huggingface.co/datasets/marccgrau/llm_filtered_customer_service_conversations.customer_service_dataset
GTBank Customer Service – Synthetic Dataset
Overview
This dataset contains 500 synthetic user_query / assistant_reply pairs modeled on common GTBank (Guaranty Trust Bank) retail banking customer support interactions, such as password resets, debit card activation, failed transactions, and KYC/transfer limits. It is entirely synthetic — no real customer data, transcripts, or GTBank systems were used — and is intended for prototyping and testing customer-support NLP… See the full description on the dataset page: https://huggingface.co/datasets/Ephraimmm/customer_service_dataset.shopeasy-customer-serviceCustomer-Service-AssistantMultilingual-Nepali-Customer-Care-Services-DatasetCustomer-Care-Services-Dataset-in-Nepalillm_filtered_customer_service_conversations_cleaned
LLM-filtered Customer Service Conversations Dataset (cleaned)
Overview
This dataset contains simulated conversations generated by our agentic simulation system.
The conversations are filtered by a LLM to ensure they are of high quality.
Each record is stored in JSON Lines (JSONL) format and includes:
Input Settings: Metadata such as selected bank, customer, agent profiles, and task details.
Messages: The full conversation messages.
Summary: A German summary of the… See the full description on the dataset page: https://huggingface.co/datasets/marccgrau/llm_filtered_customer_service_conversations_cleaned.shopeasy-customer-service-extendedcustomer_service_chatCustomer service chat
customer-service-ai-agent
Customer Service Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/customer-service-ai-agent.finance_customer_service_10000_optimizedcustomer-service-a3x9
Customer Service Queries (a3x9)
Description
This repository is an unofficial mirror of the well-known BANKING77 dataset published by PolyAI, re-uploaded to this hub for internal benchmarking. BANKING77 contains 13,083 customer service queries in the banking domain, labeled with 77 fine-grained intents.
The original data is distributed under a CC BY 4.0 license. No modifications were made to the underlying data.
Source
Source: TBD
License… See the full description on the dataset page: https://huggingface.co/datasets/Roy229/customer-service-a3x9.rtb-customer-service-instruction
Bitext-customer-support-llm-chatbot-training-dataset dataset
Red teaming Bitext-customer-support-llm-chatbot-training-dataset dataset.
Generated from https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset
Dataset Structure
Sample
{
"expected": "Invoice",
"id": "82",
"messages": [
{
"content": "You are a customer service virtual assistant. When given an instruction, you determine which category… See the full description on the dataset page: https://huggingface.co/datasets/innodatalabs/rtb-customer-service-instruction.customer-support-and-servicecustomerservicecustomer-servicelicense: cc-by-4.0
task_categories:
text-generation
question-answering
language:
zh # 修改为中文
tags:
用户手册
操作指南
故障排除
软件支持
技术问答
系统操作
pretty_name: 系统使用问答数据集
size_categories:
10K<n<100K
customer-service-requests-json
Customer Service Requests Dataset
JSON dataset containing common customer service requests
for retail service robots.
customer-service-ar-json
