datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.customer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits… See the full description on the dataset page: https://huggingface.co/datasets/Tobi-Bueck/customer-support-tickets.Customer-Support-Responsesinsuff_supported_argumentstts-eval-customer-support-202608
TTS Hard Cases — Customer Support
A compact, text-only evaluation set of hard cases for text-to-speech models, focused on the
customer support use case. Published by Gradium.
Most TTS benchmarks measure naturalness on ordinary prose, where modern models are already close to
saturated. Production voice agents fail somewhere else: on the literal payload of a support call —
the order number, the email address, the spelled-out surname, the date, the amount refunded. A voice
that… See the full description on the dataset page: https://huggingface.co/datasets/gradium/tts-eval-customer-support-202608.Support-Bot-Recommendationfetch_terminal_huggingface_9087_861abb34_novaretail_support_tickets
NovaRetail Support Ticket Logs
This dataset contains anonymized support ticket logs from the NovaRetail help desk. It is used to benchmark intent classification models.
Dataset Metadata
License: cc-by-nc-4.0
Origin: NovaRetail Help Desk
Review Status: reviewed
Data Fields
ticket_id: unique identifier of the ticket
text: anonymized support request text
category: ticket category (account / shipping / returns / technical)
customer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits your needs.… See the full description on the dataset page: https://huggingface.co/datasets/Prady06/customer-support-tickets.customer-service-robot-support
This Dialogue
Comprised of fictitious examples of dialogues between a customer encountering problems with a robotic arm and a technical support agent. Check out the example below:
"id": 1,
"description": "Robotic arm calibration issue",
"dialogue": "Customer: My robotic arm seems to be misaligned. It's not picking objects accurately. What can I do? Agent: It appears that the arm may need recalibration. Please follow the instructions in the user manual to reset the calibration… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/customer-service-robot-support.customer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits your needs.… See the full description on the dataset page: https://huggingface.co/datasets/vasu1111/customer-support-tickets.fsfh6410-support-tickets-0c136b
Support Tickets
Support ticket records for triage processing. Each row contains the ticket id,
customer name, category, priority, status, refund amount (where applicable),
created date, and a short description.
tickets.csv: the ticket records.
policy.md: the triage policy to apply.
This dataset is used by the customer support operations team for automated
triage. The output report is published to a repository with the prefix
fsfh6410-triage-report.
filesystem_huggingface_9840_z1xmjtic_support_tickets
Support Ticket Export
Fresh export of support tickets from the company's customer support system.
Each record contains ticket metadata, priority, assignment, SLA deadline and
whether the ticket has breached its SLA window. This is the source dataset for
the support-ticket triage workflow.
File: tickets.csv
github-fetch-hf-9020-support-tickets
Support Ticket Triage Data
Overview
This dataset contains anonymised support ticket records used for triage analysis and customer service quality monitoring. Each record includes ticket identifiers, creation timestamps, customer references, priority levels, current status, and resolution time in hours.
Data Fields
ticket_id (string): Unique identifier for each support ticket.
created_at (string): ISO 8601 timestamp when the ticket was created.… See the full description on the dataset page: https://huggingface.co/datasets/TianfuXinqu/github-fetch-hf-9020-support-tickets.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/abhi23457/Bitext-customer-support-llm-chatbot-training-dataset.it_support_ticket_classification_pegasus_dataset
IT Support Ticket Classification
Description: Automatically categorize and prioritize IT support tickets based on their text descriptions, enabling more efficient resolution and customer support.
How to Use
Here is how to use this model to classify text into different categories:
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model_name = "interneuronai/it_support_ticket_classification_pegasus"
model =… See the full description on the dataset page: https://huggingface.co/datasets/interneuronai/it_support_ticket_classification_pegasus_dataset.it-support-llmcustomer-support-dataset-predibaseemotional-supportcustomer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits your needs.… See the full description on the dataset page: https://huggingface.co/datasets/alibinfaizan/customer-support-tickets.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/ljoaql/Bitext-customer-support-llm-chatbot-training-dataset.Contextual_Response_Evaluation_for_ESL_and_ASD_Support
Dataset Card for "Contextual Response Evaluation for ESL and ASD Support💜💬🌐""
Dataset Description 📖
Dataset Summary 📝
Curated by Eric Soderquist, this dataset is a collection of English prompts and responses generated by the Phi-2 model, designed to evaluate and improve NLP models for supporting ESL (English as a Second Language) and ASD (Autism Spectrum Disorder) user bases. Each prompt is paired with multiple AI-generated responses and evaluated using a… See the full description on the dataset page: https://huggingface.co/datasets/yunjaeys/Contextual_Response_Evaluation_for_ESL_and_ASD_Support.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/gorges-haha/Bitext-customer-support-llm-chatbot-training-dataset.self-collected-ENEM-dataset-with-prompts-and-text-supportBitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/vshanks/Bitext-customer-support-llm-chatbot-training-dataset.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/mostafafhasjk/Bitext-customer-support-llm-chatbot-training-dataset.customer_support_auto_completionsupport-message-categorizationcustomer-support-csatclinical-discharge-plan-home-support-coherence-risk-v0.1What this repo is for
Detect when
discharge plans
and
real home support
do not match
before
readmission
or harm at home.
insurance_customer_support_conversation
Dataset Card for Insurance Customer Support Conversation Dataset
This is a synthetic dataset generated by ChatGPT-4o.
Dataset Details
Field Descriptions
conversation:
Type: StringDescription: The entire text of the conversation between the customer and the agent, including both parties' dialogues.Example:
"Customer: Hello, I need to discuss the status of my insurance claim. It's been over a month, and I haven't received any updates.
Agent: Good afternoon… See the full description on the dataset page: https://huggingface.co/datasets/aibabyshark/insurance_customer_support_conversation.
