datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.customer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits… See the full description on the dataset page: https://huggingface.co/datasets/Tobi-Bueck/customer-support-tickets.Customer_support_faqs_datasetDataset Name: Customer Support FAQs Dataset
Description:
This dataset contains a collection of 200 frequently asked questions (FAQs) and their corresponding answers, designed to assist in customer support scenarios. The questions cover a wide range of common customer inquiries related to account management, payment methods, order tracking, shipping, returns, and more. This dataset is intended for use in developing and training AI models for customer support chatbots, automated response systems… See the full description on the dataset page: https://huggingface.co/datasets/MakTek/Customer_support_faqs_dataset.unpredictable_support-google-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.Customer-Support-Responsescustomer-support-on-twitter-conversationunpredictable_support-google-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.Customer_Support_on_Twitterinsuff_supported_argumentsmo-customer-support-tweets-945k
Customer Support on Twitter Dataset 945k
Dataset Description
Context
This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions.
Content
Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/MohammadOthman/mo-customer-support-tweets-945k.synthetic-it-support-tickets
Synthetic IT Support Tickets — PII-Enriched + Redaction Ground Truth
745 synthetic IT service-management incident records for LLM wiki and
retrieval-augmented-generation experiments. Each record is a help-desk/IT-ops incident with
submitted ticket text, timestamped troubleshooting correspondence, structured diagnostics, root
cause, and resolution steps.
The free text is enriched with realistic technical detail and injected synthetic PII. The corpus
ships two authored… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthetic-it-support-tickets.task083_babi_t1_single_supporting_fact_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task083_babi_t1_single_supporting_fact_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task083_babi_t1_single_supporting_fact_answer_generation.E-Commerce_Customer_Support_Conversations
Dataset Card for "E-Commerce_Customer_Support_Conversations"
The dataset is synthetically generated with OpenAI ChatGPT model (gpt-3.5-turbo).
More Information needed
task084_babi_t1_single_supporting_fact_identify_relevant_fact
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task084_babi_t1_single_supporting_fact_identify_relevant_fact
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task084_babi_t1_single_supporting_fact_identify_relevant_fact.customer-support-client-agent-conversations
Customer Support Client-Agent Conversations Dataset
A synthetic context-summarized multi-turn customer-service question-answering dataset for banking domain conversations, designed for training and evaluating small language models on dialogue continuity and contextual understanding tasks.
Dataset Description
This dataset contains 183,337 context-summarized multi-turn customer-service conversations spanning various banking scenarios including account management… See the full description on the dataset page: https://huggingface.co/datasets/Lakshan2003/customer-support-client-agent-conversations.e-commerce-customer-support-qa
Dataset Card for Dataset Name
from: NebulaByte/E-Commerce_Customer_Support_Conversations
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/rjac/e-commerce-customer-support-qa.processed_support_ticketslingrow-support-tickets
Lingrow Support Tickets (Synthetic)
A synthetic dataset of 10,000 customer-support tickets for Lingrow,
a real-time multilingual translation and communication platform. Each ticket
contains a customer message (an error report or a how-to question), rich
metadata, and a resolution. The data is fully synthetic — no real customer
information is included.
This dataset was built as the final project for a Data Science course. It powers
the Lingrow Support Copilot: a tool that, given… See the full description on the dataset page: https://huggingface.co/datasets/adiprog14/lingrow-support-tickets.lfqa_support_docsSupport documents for building https://huggingface.co/vblagoje/bart_lfqa model
customer_support_ticketsIT_Support_V2
Mack: IT Support & Admin Dataset
📋 Dataset Description
This dataset consists of 100,000+ conversation logs focused on IT Support and IT Administration tasks. It was generated to fine-tune the "Mack" model—an AI persona designed to act as an expert Tier 1 & Tier 2 IT Helpdesk agent.
The data covers a wide range of technical domains, including Windows troubleshooting, SQL Server administration, driver issues, network diagnostics, and hardware debugging.
Curated by: [Dev… See the full description on the dataset page: https://huggingface.co/datasets/benjaminmacklin/IT_Support_V2.twitter_customer_support_weaviate_export_200000_text-embedding-3-smallecommerce-customer-support-conversationsE-Commerce Customer Support Conversations
Dataset Summary:
This dataset contains customer support queries and responses from an e-commerce context.
It is designed for training and fine-tuning AI models for automated customer service, chatbots, and natural language processing (NLP) applications.
Use Cases:
Fine-tuning conversational AI models (e.g., GPT, BERT)
Training chatbots for e-commerce support
Improving customer service automation
Sentiment and intent analysis
Dataset Format:
The… See the full description on the dataset page: https://huggingface.co/datasets/Venkatrajan247/ecommerce-customer-support-conversations.torch-tpu-model-support
torch-tpu-model-support
Support sweep results for HuggingFace models on Google Cloud TPU with torch_tpu.
Each row is one model tested in one sweep. The sweep loads the model and
generates 8 tokens with the fixed prompt "The future of AI is" in bf16.
Result values
PASS — model loaded and generated successfully
FAIL(load) — exception during model load
FAIL(gen) — load succeeded, generation raised
FAIL(output) — generation returned but the text was judged not valid… See the full description on the dataset page: https://huggingface.co/datasets/hf-gcp-tpu-internal/torch-tpu-model-support.tts-eval-customer-support-202608
TTS Hard Cases — Customer Support
A compact, text-only evaluation set of hard cases for text-to-speech models, focused on the
customer support use case. Published by Gradium.
Most TTS benchmarks measure naturalness on ordinary prose, where modern models are already close to
saturated. Production voice agents fail somewhere else: on the literal payload of a support call —
the order number, the email address, the spelled-out surname, the date, the amount refunded. A voice
that… See the full description on the dataset page: https://huggingface.co/datasets/gradium/tts-eval-customer-support-202608.saas-product-support-sharegpt-1k
SaaS/Tech Product Support — Multi-Turn SFT Dataset
A domain-specific supervised fine-tuning dataset for
SaaS and tech product support conversations, built for
LLM fine-tuning and instruction tuning.
Dataset Summary
This dataset contains 1,200 multi-turn English conversations
between a customer and a support agent, covering common
SaaS/tech support scenarios: bug reports, billing issues,
API errors, authentication problems, onboarding blockers,
integration failures… See the full description on the dataset page: https://huggingface.co/datasets/Dang-DN-VN/saas-product-support-sharegpt-1k.insurance_customer_support_conversationcustomer-support-twitter-datasetBitext-customer-support-llm-chatbot-training-dataset-spanish
Spanish Customer Support LLM Chatbot Training Dataset
Spanish-language adaptation of the Bitext Customer Support LLM Chatbot Training Dataset.
This dataset is intended for training and evaluating Spanish-language customer-support chatbots and instruction-following large language models.
Dataset Details
Dataset Description
This dataset is a Spanish translation and adaptation of the original Bitext Customer Support LLM Chatbot Training Dataset.
The… See the full description on the dataset page: https://huggingface.co/datasets/Faramir/Bitext-customer-support-llm-chatbot-training-dataset-spanish.food-delivery-support-tickets
Food Delivery Support Tickets (synthetic)
10,153 synthetic English customer-support conversations for a food delivery
platform (à la Wolt / Uber Eats / DoorDash). Each record is a realistic customer
message with structured labels and a professional agent resolution + reply.
Built for the Food Delivery Support Copilot — an assistant that classifies an
incoming ticket, retrieves similar resolved cases, and drafts a reply.
How it was made
Generated locally with the… See the full description on the dataset page: https://huggingface.co/datasets/OrSabbach/food-delivery-support-tickets.
