datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.customer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits… See the full description on the dataset page: https://huggingface.co/datasets/Tobi-Bueck/customer-support-tickets.controlled_anchor_v1_support_switch
Controlled ICIL Anchor V1 Support Switch
LeRobot conversion of the Anchor V1 controlled ICIL collection.
Source HDF5:
/ibex/project/c2090/jian/icil_openpi/ICIL/data/manifest_collection_v1/controlled_anchor_v1_60a_6p_3j_6obj_res256_lzf_merged.hdf5
OpenPI sidecars are stored under meta/controlled_icil/.
Customer_support_faqs_datasetDataset Name: Customer Support FAQs Dataset
Description:
This dataset contains a collection of 200 frequently asked questions (FAQs) and their corresponding answers, designed to assist in customer support scenarios. The questions cover a wide range of common customer inquiries related to account management, payment methods, order tracking, shipping, returns, and more. This dataset is intended for use in developing and training AI models for customer support chatbots, automated response systems… See the full description on the dataset page: https://huggingface.co/datasets/MakTek/Customer_support_faqs_dataset.unpredictable_support-google-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.SupportBench
SupportBench
A multilingual benchmark for evaluating case extraction from real-world tech support group chats.
SupportBench contains 60,000 messages across 6 datasets in 3 languages (English, Spanish, Ukrainian), spanning 6 technical domains. All messages are sourced from public Telegram support groups.
Datasets
Dataset
Language
Domain
Messages
Users
Reply%
Media
Ardupilot-UA
Ukrainian
UAV / Drones
10,000
319
51.8%
1,440
MikroTik-UA
Ukrainian
Networking
10… See the full description on the dataset page: https://huggingface.co/datasets/pavelshpagin/SupportBench.customer-support-on-twitter-conversationcustomer_support_conversations_dataset
💬 Customer Support Conversation Dataset — Powered by Syncora.ai
A free synthetic dataset for chatbot training, LLM fine-tuning, and synthetic data generation research.Created using Syncora.ai’s privacy-safe synthetic data engine, this dataset is ideal for developing, testing, and benchmarking AI customer support systems.
It serves as a dataset for chatbot training and a dataset for LLM training, offering rich, structured conversation data for real-world simulation.
🌟… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/customer_support_conversations_dataset.Customer-Support-Responsesunpredictable_support-google-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.Customer_Support_on_Twitterinsuff_supported_argumentsmo-customer-support-tweets-945k
Customer Support on Twitter Dataset 945k
Dataset Description
Context
This dataset provides a large corpus of real-world English conversations between consumers and customer support agents on Twitter, designed to drive innovation in Natural Language Processing (NLP) by providing data that better matches the actual language used in contemporary customer support interactions.
Content
Initially, the data included complex threads of conversations… See the full description on the dataset page: https://huggingface.co/datasets/MohammadOthman/mo-customer-support-tweets-945k.controlled_anchor_v0_support_switch
Controlled ICIL Anchor v0 Support Switch
LeRobot conversion of the controlled ICIL anchor-v0 atomic demonstrations, with
the challenge support-switch pair manifests attached under
meta/controlled_icil/pair_manifests/anchor_v0_challenge_split.
Counts
Episodes: 2379
Frames: 236436
Skipped source episodes: 0
Train anchors: 165
Test anchors: 20
Excluded incomplete anchors: 15
Train support-target pairs: 1980
Test support-target pairs: 240
OpenPI paths… See the full description on the dataset page: https://huggingface.co/datasets/daixianjie/controlled_anchor_v0_support_switch.pydreg-supporting-data
pydreg vs. dREG benchmark outputs
Raw benchmark artifacts backing the performance and accuracy comparisons in
pydreg, a from-scratch Python port of
dREG (Danko Lab). This dataset holds the
paired outputs of running both tools' full peak-calling pipeline
(run_dREG/pydreg) on the same 12 real PRO-seq/GRO-seq/ChRO-seq libraries,
plus the /usr/bin/time -v logs used to compare wall-clock time and peak
memory. It is data, not code — see the pydreg repo for the package itself and
for… See the full description on the dataset page: https://huggingface.co/datasets/adamyhe/pydreg-supporting-data.supportSupportsynthetic-it-support-tickets
Synthetic IT Support Tickets — PII-Enriched + Redaction Ground Truth
745 synthetic IT service-management incident records for LLM wiki and
retrieval-augmented-generation experiments. Each record is a help-desk/IT-ops incident with
submitted ticket text, timestamped troubleshooting correspondence, structured diagnostics, root
cause, and resolution steps.
The free text is enriched with realistic technical detail and injected synthetic PII. The corpus
ships two authored… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthetic-it-support-tickets.processed_support_ticketstask083_babi_t1_single_supporting_fact_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task083_babi_t1_single_supporting_fact_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task083_babi_t1_single_supporting_fact_answer_generation.task084_babi_t1_single_supporting_fact_identify_relevant_fact
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task084_babi_t1_single_supporting_fact_identify_relevant_fact
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task084_babi_t1_single_supporting_fact_identify_relevant_fact.huggingface_filesystem_terminal_12687_supportE-Commerce_Customer_Support_Conversations
Dataset Card for "E-Commerce_Customer_Support_Conversations"
The dataset is synthetically generated with OpenAI ChatGPT model (gpt-3.5-turbo).
More Information needed
lingrow-support-tickets
Lingrow Support Tickets (Synthetic)
A synthetic dataset of 10,000 customer-support tickets for Lingrow,
a real-time multilingual translation and communication platform. Each ticket
contains a customer message (an error report or a how-to question), rich
metadata, and a resolution. The data is fully synthetic — no real customer
information is included.
This dataset was built as the final project for a Data Science course. It powers
the Lingrow Support Copilot: a tool that, given… See the full description on the dataset page: https://huggingface.co/datasets/adiprog14/lingrow-support-tickets.e-commerce-customer-support-qa
Dataset Card for Dataset Name
from: NebulaByte/E-Commerce_Customer_Support_Conversations
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/rjac/e-commerce-customer-support-qa.lfqa_support_docsSupport documents for building https://huggingface.co/vblagoje/bart_lfqa model
twitter_customer_support_weaviate_export_200000_text-embedding-3-smallcustomer-support-client-agent-conversations
Customer Support Client-Agent Conversations Dataset
A synthetic context-summarized multi-turn customer-service question-answering dataset for banking domain conversations, designed for training and evaluating small language models on dialogue continuity and contextual understanding tasks.
Dataset Description
This dataset contains 183,337 context-summarized multi-turn customer-service conversations spanning various banking scenarios including account management… See the full description on the dataset page: https://huggingface.co/datasets/Lakshan2003/customer-support-client-agent-conversations.customer_support_ticketsminddistiller-support-data-20260626
MindDistiller encrypted support data
Encrypted support data for restoring MindDistiller input datasets, exported result sets, and materialized taskset directories.
Base files:
minddistiller-data-input-20260626.tar.gz.gpg: encrypted tarball of data/input as of 2026-06-26.
minddistiller-data-results_sets-20260626.tar.gz.gpg: encrypted tarball of data/results_sets as of 2026-06-26.
Incremental files added 2026-06-30:… See the full description on the dataset page: https://huggingface.co/datasets/PGCodeLLM/minddistiller-support-data-20260626.
