datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.customer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits… See the full description on the dataset page: https://huggingface.co/datasets/Tobi-Bueck/customer-support-tickets.unpredictable_support-google-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.SupportBench
SupportBench
A multilingual benchmark for evaluating case extraction from real-world tech support group chats.
SupportBench contains 60,000 messages across 6 datasets in 3 languages (English, Spanish, Ukrainian), spanning 6 technical domains. All messages are sourced from public Telegram support groups.
Datasets
Dataset
Language
Domain
Messages
Users
Reply%
Media
Ardupilot-UA
Ukrainian
UAV / Drones
10,000
319
51.8%
1,440
MikroTik-UA
Ukrainian
Networking
10… See the full description on the dataset page: https://huggingface.co/datasets/pavelshpagin/SupportBench.unpredictable_support-google-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.synthetic-it-support-tickets
Synthetic IT Support Tickets — PII-Enriched + Redaction Ground Truth
745 synthetic IT service-management incident records for LLM wiki and
retrieval-augmented-generation experiments. Each record is a help-desk/IT-ops incident with
submitted ticket text, timestamped troubleshooting correspondence, structured diagnostics, root
cause, and resolution steps.
The free text is enriched with realistic technical detail and injected synthetic PII. The corpus
ships two authored… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthetic-it-support-tickets.IT_Support_V2
Mack: IT Support & Admin Dataset
📋 Dataset Description
This dataset consists of 100,000+ conversation logs focused on IT Support and IT Administration tasks. It was generated to fine-tune the "Mack" model—an AI persona designed to act as an expert Tier 1 & Tier 2 IT Helpdesk agent.
The data covers a wide range of technical domains, including Windows troubleshooting, SQL Server administration, driver issues, network diagnostics, and hardware debugging.
Curated by: [Dev… See the full description on the dataset page: https://huggingface.co/datasets/benjaminmacklin/IT_Support_V2.polychart-shown-is-not-supported
Shown Is Not Supported
A chart is a completion claim. It renders cleanly, states a confident finding,
and the underlying data may not support it. A model reading that chart inherits
the gap: it answers from the visual impression because it never extracted the
values.
This adapter reads the data first and corrects the chart when its encoding
misleads. It ships with the first continuous score for how much a chart lies.
Trained with AutoScientist by Adaption for the AutoScientist… See the full description on the dataset page: https://huggingface.co/datasets/rodriguescarson/polychart-shown-is-not-supported.Generated-Recovery-Support-Dialogues
# Empathetic Conversations for Addiction Recovery Support Dataset
Dataset Description
This dataset contains synthetically generated conversational examples between a user discussing their addiction recovery journey and an AI assistant designed to be empathetic, supportive, non-judgmental, and encouraging. The conversations are in English and cover various stages and aspects of the recovery process, following established therapeutic guidelines and models.
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/filippo19741974/Generated-Recovery-Support-Dialogues.customer-service-robot-support
This Dialogue
Comprised of fictitious examples of dialogues between a customer encountering problems with a robotic arm and a technical support agent. Check out the example below:
"id": 1,
"description": "Robotic arm calibration issue",
"dialogue": "Customer: My robotic arm seems to be misaligned. It's not picking objects accurately. What can I do? Agent: It appears that the arm may need recalibration. Please follow the instructions in the user manual to reset the calibration… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/customer-service-robot-support.customer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits your needs.… See the full description on the dataset page: https://huggingface.co/datasets/vasu1111/customer-support-tickets.customer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits your needs.… See the full description on the dataset page: https://huggingface.co/datasets/Prady06/customer-support-tickets.amazon-customer-support
Amazon Customer Support
Derived from the TWCS corpus (Kaggle: thoughtvector/customer-support-on-twitter), this dataset
contains 200 labelled customer-support interactions for evaluation / fine-tuning purposes.
Schema
Each line is a JSON object with four fields:
Field
Type
Description
query
string
The customer's raw message (input)
action
string
Agent action taken — resolve or escalated_to_human
intent
string
Classified intent — complaint, question… See the full description on the dataset page: https://huggingface.co/datasets/iam-tsr/amazon-customer-support.TREC_Clinicial-Decision-Support
TREC Clinical Decision Support (2014, 2015 and 2016)
Dataset Description
Links
Homepage:
TREC
Paper:
2014 / 2015 / 2016
Contact (Original Authors):
Kirk Roberts (Kirk.Roberts@uth.tmc.edu )
Contact (Curator):
Artur Guimarães (artur.guimas@gmail.com)
Dataset Summary
This track focused on clinicians looking for evidence-based full-text literature to support diagnosis, treatment, and testing decisions
Data Instances
Source… See the full description on the dataset page: https://huggingface.co/datasets/araag2/TREC_Clinicial-Decision-Support.IT_Support
Mack IT Support Datasets
The Mack dataset is a collection of high-quality IT support data curated for developing and benchmarking agentic language models, digital helpdesk assistants, and troubleshooting bots.It contains seven .jsonl files with diverse coverage:
A_identity.jsonl: Agent identity and persona modeling.
B_troubleshooting.jsonl: Stepwise troubleshooting dialogs and solutions.
C_steps.jsonl: IT procedures and diagnostic workflow data.
D_reasoning.jsonl: Support agent… See the full description on the dataset page: https://huggingface.co/datasets/benjaminmacklin/IT_Support.arabic-rag-support-25K
Arabic RAG customer-support scenarios (27,927 rows)
Synthetic Modern Standard Arabic customer-support scenarios for training small
RAG answerers, distilled from unsloth/gemma-4-31B-it-NVFP4 on a local vLLM.
Built as the training set for oddadmix/Nawah-50M-RAG-Support.
Each row: a customer question + the knowledge-base chunks of one fictional
company (products, prices, policies, FAQ entries) + the ideal grounded agent
answer. One generation request invents one company KB and 4 QA… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-support-25K.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/abhi23457/Bitext-customer-support-llm-chatbot-training-dataset.it-support-l1-ticket-classification
IT Support L1 Multilingual Dataset
Dataset Summary
IT Support L1 Multilingual Dataset is a synthetic enterprise help desk dataset for ticket classification and troubleshooting response generation. It contains realistic Level 1 IT support scenarios in English and Czech, designed for experiments in structured classification, response generation, and multilingual support workflow prototyping.
This dataset contains synthetic IT Support L1 scenarios. The records were generated… See the full description on the dataset page: https://huggingface.co/datasets/w1z4rd3k/it-support-l1-ticket-classification.live-face-swap-support-qa
LiveFaceSwap AI Public Support Q&A
This dataset contains English question and answer pairs from the public LiveFaceSwap AI browser, desktop, and pricing FAQs, captured on 2026-09-17. Each row records its source page. The official website is the current source for product behavior and pricing; this snapshot can become outdated.
It is suitable for evaluating or prototyping retrieval over LiveFaceSwap AI product support content. It is not a face image dataset, a face swap training… See the full description on the dataset page: https://huggingface.co/datasets/LiveFaceSwapAI/live-face-swap-support-qa.MentalHealth-Support
Important Note
This dataset is created from merging two datasets from different sources and has been formatted according to the "messages", "role", "content" chat format. I do not claim any ownership of this dataset.
Keep in mind that this dataset is entirely synthetic. It is not fully representative of real therapy situations. If you are training an LLM therapist keep in mind the limitations of LLMs and highlight those limitations to users in a responsible manner.
Since Mental… See the full description on the dataset page: https://huggingface.co/datasets/ShivomH/MentalHealth-Support.customer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits your needs.… See the full description on the dataset page: https://huggingface.co/datasets/alibinfaizan/customer-support-tickets.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/ljoaql/Bitext-customer-support-llm-chatbot-training-dataset.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/gorges-haha/Bitext-customer-support-llm-chatbot-training-dataset.customer_support_auto_completioncustomer-support-tickets
Featuring Labeled Customer Emails and Support Responses
🔧 Synthetic IT Ticket Generator — Custom Dataset
Create a dataset tailored to your own queues & priorities (no PII).
👉 Generate custom data
Define your queues, priorities, language
Need an on-prem AI to auto-classify tickets?→ Open Ticket AI
There are 2 Versions of the dataset, the new version has more tickets, but only languages english and german. So please look at both files, to find what best fits your needs.… See the full description on the dataset page: https://huggingface.co/datasets/manojroyal23/customer-support-tickets.kazakh-customer-support-qa
Dataset Card for kazakh-customer-support-qa
Maintained by: Mäñgi ÜTM (mangi-llm)
Dataset Summary
kazakh-customer-support-qa is a small, hand-curated question–answer dataset in the Kazakh language, built to represent realistic customer-support conversations across several industries (banking, telecom, retail/service centers, sales, and general support). Each record pairs a short customer question with a concise, policy-safe answer, and many answers include… See the full description on the dataset page: https://huggingface.co/datasets/mangi-llm/kazakh-customer-support-qa.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/vshanks/Bitext-customer-support-llm-chatbot-training-dataset.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/mostafafhasjk/Bitext-customer-support-llm-chatbot-training-dataset.uzbek-customer-support-dialogs
Uzbek Customer Support Dialogs 🇺🇿
A high-quality dataset of 990 customer support conversations in Uzbek (Latin script), designed for training and fine-tuning conversational AI models.
This is one of the first large-scale customer support datasets in Uzbek, created to address the gap of low-resource NLP for Central Asian languages.
📋 Dataset Description
990 conversational dialogs in natural Uzbek (Latin script)
11 customer support categories: Order, Shipping, Cancel… See the full description on the dataset page: https://huggingface.co/datasets/AsrorAsr/uzbek-customer-support-dialogs.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/rohityadavv/Bitext-customer-support-llm-chatbot-training-dataset.
