datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.Customer_support_faqs_datasetDataset Name: Customer Support FAQs Dataset
Description:
This dataset contains a collection of 200 frequently asked questions (FAQs) and their corresponding answers, designed to assist in customer support scenarios. The questions cover a wide range of common customer inquiries related to account management, payment methods, order tracking, shipping, returns, and more. This dataset is intended for use in developing and training AI models for customer support chatbots, automated response systems… See the full description on the dataset page: https://huggingface.co/datasets/MakTek/Customer_support_faqs_dataset.autotrain-data-quality-customer-reviews
AutoTrain Dataset for project: quality-customer-reviews
Dataset Descritpion
This dataset has been automatically processed by AutoTrain for project quality-customer-reviews.
Languages
The BCP-47 code for the dataset's language is en.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"text": " Love this truck, I think it is light years better than the competition. I have driven or… See the full description on the dataset page: https://huggingface.co/datasets/Gflorent/autotrain-data-quality-customer-reviews.268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Data-Sample
Description
Arabic(Saudi) Multi-stream Spontaneous Dialogue Smartphone speech dataset-Customer Service. Transcribed with text content, speaker's ID, gender, age and other attributes. Our dataset was collected from extensive and diversify speakers(268 native speakers), geographicly speaking, enhancing model performance in real and complex tasks.
For more details, please refer to the link: https://www.nexdata.ai/datasets/speechrecog/1627?source=Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-AI/268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Data-Sample.customer-support-twitter-datasetBitext-customer-support-llm-chatbot-training-dataset-spanish
Spanish Customer Support LLM Chatbot Training Dataset
Spanish-language adaptation of the Bitext Customer Support LLM Chatbot Training Dataset.
This dataset is intended for training and evaluating Spanish-language customer-support chatbots and instruction-following large language models.
Dataset Details
Dataset Description
This dataset is a Spanish translation and adaptation of the original Bitext Customer Support LLM Chatbot Training Dataset.
The… See the full description on the dataset page: https://huggingface.co/datasets/Faramir/Bitext-customer-support-llm-chatbot-training-dataset-spanish.crros-customer-behavior-dataset
CRROS Customer Behavior Dataset
This dataset is part of my Customer Retention & Revenue Optimization System (CRROS) project. The goal of the project is to simulate realistic customer behavior and use it to build an end-to-end customer analytics workflow, from raw data all the way to business decisions.
Instead of generating completely random records, the dataset follows business-driven rules that simulate how customers interact with products, make purchases, become inactive over… See the full description on the dataset page: https://huggingface.co/datasets/nibeditans/crros-customer-behavior-dataset.268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Dataset
Description
사우디아라비아 아랍어(Arabic-Saudi) 멀티스트림 자연 대화 스마트폰 고객 서비스 음성 데이터셋입니다. 다양한 고객 서비스 상황에서 자유롭게 대화하는 방식으로 수집되었으며, 전사 텍스트와 함께 화자 ID, 성별, 연령 등의 속성 정보가 제공됩니다. 총 268명의 아랍어 원어민 화자로부터 데이터를 수집하여 다양한 화자 특성을 반영했으며, 실제 환경에서 발생하는 복잡하고 다양한 음성 상황에 대한 모델의 성능 향상을 지원합니다.
자세한 내용은 아래 링크를 참고해 주세요: https://ko.nexdata.ai/datasets/speechrecog/1627?source=hf.kr
Specifications
Format
16kHz, 16 bit, WAV, 모노 채널
Content category
정해진 주제 없이 자유롭게 진행된 자연 대화… See the full description on the dataset page: https://huggingface.co/datasets/Nexdata-kr/268-Hours-Arabic-Saudi-Full-Duplex-Multi-Channel-Customer-Service-Speech-Dataset.CUSTOMER_SUPPORT_DATASET_JSONL_V1VNOVA AI — Customer Support Dataset (100 Synthetic Scenarios)
High-quality synthetic customer support conversations designed for training and evaluating AI support agents.
About the Dataset
This dataset contains 100 fully synthetic customer–agent interactions across multiple industries, created to help developers train:
1-Customer support chatbots
2-Automated ticket resolution systems
3-LLM fine-tuning for support tasks
4-RAG workflows
5-Complaint classification models
6-Sentiment & intent… See the full description on the dataset page: https://huggingface.co/datasets/vnovaai/CUSTOMER_SUPPORT_DATASET_JSONL_V1.africa-synth-retail-and-ecommerce-customer-segmentation-data-nigeria
Customer Segmentation Data | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-retail-and-ecommerce-customer-segmentation-data-nigeria.tourism-customer-dataCustomer-Churn-Dataset-V2
Customer Churn Conversation Dataset - 500 (Benchmark-Anchored)
Free 500-record sample. Licensed CC BY-NC 4.0. Commercial use requires a license.
The generator is the product
This sample was produced by our synthetic customer-churn dialogue generator. The generator is what we license: it produces a labeled 10,000-record dataset anchored to published subscription-industry benchmarks, with a cleaner and refiner pipeline built in. Real churn conversations are locked… See the full description on the dataset page: https://huggingface.co/datasets/ConsumerDividends/Customer-Churn-Dataset-V2.Customer_Reviews_Dataset_Online_Food_Ordering_Portal_Bangladesh
Citation Requirement:
Anyone using this dataset must provide proper citation in their work.
Cite Our Paper at : https://ieeexplore.ieee.org/document/10871506
About Dataset
The dataset is a collection of customer reviews on a food delivery website for Bangladesh named Foodpanda Bangladesh. The dataset includes 43,061 reviews. It contains customer feedback in a mix of Bangla, English, and Romanized Bangla, translated into English along with their respective ratings.
The… See the full description on the dataset page: https://huggingface.co/datasets/farhanenzo/Customer_Reviews_Dataset_Online_Food_Ordering_Portal_Bangladesh.rooyai-customer-service-datasetBitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/abhi23457/Bitext-customer-support-llm-chatbot-training-dataset.customer-support-training-dataset
dataset_info: features: - name: flags dtype: string - name: instruction dtype: string - name: categorydtype: string - name: intent dtype: string - name: response dtype: string splits: - name: train num_bytes: 19526505 num_examples: 26872 download_size: 6048908 dataset_size: 19526505configs:- config_name: default data_files: - split: train path: data/train-*license: mittask_categories:- text-generationlanguage:- entags:- financepretty_name:… See the full description on the dataset page: https://huggingface.co/datasets/Victorano/customer-support-training-dataset.customer-service
Customer Service Conversations Dataset
This dataset contains 100 realistic customer service conversations between customers and support agents. Each dialogue is 11 turns long and covers a variety of common issues such as late deliveries, billing errors, account problems, and more. It is ideal for training and evaluating AI assistants, chatbots, and customer support models.
Dataset Structure
Each conversation is stored as a JSON object with the following fields:
id:… See the full description on the dataset page: https://huggingface.co/datasets/ai-training-datasets/customer-service.bitext-customer-support-llm-chatbot-training-datasetcustomer-support-dataset-predibaseBitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/ljoaql/Bitext-customer-support-llm-chatbot-training-dataset.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/gorges-haha/Bitext-customer-support-llm-chatbot-training-dataset.Customer_support_faqs_datasetDataset Name: Customer Support FAQs Dataset
Description:
This dataset contains a collection of 200 frequently asked questions (FAQs) and their corresponding answers, designed to assist in customer support scenarios. The questions cover a wide range of common customer inquiries related to account management, payment methods, order tracking, shipping, returns, and more. This dataset is intended for use in developing and training AI models for customer support chatbots, automated response systems… See the full description on the dataset page: https://huggingface.co/datasets/Yatzo/Customer_support_faqs_dataset.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/vshanks/Bitext-customer-support-llm-chatbot-training-dataset.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/mostafafhasjk/Bitext-customer-support-llm-chatbot-training-dataset.Matan-Kriel-Customer-Churn-dataset
Project Overview:
Uncovering Critical Churn Drivers
This project employed Exploratory Data Analysis (EDA) techniques on a comprehensive customer dataset with the primary objective of diagnosing the root causes of severe customer attrition.
The Business Challenge:
The initial analysis revealed an alarmingly high baseline churn rate of 56.7%, indicating a massive challenge in customer retention and significant revenue leakage. The core mandate was to move beyond… See the full description on the dataset page: https://huggingface.co/datasets/MatanKriel/Matan-Kriel-Customer-Churn-dataset.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/rohityadavv/Bitext-customer-support-llm-chatbot-training-dataset.Customer_service_datapersian_customer-service_datasetsafrica-synth-retail-and-ecommerce-customer-review-and-ratings-data-nigeria
Customer Review And Ratings Data | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-retail-and-ecommerce-customer-review-and-ratings-data-nigeria.Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/Nilesh987/Bitext-customer-support-llm-chatbot-training-dataset.
