datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chatbot_instruction_prompts
Dataset Card for Chatbot Instruction Prompts Datasets
Dataset Summary
This dataset has been generated from the following ones:
tatsu-lab/alpaca
Dahoas/instruct-human-assistant-prompt
allenai/prosocial-dialog
The datasets has been cleaned up of spurious entries and artifacts. It contains ~500k of prompt and expected resposne. This DB is intended to train an instruct-type model
Bread-chatbot-dataset-test
Dataset Card for "Bread-chatbot-dataset-test"
More Information needed
mental_health_chatbot_dataset
Dataset Card for "heliosbrahma/mental_health_chatbot_dataset"
Dataset Description
Dataset Summary
This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters.
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/heliosbrahma/mental_health_chatbot_dataset.cukurova_university_chatbot
Çukurova University Computer Engineering Chatbot Dataset
📊 Dataset Overview
This dataset contains 22,524 high-quality question-answer pairs specifically designed for training an AI chatbot that serves the Computer Engineering Department at Çukurova University. The dataset is part of the CengBot project, a sophisticated multilingual Telegram chatbot that provides automated assistance to students regarding courses, programs, and departmental information.
🔢… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/cukurova_university_chatbot.Chatbot-Url
Dataset Card for Wikimedia Wikipedia
Dataset Summary
Wikipedia dataset containing cleaned articles of all languages.
The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/)
with one subset per language, each containing a single train split.
Each example contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).
All language subsets have already been processed for recent dump… See the full description on the dataset page: https://huggingface.co/datasets/zabr946/Chatbot-Url.Full-Ecom-Chatbot-Dataset
E-commerce Chatbot Training Data
A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains.
Dataset Summary
Split
Records
Train
35,213
Test
8,818
Total
44,031
The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.NBRO-Chatbot-V1
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: National Building Research Organisation (NBRO)
Funded by : NBRO
Language(s) (NLP): English (en)
License: Creative Commons Attribution 4.0 International (CC BY 4.0)
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/nimesh7814/NBRO-Chatbot-V1.Ecom-Chatbot-Finetuning-Dataset
Ecom Chatbot Finetuning Dataset
A unified instruction-following dataset for fine-tuning e-commerce customer service chatbots. It covers a wide range of real-world retail scenarios — from product discovery and order management to returns, complaints, and account support.
Dataset Summary
Field
Value
Total records
40,098
Language
English
Sources
Amazon Reviews 2023, Amazon Meta 2023, ASOS, Bitext
Response types
Text, Tool Call, Mixed
Difficulty levels
1… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Ecom-Chatbot-Finetuning-Dataset.saas-chatbot-v4
SaaS Chatbot V4 Dataset
Multi-industry, multilingual conversational dataset for fine-tuning LLMs as SaaS AI chatbot agents with tool calling.
Stats
Metric
Value
Train
4,043
Test
450
Total messages
64,645
Avg msgs/conv
14.4
Think blocks
29,345 (21% empty)
Tool calls
15,215
Tool responses
15,387
Industries (8)
E-commerce (1,301), Travel (641), Services (504), Food (490), Beauty (478), Healthcare (404), Education (357), Real Estate… See the full description on the dataset page: https://huggingface.co/datasets/huutho13254/saas-chatbot-v4.health-chatbot
Dataset Card for Dataset Name
Health Question and Answer Clean Dataset
Dataset Details
Dataset Description
This dataset provides a detailed overview of health question & answer pairs. It includes data on health problems and corresponding answers, making it suitable for variable tasks like healthcare chatbot training.
Language(s) (NLP): English
License: Apache-2.0
Dataset Sources [optional]
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/shaneperry0101/health-chatbot.customer_service_chatbotsiddha_vaithiyam_question_answering_chatbot
Medical Home Remedy Chatbot Dataset
Overview
This dataset is designed for a chatbot that answers questions related to medical problems with simple home remedies. The information in this dataset has been sourced from old books containing traditional remedies used in the past.
Contents
Dataset Files:
dataset.csv : The main dataset file containing questions and corresponding home remedy answers.
Data Structure:
Each row in the CSV file… See the full description on the dataset page: https://huggingface.co/datasets/RahulS3/siddha_vaithiyam_question_answering_chatbot.mental_health_Chatbot
Amod/mental_health_counseling_conversations
This dataset is a compilation of high-quality, real one-on-one mental health counseling conversations between individuals and licensed professionals. Each exchange is structured as a clear question–answer pair, making it directly suitable for fine-tuning or instruction-tuning language models that need to handle sensitive, empathetic, and contextually aware dialogue.
Since its public release in 2023, it has been downloaded over 100,000… See the full description on the dataset page: https://huggingface.co/datasets/Iamzoo/mental_health_Chatbot.medical_chatbot_datasetMental_Health_Support_ChatBOT_Conversation
Mental Health Support Dataset
Instruction–response pairs for training supportive, non-diagnostic,
safety-aware mental health chatbots.
Fields
instruction: user message
response: Bot reposne
category: intent label
Safety
This dataset includes crisis escalation examples and refusal patterns.
Not a replacement for professional care.
chatbotbanking-chatbot-enquiriesKhudra_Chatbot
Dubai Property AI Dataset
Description
Fine-tuning dataset for Dubai property management AI, covering:
Emergency alerts (water leaks, AC failures)
Financial optimization
DEWA/RERA compliance
Vendor dispatch protocols
Usage
from datasets import load_dataset
dataset = load_dataset("your-username/dubai-property-ai")
fitness-chatbot-dataset
🏋️♂️ Fitness Chatbot Space
💬 Chatbot em português especializado em treino, nutrição e bem-estar físico, com respostas naturais e ajustáveis via parâmetros interativos.
🚀 Demonstração
🔗 Acesse o Space: fitness-chatbot-space📸 Interface estilo chat com histórico de conversa, sliders de controle e exemplos prontos para testar.
🧠 Modelo Utilizado
Nome
Tipo
Idioma
Base
wpbcpaz/fitness-chatbot-model
Causal LM
Português
Ajustado com… See the full description on the dataset page: https://huggingface.co/datasets/wpbcpaz/fitness-chatbot-dataset.mental_health_chatbot_dataset
Dataset Card for "heliosbrahma/mental_health_chatbot_dataset"
Dataset Description
Dataset Summary
This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters.
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/manishaNeura/mental_health_chatbot_dataset.data-science-chatbot
📊 Data Science Chatbot Dataset (2000 Samples)
🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts.
This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way.
🎯 Objective
The goal of this dataset is to:
Train LLMs to act as a Data Science Tutor
Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.Ecom-Chatbot-Test-Set
Ecom Chatbot Synthetic Test Set
A 2,000-sample fully synthetic test set for evaluating e-commerce chatbot models fine-tuned on
rescommons/Ecom-Chatbot-Finetuning-Dataset.
Designed for zero-contamination evaluation — all products, orders, customer names, and responses
are synthetically generated and do not overlap with the training data.
Dataset Summary
Split
Samples
test
2,000
Group Distribution
Group
Count
Description
A
667… See the full description on the dataset page: https://huggingface.co/datasets/V1rtucious/Ecom-Chatbot-Test-Set.smolified-mental-health-chatbot
🤏 smolified-mental-health-chatbot
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-mental-health-chatbot.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 8f18325b)
Records: 9686
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
pashto-restaurant-chatbot
🇦🇫 Pashto Restaurant Chatbot Training Dataset (With Love for Pashto AI)
بیا رغونه او پښتو ژبې ته ځانګړې پاملرنه؛ د پښتو مصنوعي ځیرکتیا (Pashto AI) د بډاینې او ودې لپاره په مینه چمتو شوی کڅوړه.
This dataset is a high-quality, state-resilient Pashto translation of the widely used bitext/Bitext-restaurants-llm-chatbot-training-dataset. It contains approximately 30,000 conversational instruction-response pairs meticulously optimized for domain-specific fine-tuning in the hospitality… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-restaurant-chatbot.chatbotmental_health_chatbot_dataset
Dataset Card for "heliosbrahma/mental_health_chatbot_dataset"
Dataset Description
Dataset Summary
This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters.
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/VedxntR18/mental_health_chatbot_dataset.medical_chatbot_datasetTrix-Chatbot-Prompt-Response
Dataset Creation Process
Overview
This dataset was created to train and evaluate a chatbot focused on answering questions about Pooria Roy, his background, projects, and related topics. The goal was to build a dataset grounded in real user behavior while maintaining sufficient diversity and coverage of edge cases.
The final dataset contains 2,105 prompt-response examples, including a small portion of multi-turn conversations.
Data Collection Pipeline… See the full description on the dataset page: https://huggingface.co/datasets/regularpooria/Trix-Chatbot-Prompt-Response.pashto-chatbot-greetings
Pashto Chatbot Greetings Dataset
Multi-turn chatbot greeting configurations translated from Italian into Pashto (ps).
Details
Total Records: 8,846
Format: Streaming JSON Lines (.jsonl)
Maintained by: nassimjp
ai-hindi-chatbot
