datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chatbot_instruction_prompts
Dataset Card for Chatbot Instruction Prompts Datasets
Dataset Summary
This dataset has been generated from the following ones:
tatsu-lab/alpaca
Dahoas/instruct-human-assistant-prompt
allenai/prosocial-dialog
The datasets has been cleaned up of spurious entries and artifacts. It contains ~500k of prompt and expected resposne. This DB is intended to train an instruct-type model
Bread-chatbot-dataset-test
Dataset Card for "Bread-chatbot-dataset-test"
More Information needed
mental_health_chatbot_dataset
Dataset Card for "heliosbrahma/mental_health_chatbot_dataset"
Dataset Description
Dataset Summary
This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters.
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/heliosbrahma/mental_health_chatbot_dataset.Full-Ecom-Chatbot-Dataset
E-commerce Chatbot Training Data
A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains.
Dataset Summary
Split
Records
Train
35,213
Test
8,818
Total
44,031
The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.cukurova_university_chatbot
Çukurova University Computer Engineering Chatbot Dataset
📊 Dataset Overview
This dataset contains 22,524 high-quality question-answer pairs specifically designed for training an AI chatbot that serves the Computer Engineering Department at Çukurova University. The dataset is part of the CengBot project, a sophisticated multilingual Telegram chatbot that provides automated assistance to students regarding courses, programs, and departmental information.
🔢… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/cukurova_university_chatbot.NBRO-Chatbot-V1
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: National Building Research Organisation (NBRO)
Funded by : NBRO
Language(s) (NLP): English (en)
License: Creative Commons Attribution 4.0 International (CC BY 4.0)
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More… See the full description on the dataset page: https://huggingface.co/datasets/nimesh7814/NBRO-Chatbot-V1.cje-chatbot-arena
CJE Chatbot Arena Dataset
Dataset from Causal Judge Evaluation experiments on Chatbot Arena data.
Dataset Structure
cje_dataset.jsonl - Main dataset with judge scores and oracle labels (4,961 prompts)
prompts.jsonl - Original Chatbot Arena prompts
responses/ - Model responses for each policy variant
logprobs/ - Token logprobs for importance sampling estimators
Policies
5 system prompt variants evaluated:
base - No system prompt
clone - "Respond exactly as… See the full description on the dataset page: https://huggingface.co/datasets/elandy/cje-chatbot-arena.Chatbot-Url
Dataset Card for Wikimedia Wikipedia
Dataset Summary
Wikipedia dataset containing cleaned articles of all languages.
The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/)
with one subset per language, each containing a single train split.
Each example contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).
All language subsets have already been processed for recent dump… See the full description on the dataset page: https://huggingface.co/datasets/zabr946/Chatbot-Url.llm-jp-chatbot-arena-conversations
LLM-jp Chatbot Arena Conversations Dataset
This dataset contains approximately 1,000 conversations with pairwise human preferences, most of which are in Japanese.
The data was collected during the trial phase of the LLM-jp Chatbot Arena (January–February 2025), where users compared responses from two different models in a head-to-head format.
Each sample includes a question ID, the names of the two models, their conversation transcripts, the user's vote, an anonymized user ID, a… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/llm-jp-chatbot-arena-conversations.customer_service_chatbotEcom-Chatbot-Finetuning-Dataset
Ecom Chatbot Finetuning Dataset
A unified instruction-following dataset for fine-tuning e-commerce customer service chatbots. It covers a wide range of real-world retail scenarios — from product discovery and order management to returns, complaints, and account support.
Dataset Summary
Field
Value
Total records
40,098
Language
English
Sources
Amazon Reviews 2023, Amazon Meta 2023, ASOS, Bitext
Response types
Text, Tool Call, Mixed
Difficulty levels
1… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Ecom-Chatbot-Finetuning-Dataset.saas-chatbot-v4
SaaS Chatbot V4 Dataset
Multi-industry, multilingual conversational dataset for fine-tuning LLMs as SaaS AI chatbot agents with tool calling.
Stats
Metric
Value
Train
4,043
Test
450
Total messages
64,645
Avg msgs/conv
14.4
Think blocks
29,345 (21% empty)
Tool calls
15,215
Tool responses
15,387
Industries (8)
E-commerce (1,301), Travel (641), Services (504), Food (490), Beauty (478), Healthcare (404), Education (357), Real Estate… See the full description on the dataset page: https://huggingface.co/datasets/huutho13254/saas-chatbot-v4.tenant-law-chicago-chatbot-conv
Dataset Card for Dataset Name
The dataset is collected from human-chatbot interaction on tenant law consultation in the Chicago area.
Dataset Details
health-chatbot
Dataset Card for Dataset Name
Health Question and Answer Clean Dataset
Dataset Details
Dataset Description
This dataset provides a detailed overview of health question & answer pairs. It includes data on health problems and corresponding answers, making it suitable for variable tasks like healthcare chatbot training.
Language(s) (NLP): English
License: Apache-2.0
Dataset Sources [optional]
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/shaneperry0101/health-chatbot.siddha_vaithiyam_question_answering_chatbot
Medical Home Remedy Chatbot Dataset
Overview
This dataset is designed for a chatbot that answers questions related to medical problems with simple home remedies. The information in this dataset has been sourced from old books containing traditional remedies used in the past.
Contents
Dataset Files:
dataset.csv : The main dataset file containing questions and corresponding home remedy answers.
Data Structure:
Each row in the CSV file… See the full description on the dataset page: https://huggingface.co/datasets/RahulS3/siddha_vaithiyam_question_answering_chatbot.mental_health_Chatbot
Amod/mental_health_counseling_conversations
This dataset is a compilation of high-quality, real one-on-one mental health counseling conversations between individuals and licensed professionals. Each exchange is structured as a clear question–answer pair, making it directly suitable for fine-tuning or instruction-tuning language models that need to handle sensitive, empathetic, and contextually aware dialogue.
Since its public release in 2023, it has been downloaded over 100,000… See the full description on the dataset page: https://huggingface.co/datasets/Iamzoo/mental_health_Chatbot.Mental_Health_Support_ChatBOT_Conversation
Mental Health Support Dataset
Instruction–response pairs for training supportive, non-diagnostic,
safety-aware mental health chatbots.
Fields
instruction: user message
response: Bot reposne
category: intent label
Safety
This dataset includes crisis escalation examples and refusal patterns.
Not a replacement for professional care.
medical_chatbot_datasetdemo_dataset_shareGPT_chatbot
888流感灵对话数据集使用说明
数据集概述
本数据集基于888流感灵产品资料手册内容,按照shareGPT格式创建,用于微调Qwen-7B-instruct模型。数据集包含81条多样化的对话,涵盖了产品咨询、症状询问、用药指导、售后服务等多种场景。
数据集特点
多种风格:包含正式、随意、专业等不同风格的对话
语言多样性:中英文混合,以中文为主
对话长度:包含短对话、长对话
对话结构:包含单轮对话、多轮对话和追问对话
内容全面:涵盖产品信息、用药指导、注意事项、公司背景等多方面内容
文件格式
数据集采用jsonl格式,完全兼容HuggingFace上传标准。每条对话的格式如下:
[
{"from": "human", "value": "用户问题"},
{"from": "gpt", "value": "助手回答"}
]
使用方法
上传至HuggingFace:
登录HuggingFace账户
创建新的数据集仓库… See the full description on the dataset page: https://huggingface.co/datasets/wangjiangcheng/demo_dataset_shareGPT_chatbot.chatbotmicrosoft-phi-3-5-mini-instruct__llm-inference-chatbot-short__019e3b5cf170
microsoft/Phi-3.5-mini-instruct on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
TTFT P50
18.4079
ms
TTFT P99
265.3733
ms
TPOT P50
4.6615
ms
TPOT P99
6.5718
ms
Total P50 Ms
531.3797
Total P99 Ms
786.9971
Req Per S Passing
6.5257
Req Per S All
6.6729
Compliance Rate
0.9779
Ok Rate
1
Throughput Tok Per S
716.0655
Power Avg W
861.866
Power Peak W
897.907
Energy Joules… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/microsoft-phi-3-5-mini-instruct__llm-inference-chatbot-short__019e3b5cf170.banking-chatbot-enquiriesdeepseek-ai-deepseek-coder-v2-lite-instruct__llm-inference-chatbot-short__019e3b6f22ca
deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
TTFT P50
74.4268
ms
TTFT P99
521.1922
ms
TPOT P50
21.0371
ms
TPOT P99
23.458
ms
Total P50 Ms
2661.4077
Total P99 Ms
3136.5972
Req Per S Passing
1.0172
Req Per S All
1.1559
Compliance Rate
0.88
Ok Rate
1
Throughput Tok Per S
134.316
Power Avg W
808.3085
Power Peak W
854.62… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/deepseek-ai-deepseek-coder-v2-lite-instruct__llm-inference-chatbot-short__019e3b6f22ca.google-gemma-2-9b-it__llm-inference-chatbot-short__019e3b973b4f
google/gemma-2-9b-it on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
TTFT P50
30.0542
ms
TTFT P99
183.9538
ms
TPOT P50
8.648
ms
TPOT P99
10.7976
ms
Total P50 Ms
1095.4794
Total P99 Ms
1422.5879
Req Per S Passing
3.3476
Req Per S All
3.3947
Compliance Rate
0.9861
Ok Rate
1
Throughput Tok Per S
385.0208
Power Avg W
901.0112
Power Peak W
937.794
Energy Joules… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/google-gemma-2-9b-it__llm-inference-chatbot-short__019e3b973b4f.meb-ogrenci-chatbot-benchmark
MEB Öğrenci Chatbotu Benchmark (v1)
Ortaokul (5-8. sınıf) öğrencisine ders anlatan bir chatbot modelini değerlendirmek için
hazırladığım özel benchmark veri seti. Toplam 101 soru ve her soru için bir referans
cevap içerir.
Amaç
Bu benchmark, genel amaçlı bir sınav (MMLU gibi) değildir. Kendi senaryoma özeldir:
model, bir ortaokul öğrencisinin sorduğu ders sorusuna, seviyesine uygun, doğru ve
açıklayıcı bir cevap verebiliyor mu?
İçerik
Toplam soru:… See the full description on the dataset page: https://huggingface.co/datasets/nursimakgul/meb-ogrenci-chatbot-benchmark.Ecom-Chatbot-Test-Set
Ecom Chatbot Synthetic Test Set
A 2,000-sample fully synthetic test set for evaluating e-commerce chatbot models fine-tuned on
rescommons/Ecom-Chatbot-Finetuning-Dataset.
Designed for zero-contamination evaluation — all products, orders, customer names, and responses
are synthetically generated and do not overlap with the training data.
Dataset Summary
Split
Samples
test
2,000
Group Distribution
Group
Count
Description
A
667… See the full description on the dataset page: https://huggingface.co/datasets/V1rtucious/Ecom-Chatbot-Test-Set.smolified-mental-health-chatbot
🤏 smolified-mental-health-chatbot
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-mental-health-chatbot.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 8f18325b)
Records: 9686
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
data-science-chatbot
📊 Data Science Chatbot Dataset (2000 Samples)
🚀 A high-quality instruction-style dataset designed for fine-tuning Large Language Models (LLMs) on Data Science concepts.
This dataset contains ~2000 curated question-answer pairs in ChatML format, enabling models to learn how to explain, define, and discuss core data science topics in a clear and beginner-friendly way.
🎯 Objective
The goal of this dataset is to:
Train LLMs to act as a Data Science Tutor
Provide clear… See the full description on the dataset page: https://huggingface.co/datasets/Hamzasajjad38/data-science-chatbot.mental_health_chatbot_dataset
Dataset Card for "heliosbrahma/mental_health_chatbot_dataset"
Dataset Description
Dataset Summary
This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters.
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/manishaNeura/mental_health_chatbot_dataset.qwen-qwen2-5-7b-instruct__llm-inference-chatbot-short__019e3b395d04
Qwen/Qwen2.5-7B-Instruct on llm.inference.chatbot-short (NVIDIA H100 80GB HBM3)
Back to leaderboard
Headline metrics
Metric
Value
Unit
TTFT P50
26.8358
ms
TTFT P99
52.9781
ms
TPOT P50
6.2672
ms
TPOT P99
6.504
ms
Total P50 Ms
822.2015
Total P99 Ms
929.0652
Req Per S Passing
4.585
Req Per S All
4.6332
Compliance Rate
0.9896
Ok Rate
1
Throughput Tok Per S
540.8831
Power Avg W
908.6194
Power Peak W
947.635
Energy Joules… See the full description on the dataset page: https://huggingface.co/datasets/Yobitel/qwen-qwen2-5-7b-instruct__llm-inference-chatbot-short__019e3b395d04.
