datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-medical-chatbot
AI Medical Chatbot Dataset
This is an experimental Dataset designed to run a Medical Chatbot
It contains at least 250k dialogues between a Patient and a Doctor.
Playground ChatBot
ruslanmv/AI-Medical-Chatbot
For furter information visit the project here:
https://github.com/ruslanmv/ai-medical-chatbot
chatbot_arena_conversations
Chatbot Arena Conversations Dataset
This dataset contains 33K cleaned conversations with pairwise human preferences.
It is collected from 13K unique IP addresses on the Chatbot Arena from April to June 2023.
Each sample includes a question ID, two model names, their full conversation text in OpenAI API JSON format, the user vote, the anonymized user ID, the detected language tag, the OpenAI moderation API tag, the additional toxic tag, and the timestamp.
To ensure the safe release… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/chatbot_arena_conversations.chatbot_instruction_prompts
Dataset Card for Chatbot Instruction Prompts Datasets
Dataset Summary
This dataset has been generated from the following ones:
tatsu-lab/alpaca
Dahoas/instruct-human-assistant-prompt
allenai/prosocial-dialog
The datasets has been cleaned up of spurious entries and artifacts. It contains ~500k of prompt and expected resposne. This DB is intended to train an instruct-type model
lmsys-chatbot_arena_conversations
Dataset Card for "lmsys-chatbot_arena_conversations"
More Information needed
Bread-chatbot-dataset-test
Dataset Card for "Bread-chatbot-dataset-test"
More Information needed
mental_health_chatbot_dataset
Dataset Card for "heliosbrahma/mental_health_chatbot_dataset"
Dataset Description
Dataset Summary
This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters.
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/heliosbrahma/mental_health_chatbot_dataset.Bitext-retail-banking-llm-chatbot-training-dataset
Bitext - Retail Banking Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail Banking] sector can be easily achieved using our two-step approach to LLM Fine-Tuning.… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-banking-llm-chatbot-training-dataset.arabic-univeristy-chatbot-qa
Arabic University Chatbot QA
A multilingual, multi-label intent-routing dataset for a university chatbot: given a student's
message, predict which of 20 intent categories it should route to. This is
routing, not question answering — the dataset contains no answers.
Release v0.8.0 — pinned as a Hub tag, so revision="v0.8.0" always resolves to exactly these rows.
This release holds 50,000 question rows in 22,600 scenario
groups. Every row has accepted == true; the classifier… See the full description on the dataset page: https://huggingface.co/datasets/NajahUniv/arabic-univeristy-chatbot-qa.spml-chatbot-prompt-injection-malicious-refusalschatbot-datachatbot-arena-spoken-voicesMedical-ChatBot-DPOchatbot_zh_datasettw_chatbot_arena
TW Chatbot Arena 資料集說明
概述
TW Chatbot Arena 資料集是一個開源資料集,旨在促進台灣聊天機器人競技場 https://arena.twllm.com/ 的人類回饋強化學習資料(RLHF)。這個資料集包含英文和中文的對話資料,主要聚焦於繁體中文,以支援語言模型的開發和評估。
資料集摘要
授權: Apache-2.0
語言: 主要為繁體中文
規模: 3.6k 筆資料(2024/08/02)
內容: 使用者與聊天機器人的互動,每筆互動都根據回應品質標記為被選擇或被拒絕。
贊助
本計畫由「【g0v 零時小學校】繁體中文AI 開源實踐計畫」(https://sch001.g0v.tw/dash/brd/2024TC-AI-OS-Grant/list)贊助。
資料集結構
資料集包含以下欄位:
question_id: 每次互動的唯一隨機識別碼。
model_a: 左側模型的名稱。
model_b: 右側模型的名稱。
winner:… See the full description on the dataset page: https://huggingface.co/datasets/aigrant/tw_chatbot_arena.Chatbot-Url
Dataset Card for Wikimedia Wikipedia
Dataset Summary
Wikipedia dataset containing cleaned articles of all languages.
The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/)
with one subset per language, each containing a single train split.
Each example contains the content of one full Wikipedia article with cleaning to strip
markdown and unwanted sections (references, etc.).
All language subsets have already been processed for recent dump… See the full description on the dataset page: https://huggingface.co/datasets/zabr946/Chatbot-Url.chatbot_arena_completionstalking-to-chatbots-unwrapped-chatsThis work-in-progress dataset contains conversations with various LLM tools, sourced by the author of the website Talking to Chatbots.
A simplified version of this dataset can be found at reddgr/talking-to-chatbots-chats, where messages belonging to a same conversation are 'wrapped' inside a single record. In this extended dataset, each conversation turn (pair of messages consisting of a user prompt and a response by the LLM) is presented as an individual record, with additional metrics and… See the full description on the dataset page: https://huggingface.co/datasets/reddgr/talking-to-chatbots-unwrapped-chats.lmsys_chatbot_arena_conversations
Dataset Card for "lmsys_chatbot_arena_conversations"
More Information needed
chatbot_emotion챗봇 학습용 문답 페어 11,876개로 구성되었습니다.
https://github.com/songys/Chatbot_data
dataset_info:
features:
- name: index
dtype: int64
- name: Q
dtype: string
- name: A
dtype: string
splits:
- name: train
num_bytes: 773618
num_examples: 9465
- name: test
num_bytes: 246115
num_examples: 2358
download_size: 557106
dataset_size: 1019733
Dataset Card for "chatbot_emotion"
More Information needed
Ecom-Chatbot-Finetuning-Dataset
Ecom Chatbot Fine-Tuning Dataset
A unified e-commerce chatbot fine-tuning dataset combining 5 source datasets (40,098 examples total), covering product discovery, order management, customer support, returns, and more.
Splits
Split
Source
Examples
amazon_meta
Amazon product metadata
5,000
amazon_reviews
Amazon product reviews
23,100
asos_ecom_dataset
ASOS fashion e-commerce
2,000
bitext_customer_support
Bitext customer support (placeholder-free)
5,000… See the full description on the dataset page: https://huggingface.co/datasets/V1rtucious/Ecom-Chatbot-Finetuning-Dataset.Full-Ecom-Chatbot-Dataset
E-commerce Chatbot Training Data
A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains.
Dataset Summary
Split
Records
Train
35,213
Test
8,818
Total
44,031
The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.speech-chatbot-alpaca-eval
This dataset only contains test data, which is integrated into UltraEval-Audio(https://github.com/OpenBMB/UltraEval-Audio) framework.
python audio_evals/main.py --dataset speech-chatbot-alpaca-eval --model gpt4o_speech
🚀超凡体验,尽在UltraEval-Audio🚀
UltraEval-Audio——全球首个同时支持语音理解和语音生成评估的开源框架,专为语音大模型评估打造,集合了34项权威Benchmark,覆盖语音、声音、医疗及音乐四大领域,支持十种语言,涵盖十二类任务。选择UltraEval-Audio,您将体验到前所未有的便捷与高效:
一键式基准管理 📥:告别繁琐的手动下载与数据处理,UltraEval-Audio为您自动化完成这一切,轻松获取所需基准测试数据。
内置评估利器… See the full description on the dataset page: https://huggingface.co/datasets/TwinkStart/speech-chatbot-alpaca-eval.patient_doctor_chatbotEcommerce-dataset-chatbotMedical-ChatBot-DPO
Medical-ChatBot-DPO 数据集
数据集概述
本数据集是一个用于 DPO (Direct Preference Optimization) 训练的偏好对齐数据集,专门为医疗对话机器人设计。数据集包含 40,672 条样本,融合了通用对话安全性、人类偏好对齐和医疗领域专业知识。
数据来源与处理
1. Anthropic/hh-rlhf (harmless-base)
数据量: 10,000 条
来源: Anthropic/hh-rlhf
子集: harmless-base (无害对话子集)
处理方式:
从原始对话中提取最后一轮 Human-Assistant 对话
从 chosen 字段提取最后的 Assistant 回复作为 chosen(安全的拒绝回复)
从 rejected 字段提取最后的 Assistant 回复作为 rejected(有帮助但可能有害的回复)
过滤掉无效样本(prompt 为空的样本)
随机采样 10,000 条(seed=42)
用途:… See the full description on the dataset page: https://huggingface.co/datasets/bootscoder/Medical-ChatBot-DPO.Medical-ChatBot-SFTchatbot-arena-conversations-Embeddings
Chatbot Arena Conversations Embeddings
Embeddings of agie-ai/lmsys-chatbot_arena_conversations, produced with amkdg/Qwen3-Embedding-8B-NVFP4 — 4096-d,
L2-normalized float16 (cosine = dot product).
65,960 conversations → 65,960 vectors
emb.npy — float16 [65960, 4096]
meta.parquet — one row per vector, aligned with emb.npy: id, uuid, tag, chunk, n_chunks, count, source_ref
manifest.json — counts and provenance
Usage
import numpy as np, pyarrow.parquet as pq
emb… See the full description on the dataset page: https://huggingface.co/datasets/amkdg/chatbot-arena-conversations-Embeddings.chatbotarena-spoken-all-7824
ChatbotArena-Spoken Dataset
Based on ChatbotArena, we employ GPT-4o-mini to select dialogue turns that are well-suited to spoken-conversation analysis, yielding 7824 data points. To obtain audio, we synthesize every utterance, user prompt and model responses, using one of 12 voices from KokoroTTS (v0.19), chosen uniformly at random. Because the original human annotations assess only lexical content in text, we keep labels unchanged and treat them as ground truth for the spoken… See the full description on the dataset page: https://huggingface.co/datasets/potsawee/chatbotarena-spoken-all-7824.talking-to-chatbots-chatsThis work-in-progress dataset contains conversations with various LLM tools, sourced by the author of the website Talking to Chatbots.
The format chosen for structuring this dataset is similar to that of lmsys/lmsys-chat-1m.
Conversations are identified by a UUID (v4) and 'wrapped' in a JSON format where each message is contained in the 'content' key. The 'role' key identifies whether the message is a prompt ('user') or a response by the LLM ('assistant'). For each dictionary, 'turn'… See the full description on the dataset page: https://huggingface.co/datasets/reddgr/talking-to-chatbots-chats.Ecom-Chatbot-Finetuning-Dataset
Ecom Chatbot Finetuning Dataset
A unified instruction-following dataset for fine-tuning e-commerce customer service chatbots. It covers a wide range of real-world retail scenarios — from product discovery and order management to returns, complaints, and account support.
Dataset Summary
Field
Value
Total records
40,098
Language
English
Sources
Amazon Reviews 2023, Amazon Meta 2023, ASOS, Bitext
Response types
Text, Tool Call, Mixed
Difficulty levels
1… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Ecom-Chatbot-Finetuning-Dataset.
