CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ruslanmv /ai-medical-chatbot AI Medical Chatbot Dataset This is an experimental Dataset designed to run a Medical Chatbot It contains at least 250k dialogues between a Patient and a Doctor. Playground ChatBot ruslanmv/AI-Medical-Chatbot For furter information visit the project here: https://github.com/ruslanmv/ai-medical-chatbot text100K<n<1M253 likes4.9k downloads3y agoHugging Face02lmsys /chatbot_arena_conversationsgated Chatbot Arena Conversations Dataset This dataset contains 33K cleaned conversations with pairwise human preferences. It is collected from 13K unique IP addresses on the Chatbot Arena from April to June 2023. Each sample includes a question ID, two model names, their full conversation text in OpenAI API JSON format, the user vote, the anonymized user ID, the detected language tag, the OpenAI moderation API tag, the additional toxic tag, and the timestamp. To ensure the safe release… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/chatbot_arena_conversations.tabular10K<n<100K491 likes2.5k downloads3y agoHugging Face03alespalla /chatbot_instruction_prompts Dataset Card for Chatbot Instruction Prompts Datasets Dataset Summary This dataset has been generated from the following ones: tatsu-lab/alpaca Dahoas/instruct-human-assistant-prompt allenai/prosocial-dialog The datasets has been cleaned up of spurious entries and artifacts. It contains ~500k of prompt and expected resposne. This DB is intended to train an instruct-type model textquestion-answering100K<n<1M64 likes911 downloads2y agoHugging Face04agie-ai /lmsys-chatbot_arena_conversations Dataset Card for "lmsys-chatbot_arena_conversations" More Information needed tabular10K<n<100K0 likes449 downloads3y agoHugging Face05breadlicker45 /Bread-chatbot-dataset-test Dataset Card for "Bread-chatbot-dataset-test" More Information needed texttext-generation1M<n<10M0 likes401 downloads3y agoHugging Face06heliosbrahma /mental_health_chatbot_dataset Dataset Card for "heliosbrahma/mental_health_chatbot_dataset" Dataset Description Dataset Summary This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters. Languages The… See the full description on the dataset page: https://huggingface.co/datasets/heliosbrahma/mental_health_chatbot_dataset.texttext-generationn<1K94 likes356 downloads3y agoHugging Face07bitext /Bitext-retail-banking-llm-chatbot-training-dataset Bitext - Retail Banking Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail Banking] sector can be easily achieved using our two-step approach to LLM Fine-Tuning.… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-banking-llm-chatbot-training-dataset.textquestion-answering10K<n<100K17 likes338 downloads2y agoHugging Face08NajahUniv /arabic-univeristy-chatbot-qa Arabic University Chatbot QA A multilingual, multi-label intent-routing dataset for a university chatbot: given a student's message, predict which of 20 intent categories it should route to. This is routing, not question answering — the dataset contains no answers. Release v0.8.0 — pinned as a Hub tag, so revision="v0.8.0" always resolves to exactly these rows. This release holds 50,000 question rows in 22,600 scenario groups. Every row has accepted == true; the classifier… See the full description on the dataset page: https://huggingface.co/datasets/NajahUniv/arabic-univeristy-chatbot-qa.tabulartext-classification10K<n<100K0 likes301 downloads23d agoHugging Face09hirundo-io /spml-chatbot-prompt-injection-malicious-refusalstext10K<n<100K0 likes218 downloads3mo agoHugging Face10scythe327 /chatbot-datatabular10K<n<100K0 likes182 downloads4d agoHugging Face11typhoon-ai /chatbot-arena-spoken-voicesaudio1K<n<10K0 likes139 downloads2y agoHugging Face12Julian2002 /Medical-ChatBot-DPOtext100K<n<1M2 likes134 downloads2y agoHugging Face13feecha /chatbot_zh_datasettext1M<n<10M0 likes133 downloads2y agoHugging Face14aigrant /tw_chatbot_arena TW Chatbot Arena 資料集說明 概述 TW Chatbot Arena 資料集是一個開源資料集,旨在促進台灣聊天機器人競技場 https://arena.twllm.com/ 的人類回饋強化學習資料(RLHF)。這個資料集包含英文和中文的對話資料,主要聚焦於繁體中文,以支援語言模型的開發和評估。 資料集摘要 授權: Apache-2.0 語言: 主要為繁體中文 規模: 3.6k 筆資料(2024/08/02) 內容: 使用者與聊天機器人的互動,每筆互動都根據回應品質標記為被選擇或被拒絕。 贊助 本計畫由「【g0v 零時小學校】繁體中文AI 開源實踐計畫」(https://sch001.g0v.tw/dash/brd/2024TC-AI-OS-Grant/list)贊助。 資料集結構 資料集包含以下欄位: question_id: 每次互動的唯一隨機識別碼。 model_a: 左側模型的名稱。 model_b: 右側模型的名稱。 winner:… See the full description on the dataset page: https://huggingface.co/datasets/aigrant/tw_chatbot_arena.tabular10K<n<100K18 likes129 downloads1y agoHugging Face15zabr946 /Chatbot-Url Dataset Card for Wikimedia Wikipedia Dataset Summary Wikipedia dataset containing cleaned articles of all languages. The dataset is built from the Wikipedia dumps (https://dumps.wikimedia.org/) with one subset per language, each containing a single train split. Each example contains the content of one full Wikipedia article with cleaning to strip markdown and unwanted sections (references, etc.). All language subsets have already been processed for recent dump… See the full description on the dataset page: https://huggingface.co/datasets/zabr946/Chatbot-Url.texttext-generation10M<n<100M0 likes119 downloads4mo agoHugging Face16trl-lib /chatbot_arena_completionstext10K<n<100K4 likes116 downloads1y agoHugging Face17reddgr /talking-to-chatbots-unwrapped-chatsThis work-in-progress dataset contains conversations with various LLM tools, sourced by the author of the website Talking to Chatbots. A simplified version of this dataset can be found at reddgr/talking-to-chatbots-chats, where messages belonging to a same conversation are 'wrapped' inside a single record. In this extended dataset, each conversation turn (pair of messages consisting of a user prompt and a response by the LLM) is presented as an individual record, with additional metrics and… See the full description on the dataset page: https://huggingface.co/datasets/reddgr/talking-to-chatbots-unwrapped-chats.tabular10K<n<100K2 likes99 downloads2y agoHugging Face18dim /lmsys_chatbot_arena_conversations Dataset Card for "lmsys_chatbot_arena_conversations" More Information needed tabular10K<n<100K0 likes98 downloads3y agoHugging Face19jeongah /chatbot_emotion챗봇 학습용 문답 페어 11,876개로 구성되었습니다. https://github.com/songys/Chatbot_data dataset_info: features: - name: index dtype: int64 - name: Q dtype: string - name: A dtype: string splits: - name: train num_bytes: 773618 num_examples: 9465 - name: test num_bytes: 246115 num_examples: 2358 download_size: 557106 dataset_size: 1019733 Dataset Card for "chatbot_emotion" More Information needed text10K<n<100K5 likes87 downloads4y agoHugging Face20V1rtucious /Ecom-Chatbot-Finetuning-Dataset Ecom Chatbot Fine-Tuning Dataset A unified e-commerce chatbot fine-tuning dataset combining 5 source datasets (40,098 examples total), covering product discovery, order management, customer support, returns, and more. Splits Split Source Examples amazon_meta Amazon product metadata 5,000 amazon_reviews Amazon product reviews 23,100 asos_ecom_dataset ASOS fashion e-commerce 2,000 bitext_customer_support Bitext customer support (placeholder-free) 5,000… See the full description on the dataset page: https://huggingface.co/datasets/V1rtucious/Ecom-Chatbot-Finetuning-Dataset.tabular10K<n<100K0 likes75 downloads6mo agoHugging Face21rescommons /Full-Ecom-Chatbot-Dataset E-commerce Chatbot Training Data A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains. Dataset Summary Split Records Train 35,213 Test 8,818 Total 44,031 The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.tabularquestion-answering10K<n<100K0 likes72 downloads6mo agoHugging Face22TwinkStart /speech-chatbot-alpaca-eval This dataset only contains test data, which is integrated into UltraEval-Audio(https://github.com/OpenBMB/UltraEval-Audio) framework. python audio_evals/main.py --dataset speech-chatbot-alpaca-eval --model gpt4o_speech 🚀超凡体验,尽在UltraEval-Audio🚀 UltraEval-Audio——全球首个同时支持语音理解和语音生成评估的开源框架,专为语音大模型评估打造,集合了34项权威Benchmark,覆盖语音、声音、医疗及音乐四大领域,支持十种语言,涵盖十二类任务。选择UltraEval-Audio,您将体验到前所未有的便捷与高效: 一键式基准管理 📥:告别繁琐的手动下载与数据处理,UltraEval-Audio为您自动化完成这一切,轻松获取所需基准测试数据。 内置评估利器… See the full description on the dataset page: https://huggingface.co/datasets/TwinkStart/speech-chatbot-alpaca-eval.audion<1K0 likes58 downloads2y agoHugging Face23Sohaibsoussi /patient_doctor_chatbottext100K<n<1M4 likes53 downloads2y agoHugging Face24MartinDang /Ecommerce-dataset-chatbottext10K<n<100K1 likes52 downloads2y agoHugging Face25bootscoder /Medical-ChatBot-DPO Medical-ChatBot-DPO 数据集 数据集概述 本数据集是一个用于 DPO (Direct Preference Optimization) 训练的偏好对齐数据集,专门为医疗对话机器人设计。数据集包含 40,672 条样本,融合了通用对话安全性、人类偏好对齐和医疗领域专业知识。 数据来源与处理 1. Anthropic/hh-rlhf (harmless-base) 数据量: 10,000 条 来源: Anthropic/hh-rlhf 子集: harmless-base (无害对话子集) 处理方式: 从原始对话中提取最后一轮 Human-Assistant 对话 从 chosen 字段提取最后的 Assistant 回复作为 chosen(安全的拒绝回复) 从 rejected 字段提取最后的 Assistant 回复作为 rejected(有帮助但可能有害的回复) 过滤掉无效样本(prompt 为空的样本) 随机采样 10,000 条(seed=42) 用途:… See the full description on the dataset page: https://huggingface.co/datasets/bootscoder/Medical-ChatBot-DPO.text10K<n<100K1 likes52 downloads11mo agoHugging Face26Julian2002 /Medical-ChatBot-SFTtext100K<n<1M0 likes51 downloads2y agoHugging Face27amkdg /chatbot-arena-conversations-Embeddings Chatbot Arena Conversations Embeddings Embeddings of agie-ai/lmsys-chatbot_arena_conversations, produced with amkdg/Qwen3-Embedding-8B-NVFP4 — 4096-d, L2-normalized float16 (cosine = dot product). 65,960 conversations → 65,960 vectors emb.npy — float16 [65960, 4096] meta.parquet — one row per vector, aligned with emb.npy: id, uuid, tag, chunk, n_chunks, count, source_ref manifest.json — counts and provenance Usage import numpy as np, pyarrow.parquet as pq emb… See the full description on the dataset page: https://huggingface.co/datasets/amkdg/chatbot-arena-conversations-Embeddings.tabular10K<n<100K0 likes49 downloads3mo agoHugging Face28potsawee /chatbotarena-spoken-all-7824 ChatbotArena-Spoken Dataset Based on ChatbotArena, we employ GPT-4o-mini to select dialogue turns that are well-suited to spoken-conversation analysis, yielding 7824 data points. To obtain audio, we synthesize every utterance, user prompt and model responses, using one of 12 voices from KokoroTTS (v0.19), chosen uniformly at random. Because the original human annotations assess only lexical content in text, we keep labels unchanged and treat them as ground truth for the spoken… See the full description on the dataset page: https://huggingface.co/datasets/potsawee/chatbotarena-spoken-all-7824.audio1K<n<10K0 likes47 downloads1y agoHugging Face29reddgr /talking-to-chatbots-chatsThis work-in-progress dataset contains conversations with various LLM tools, sourced by the author of the website Talking to Chatbots. The format chosen for structuring this dataset is similar to that of lmsys/lmsys-chat-1m. Conversations are identified by a UUID (v4) and 'wrapped' in a JSON format where each message is contained in the 'content' key. The 'role' key identifies whether the message is a prompt ('user') or a response by the LLM ('assistant'). For each dictionary, 'turn'… See the full description on the dataset page: https://huggingface.co/datasets/reddgr/talking-to-chatbots-chats.tabular1K<n<10K2 likes46 downloads2y agoHugging Face30rescommons /Ecom-Chatbot-Finetuning-Dataset Ecom Chatbot Finetuning Dataset A unified instruction-following dataset for fine-tuning e-commerce customer service chatbots. It covers a wide range of real-world retail scenarios — from product discovery and order management to returns, complaints, and account support. Dataset Summary Field Value Total records 40,098 Language English Sources Amazon Reviews 2023, Amazon Meta 2023, ASOS, Bitext Response types Text, Tool Call, Mixed Difficulty levels 1… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Ecom-Chatbot-Finetuning-Dataset.tabularquestion-answering10K<n<100K0 likes46 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.