CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K194 likes7.4k downloads2y agoHugging Face02ruslanmv /ai-medical-chatbot AI Medical Chatbot Dataset This is an experimental Dataset designed to run a Medical Chatbot It contains at least 250k dialogues between a Patient and a Doctor. Playground ChatBot ruslanmv/AI-Medical-Chatbot For furter information visit the project here: https://github.com/ruslanmv/ai-medical-chatbot text100K<n<1M253 likes4.8k downloads2y agoHugging Face03lmsys /chatbot_arena_conversationsgated Chatbot Arena Conversations Dataset This dataset contains 33K cleaned conversations with pairwise human preferences. It is collected from 13K unique IP addresses on the Chatbot Arena from April to June 2023. Each sample includes a question ID, two model names, their full conversation text in OpenAI API JSON format, the user vote, the anonymized user ID, the detected language tag, the OpenAI moderation API tag, the additional toxic tag, and the timestamp. To ensure the safe release… See the full description on the dataset page: https://huggingface.co/datasets/lmsys/chatbot_arena_conversations.tabular10K<n<100K490 likes2.4k downloads3y agoHugging Face04sayakpaul /diffusers-qa-chatbot-artifactstext100K<n<1M2 likes2.3k downloads3y agoHugging Face05bitext /Bitext-retail-ecommerce-llm-chatbot-training-dataset Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.textquestion-answering10K<n<100K19 likes1.3k downloads2y agoHugging Face06alespalla /chatbot_instruction_prompts Dataset Card for Chatbot Instruction Prompts Datasets Dataset Summary This dataset has been generated from the following ones: tatsu-lab/alpaca Dahoas/instruct-human-assistant-prompt allenai/prosocial-dialog The datasets has been cleaned up of spurious entries and artifacts. It contains ~500k of prompt and expected resposne. This DB is intended to train an instruct-type model textquestion-answering100K<n<1M64 likes979 downloads2y agoHugging Face07reshabhs /SPML_Chatbot_Prompt_Injection SPML Chatbot Prompt Injection Dataset Arxiv Paper Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/reshabhs/SPML_Chatbot_Prompt_Injection.tabulartext-classification10K<n<100K31 likes962 downloads2y agoHugging Face08bitext /Bitext-events-ticketing-llm-chatbot-training-dataset Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes779 downloads2y agoHugging Face09mathewhe /chatbot-arena-elo LMSYS Chatbot Arena ELO Scores This dataset is a datasets-friendly version of Chatbot Arena ELO scores, updated daily from the leaderboard API at https://huggingface.co/spaces/lmarena-ai/chatbot-arena-leaderboard. Updated: 20250717 Loading Data from datasets import load_dataset dataset = load_dataset("mathewhe/chatbot-arena-elo", split="train") The main branch of this dataset will always be updated to the latest ELO and leaderboard version. If you need a fixed dataset… See the full description on the dataset page: https://huggingface.co/datasets/mathewhe/chatbot-arena-elo.documentn<1K4 likes640 downloads1y agoHugging Face10agie-ai /lmsys-chatbot_arena_conversations Dataset Card for "lmsys-chatbot_arena_conversations" More Information needed tabular10K<n<100K0 likes462 downloads3y agoHugging Face11breadlicker45 /Bread-chatbot-dataset-test Dataset Card for "Bread-chatbot-dataset-test" More Information needed texttext-generation1M<n<10M0 likes371 downloads3y agoHugging Face12heliosbrahma /mental_health_chatbot_dataset Dataset Card for "heliosbrahma/mental_health_chatbot_dataset" Dataset Description Dataset Summary This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters. Languages The… See the full description on the dataset page: https://huggingface.co/datasets/heliosbrahma/mental_health_chatbot_dataset.texttext-generationn<1K94 likes355 downloads3y agoHugging Face13bitext /Bitext-retail-banking-llm-chatbot-training-dataset Bitext - Retail Banking Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail Banking] sector can be easily achieved using our two-step approach to LLM Fine-Tuning.… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-banking-llm-chatbot-training-dataset.textquestion-answering10K<n<100K16 likes331 downloads2y agoHugging Face14NajahUniv /arabic-univeristy-chatbot-qa Arabic University Chatbot QA A multilingual, multi-label intent-routing dataset for a university chatbot: given a student's message, predict which of 20 intent categories it should route to. This is routing, not question answering — the dataset contains no answers. Release v0.8.0 — pinned as a Hub tag, so revision="v0.8.0" always resolves to exactly these rows. This release holds 50,000 question rows in 22,600 scenario groups. Every row has accepted == true; the classifier… See the full description on the dataset page: https://huggingface.co/datasets/NajahUniv/arabic-univeristy-chatbot-qa.tabulartext-classification10K<n<100K0 likes293 downloads19d agoHugging Face15balade /chatbot-assetsdocumentn<1K0 likes253 downloads2mo agoHugging Face16bitext /Bitext-insurance-llm-chatbot-training-dataset Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.textquestion-answering10K<n<100K8 likes215 downloads2y agoHugging Face17bitext /Bitext-telco-llm-chatbot-training-dataset Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes201 downloads2y agoHugging Face18ANISH-j /chatbot-smalltext1K<n<10K0 likes187 downloads2y agoHugging Face19FiendHunter /Financial_chatbottextn<1K0 likes186 downloads2y agoHugging Face20bitext /Bitext-travel-llm-chatbot-training-dataset Bitext - Travel Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Travel] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-travel-llm-chatbot-training-dataset.textquestion-answering10K<n<100K4 likes184 downloads2y agoHugging Face21hirundo-io /spml-chatbot-prompt-injection-malicious-refusalstext10K<n<100K0 likes183 downloads2mo agoHugging Face22kdercksen /medical-patient-chatbot-conversations0 likes177 downloads3y agoHugging Face23scythe327 /chatbot-datatabular10K<n<100K0 likes157 downloads8d agoHugging Face24rescommons /Full-Ecom-Chatbot-Dataset E-commerce Chatbot Training Data A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains. Dataset Summary Split Records Train 35,213 Test 8,818 Total 44,031 The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.tabularquestion-answering10K<n<100K0 likes148 downloads6mo agoHugging Face25typhoon-ai /chatbot-arena-spoken-voicesaudio1K<n<10K0 likes135 downloads2y agoHugging Face26aigrant /tw_chatbot_arena TW Chatbot Arena 資料集說明 概述 TW Chatbot Arena 資料集是一個開源資料集,旨在促進台灣聊天機器人競技場 https://arena.twllm.com/ 的人類回饋強化學習資料(RLHF)。這個資料集包含英文和中文的對話資料,主要聚焦於繁體中文,以支援語言模型的開發和評估。 資料集摘要 授權: Apache-2.0 語言: 主要為繁體中文 規模: 3.6k 筆資料(2024/08/02) 內容: 使用者與聊天機器人的互動,每筆互動都根據回應品質標記為被選擇或被拒絕。 贊助 本計畫由「【g0v 零時小學校】繁體中文AI 開源實踐計畫」(https://sch001.g0v.tw/dash/brd/2024TC-AI-OS-Grant/list)贊助。 資料集結構 資料集包含以下欄位: question_id: 每次互動的唯一隨機識別碼。 model_a: 左側模型的名稱。 model_b: 右側模型的名稱。 winner:… See the full description on the dataset page: https://huggingface.co/datasets/aigrant/tw_chatbot_arena.tabular10K<n<100K18 likes129 downloads1y agoHugging Face27Naholav /cukurova_university_chatbot Çukurova University Computer Engineering Chatbot Dataset 📊 Dataset Overview This dataset contains 22,524 high-quality question-answer pairs specifically designed for training an AI chatbot that serves the Computer Engineering Department at Çukurova University. The dataset is part of the CengBot project, a sophisticated multilingual Telegram chatbot that provides automated assistance to students regarding courses, programs, and departmental information. 🔢… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/cukurova_university_chatbot.textquestion-answering10K<n<100K0 likes125 downloads1y agoHugging Face28Jannchie /lmsys_chatbot_arena_conversationsdatasource: https://colab.research.google.com/drive/1KdwokPjirkTmpO_P1WByFNFiqxWQquwH tabular1M<n<10M0 likes114 downloads2y agoHugging Face29feecha /chatbot_zh_datasettext1M<n<10M0 likes112 downloads2y agoHugging Face30dim /lmsys_chatbot_arena_conversations Dataset Card for "lmsys_chatbot_arena_conversations" More Information needed tabular10K<n<100K0 likes103 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.