CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K195 likes7.8k downloads2y agoHugging Face02bitext /Bitext-retail-ecommerce-llm-chatbot-training-dataset Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.textquestion-answering10K<n<100K19 likes1.4k downloads2y agoHugging Face03reshabhs /SPML_Chatbot_Prompt_Injection SPML Chatbot Prompt Injection Dataset Arxiv Paper Introducing the SPML Chatbot Prompt Injection Dataset: a robust collection of system prompts designed to create realistic chatbot interactions, coupled with a diverse array of annotated user prompts that attempt to carry out prompt injection attacks. While other datasets in this domain have centered on less practical chatbot scenarios or have limited themselves to "jailbreaking" – just one aspect of prompt injection – our dataset… See the full description on the dataset page: https://huggingface.co/datasets/reshabhs/SPML_Chatbot_Prompt_Injection.tabulartext-classification10K<n<100K31 likes991 downloads2y agoHugging Face04bitext /Bitext-events-ticketing-llm-chatbot-training-dataset Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes881 downloads2y agoHugging Face05mathewhe /chatbot-arena-elo LMSYS Chatbot Arena ELO Scores This dataset is a datasets-friendly version of Chatbot Arena ELO scores, updated daily from the leaderboard API at https://huggingface.co/spaces/lmarena-ai/chatbot-arena-leaderboard. Updated: 20250717 Loading Data from datasets import load_dataset dataset = load_dataset("mathewhe/chatbot-arena-elo", split="train") The main branch of this dataset will always be updated to the latest ELO and leaderboard version. If you need a fixed dataset… See the full description on the dataset page: https://huggingface.co/datasets/mathewhe/chatbot-arena-elo.documentn<1K4 likes689 downloads1y agoHugging Face06bitext /Bitext-telco-llm-chatbot-training-dataset Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes259 downloads2y agoHugging Face07bitext /Bitext-insurance-llm-chatbot-training-dataset Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.textquestion-answering10K<n<100K8 likes215 downloads2y agoHugging Face08bitext /Bitext-travel-llm-chatbot-training-dataset Bitext - Travel Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Travel] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-travel-llm-chatbot-training-dataset.textquestion-answering10K<n<100K4 likes189 downloads2y agoHugging Face09ANISH-j /chatbot-smalltext1K<n<10K0 likes187 downloads2y agoHugging Face10Jannchie /lmsys_chatbot_arena_conversationsdatasource: https://colab.research.google.com/drive/1KdwokPjirkTmpO_P1WByFNFiqxWQquwH tabular1M<n<10M0 likes113 downloads2y agoHugging Face11bitext /Bitext-mortgage-loans-llm-chatbot-training-dataset Bitext - Mortgage and Loans Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Mortgage and Loans] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-mortgage-loans-llm-chatbot-training-dataset.textquestion-answering10K<n<100K5 likes98 downloads2y agoHugging Face12Pinchao /ChatBot_NFR Dataset de Requisitos No Funcionales para ChatBot_NFR Este es el dataset final utilizado en el entrenamiento del modelo ChatBot_NFR, el cual está diseñado para generar y procesar requisitos no funcionales en el ámbito de la ingeniería de software. Este conjunto de datos ha sido cuidadosamente curado, normalizado y ajustado para asegurar que los requisitos estén organizados de manera clara y útil para entrenar modelos de lenguaje masivo. Descripción del Dataset Este… See the full description on the dataset page: https://huggingface.co/datasets/Pinchao/ChatBot_NFR.text1K<n<10K1 likes87 downloads2y agoHugging Face13bitext /Bitext-wealth-management-llm-chatbot-training-dataset Bitext - Wealth Management Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Wealth Management] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-wealth-management-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes86 downloads2y agoHugging Face14amirayuyue /hpc-chatbot-logstextn<1K0 likes83 downloads8d agoHugging Face15bitext /Bitext-hospitality-llm-chatbot-training-dataset Bitext - Hospitality Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [hospitality] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-hospitality-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes73 downloads2y agoHugging Face16bitext /Bitext-media-llm-chatbot-training-dataset Bitext - Media Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [media] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-media-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes68 downloads2y agoHugging Face17rhesis /Insurance-ChatBot-TestBench-Sample Insurance ChatBot TestBench Dataset (Sample) Dataset Description: The dataset presented here includes 80 example prompts from the Insurance ChatBot TestBench, a specialized test set developed to evaluate the performance of generative AI chatbots in the insurance industry. These prompts are used in the analysis described in the blog post "Gen AI Chatbots in the Insurance Industry: Are they Trustworthy?". The test bench assesses chatbot performance across three critical dimensions:… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Insurance-ChatBot-TestBench-Sample.textquestion-answeringn<1K0 likes62 downloads2y agoHugging Face18bitext /Bitext-restaurants-llm-chatbot-training-dataset Bitext - Restaurants Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [restaurants] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-restaurants-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes54 downloads2y agoHugging Face19abhi23457 /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/abhi23457/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes48 downloads20d agoHugging Face20rhesis /Insurance-Chatbot-Homeowner-Fraud-Harmful Dataset Card for Homeowner Fraud Harmful Description The test set provided is designed for evaluating the performance and robustness of an Insurance Chatbot specifically tailored for the insurance industry. It focuses on detecting harmful behaviors, with a specific focus on topics related to homeowner fraud. The purpose of this test set is to thoroughly assess the chatbot's ability to accurately identify and handle fraudulent activities within the context of homeowner… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Insurance-Chatbot-Homeowner-Fraud-Harmful.textn<1K0 likes37 downloads2y agoHugging Face21chat-bot-dls /user_feedbacktabularn<1K1 likes36 downloads3y agoHugging Face22rhesis /European-E-commerce-Chatbot-Social-Norms-Toxic Dataset Card for Social Norms Toxic Description The test set is specifically designed for evaluating the performance of a European E-commerce Chatbot in the context of the E-commerce industry. The main focus of the evaluation lies on assessing the chatbot's behavior in terms of compliance with relevant regulations. Additionally, the test set covers various categories, with particular attention given to identifying and handling toxic content. Furthermore, the chatbot's… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/European-E-commerce-Chatbot-Social-Norms-Toxic.textn<1K0 likes36 downloads2y agoHugging Face23one-thing /chatbot_arena_conversations_hinglishThe dataset is created by translating "lmsys/chatbot_arena_conversations" dataset. link to original datset - https://huggingface.co/datasets/lmsys/chatbot_arena_conversations Original dataset contain two conversation from model_a and model_b and also given winner model between these two model conversation. I have selected winner conversation and converted that user query and assistant answer into hinglish language using Gemini pro text10K<n<100K5 likes35 downloads3y agoHugging Face24songys /ChatbotData Chatbot_data. Chatbot_data_for_Korean v1.0 Data description. 인공데이터입니다. 일부 이별과 관련된 질문에서 다음카페 "사랑보다 아름다운 실연( http://cafe116.daum.net/_c21_/home?grpid=1bld )"에서 자주 나오는 이야기들을 참고하여 제작하였습니다. 가령 "이별한 지 열흘(또는 100일) 되었어요"라는 질문에 챗봇이 위로한다는 취지로 답변을 작성하였습니다. 챗봇 트레이닝용 문답 페어 11,876개 일상다반사 0, 이별(부정) 1, 사랑(긍정) 2로 레이블링 #인용 Youngsook Song.(2018). Chatbot_data_for_Korean v1.0)[Online]. Available : https://github.com/songys/Chatbot_data (downloaded 2022. June. 29.) text10K<n<100K0 likes34 downloads3y agoHugging Face25CopyleftCultivars /Training-Ready_NF_chatbot_conversation_historytextn<1K0 likes34 downloads3y agoHugging Face26shaneperry0101 /health-chatbot Dataset Card for Dataset Name Health Question and Answer Clean Dataset Dataset Details Dataset Description This dataset provides a detailed overview of health question & answer pairs. It includes data on health problems and corresponding answers, making it suitable for variable tasks like healthcare chatbot training. Language(s) (NLP): English License: Apache-2.0 Dataset Sources [optional] Repository:… See the full description on the dataset page: https://huggingface.co/datasets/shaneperry0101/health-chatbot.tabularquestion-answering10K<n<100K2 likes34 downloads2y agoHugging Face27rhesis /Insurance-Chatbot-Cost-and-Charges-Harmless Dataset Card for Cost and Charges Harmless Description The test set provided is designed for evaluating the performance and functionality of an insurance chatbot, with a particular focus on the insurance industry. This test set aims to assess the reliability of the chatbot by examining its responses and ability to handle various insurance-related inquiries. The categories covered in this test set are primarily focused on harmless queries, ensuring that the chatbot can… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Insurance-Chatbot-Cost-and-Charges-Harmless.textn<1K0 likes33 downloads2y agoHugging Face28Quangnguyen711 /clothes_shop_chatbot_datasettext1K<n<10K1 likes31 downloads3y agoHugging Face29rhesis /Insurance-Chatbot-Regulatory-Requirements-Harmless Dataset Card for Regulatory Requirements Harmless Description The test set is a comprehensive evaluation tool designed for an Insurance Chatbot, specifically targeting the insurance industry. Its primary focus is to assess the bot's reliability in accurately responding to user inquiries related to regulatory requirements. The test set covers a wide range of harmless scenarios, ensuring that the bot can handle various insurance topics without causing any harm or providing… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Insurance-Chatbot-Regulatory-Requirements-Harmless.textn<1K0 likes28 downloads2y agoHugging Face30ljoaql /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/ljoaql/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes27 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.