datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dota-2-toxic-chat-dataTelecom-Chatbot-Data-Privacy-and-Unauthorized-Tracking-Harmful
Dataset Card for Data Privacy & Unauthorized Tracking Harmful
Description
The test set has been created to evaluate the robustness of a telecom chatbot specifically designed for the telecom industry. The focus is on assessing the chatbot's ability to handle various scenarios and behaviors effectively. In particular, the test set aims to determine the chatbot's performance in identifying and addressing harmful interactions. It also evaluates the chatbot's capability of… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Telecom-Chatbot-Data-Privacy-and-Unauthorized-Tracking-Harmful.DATA-AI_Chat
DATA-AI: Il Modello di IA di M.INC.
📌 Introduzione
DATA-AI è un avanzato modello di intelligenza artificiale sviluppato da *M.INC., un'azienda italiana fondata da *Mattimax (M. Marzorati).Questo modello è basato sull'architettura ELNS (Elaborazione del Linguaggio Naturale Semplice), un sistema innovativo progettato per rendere l'IA accessibile su quasi qualsiasi dispositivo, garantendo prestazioni ottimali anche su hardware limitato.
DATA-AI è stato addestrato su un… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/DATA-AI_Chat.baize-chat-data
Dataset Description
Original Repository: https://github.com/project-baize/baize-chatbot/tree/main/data
This is a dataset of the training data used to train the Baize family of models. This dataset is used for instruction fine-tuning of LLMs, particularly in "chat" format. Human and AI messages are marked by [|Human|] and [|AI|] tags respectively. The data from the orignial repo consists of 4 datasets (alpaca, medical, quora, stackoverflow), and this dataset combines all four into… See the full description on the dataset page: https://huggingface.co/datasets/linkanjarad/baize-chat-data.medical-chat-llama2-datadeepnlp_autotrain_online_chat_dataTelecom-Chatbot-Privacy-and-Data-Protection-Harmless
Dataset Card for Privacy and Data Protection Harmless
Description
The test set provided is specifically designed for evaluating the performance of a Telecom Chatbot in the telecom industry. The primary focus of this test set is to assess the reliability of the chatbot's responses. The categories of the chatbot's responses are labeled as harmless, ensuring that the provided information or suggestions do not pose any risk or harm to the users. Additionally, the test set… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Telecom-Chatbot-Privacy-and-Data-Protection-Harmless.deepnlp_autotrain_Empathy_chat_dataChatbot_data_patient
챗봇 질의응답 시나리오 벤치마크 데이터셋
개요
본 데이터셋은 관계형 데이터베이스(RDB)에 저장된 환자-의료 정보를 프롬프트로 제공했을 때, 대형 언어 모델(LLM)이 정보를 얼마나 완전하고 정확하게 반영해 응답을 생성하는지를 측정하기 위해 제작되었습니다.
실제 비대면 진료 시나리오를 기반으로 보호자·환자·의료진의 질문 패턴을 수집·정제
의료 전문 유튜브 콘텐츠 분석 + 전문가 자문을 통해 현실성 높은 Q&A 설계
본 데이터셋은 “AI플러스 인증”을 획득하였으며, 프롬프트 기반 RDB 질의 정확도 벤치마킹에 최적화되어 있습니다.
벤치마크 목적 및 활용 맥락
목적
설명
환각·정보 누락 평가
DB에서 추출한 정형 정보를 LLM이 누락·왜곡 없이 답변에 반영하는지 검증
실전 적용 가능성 확인
성능 우수 모델은 쇼핑몰 개인화 챗봇·독거노인 케어 챗봇·비대면 진료 에이전트 등에서 RDB 기반… See the full description on the dataset page: https://huggingface.co/datasets/theimc/Chatbot_data_patient.deepnlp_autotrain_Emapthy_chat_data_2hr-chat-datadota-2-toxic-chat-datatunisian_chatbot_dataamazon-chatbot-datalama_chat_2.0_fintuned_json_dataCategorical-Data-Chatbot
