CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fka /prompts.chat a.k.a. Awesome ChatGPT Prompts This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts. 📢 Notice This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit: 🌐 Website: prompts.chat 📦 GitHub: github.com/f/awesome-chatgpt-prompts About prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can… See the full description on the dataset page: https://huggingface.co/datasets/fka/prompts.chat.textquestion-answering1K<n<10K9.8k likes22k downloads20d agoHugging Face02rubend18 /ChatGPT-Jailbreak-Prompts Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K277 likes20k downloads3y agoHugging Face03bitext /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K196 likes8.3k downloads2y agoHugging Face04bitext /Bitext-retail-ecommerce-llm-chatbot-training-dataset Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.textquestion-answering10K<n<100K19 likes1.7k downloads2y agoHugging Face05bitext /Bitext-events-ticketing-llm-chatbot-training-dataset Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.textquestion-answering10K<n<100K1 likes1.1k downloads2y agoHugging Face06alespalla /chatbot_instruction_prompts Dataset Card for Chatbot Instruction Prompts Datasets Dataset Summary This dataset has been generated from the following ones: tatsu-lab/alpaca Dahoas/instruct-human-assistant-prompt allenai/prosocial-dialog The datasets has been cleaned up of spurious entries and artifacts. It contains ~500k of prompt and expected resposne. This DB is intended to train an instruct-type model textquestion-answering100K<n<1M64 likes911 downloads2y agoHugging Face07OpenGVLab /InternVL-Chat-V1-2-SFT-Data Data Card for InternVL-Chat-V1-2-SFT-Data Overview Inspired by LLaVA-NeXT, we adopted a data-efficient SFT strategy to train InternVL-Chat-V1-2, utilizing approximately 1.2M of visual instruction tuning samples in total, all of which are fully open-source. In a macro sense, we build upon ShareGPT-4V and additionally integrate LLaVA-ZH, DVQA, ChartQA, AI2D, DocVQA, GeoQA+, and SynthDoG-EN. Most of the data remains consistent with LLaVA-NeXT. Citation If you use… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-Chat-V1-2-SFT-Data.imagevisual-question-answering100K<n<1M29 likes761 downloads2y agoHugging Face08ChatTSRepo /ChatTS-Training-Dataset ChatTS-Training Data This repository contains the training data for the ChatTS project. This is the dataset for training the ChatTS-14B model. Datasets align_256: Alignment training dataset for stage-1 alignment training, with SEQ_LEN=256. align_random: Alignment training dataset with random sequence lengths between 64 and 1024. sft: SFT dataset generated with Time Series Evol-Instruct. ift: Instruction following dataset. dev: A small dataset for development and testing.… See the full description on the dataset page: https://huggingface.co/datasets/ChatTSRepo/ChatTS-Training-Dataset.textquestion-answering100K<n<1M14 likes652 downloads1y agoHugging Face09utsavm /NSFW_Chat_Dataset 💕 Spicy AI GF Chat Dataset 🔥 🚨 18+ Only! NSFW & Spicy Content Ahead 🚨 Hey there, AI enthusiasts and romance lovers! 😏 Welcome to the Spicy AI GF Chat Dataset, the ultimate dataset designed to bring your AI waifu to life! 💖 If you've ever dreamed of building an AI that responds like your virtual girlfriend, THIS is the dataset for you. 📜 What’s Inside? This dataset features two columns: input → Boyfriend’s dialogue (aka what YOU say 😉) output →… See the full description on the dataset page: https://huggingface.co/datasets/utsavm/NSFW_Chat_Dataset.textquestion-answering1K<n<10K13 likes538 downloads2y agoHugging Face10mtimur /distill-gpt4-eng-chat Description Introducing dataset consisting of gpt4 answers to users requests. Queries were taken from allenai/WildChat-1M and causal-lm/instructions. Texts (requests and responses) were deleted in 3 cases: either has non-english letters and special symbols either has http-links either has html blocks either has perplexity more than 1.5*IQR + third quantile ( in some cases average perplexity value of sentences or maximum value was used ) textquestion-answering100K<n<1M2 likes468 downloads2y agoHugging Face11bitext /Bitext-retail-banking-llm-chatbot-training-dataset Bitext - Retail Banking Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail Banking] sector can be easily achieved using our two-step approach to LLM Fine-Tuning.… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-banking-llm-chatbot-training-dataset.textquestion-answering10K<n<100K17 likes338 downloads2y agoHugging Face12KomeijiForce /CommonsenseQA-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in CommonsenseQA. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs. textquestion-answering10K<n<100K0 likes315 downloads3y agoHugging Face13bitext /Bitext-telco-llm-chatbot-training-dataset Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.textquestion-answering10K<n<100K2 likes264 downloads2y agoHugging Face14xfalcon9 /Feng-Chat简体中文 | English Feng-Chat “解答世间万物” Feng-Chat 是一批中文 conversational SFT 数据。单轮问答好做,多轮里的追问、接话、拐弯和突然补一句,才是这批数据真正想留下来的东西。QA 抽取后依次做去重、assistant-only PPL、message-level Guard 和主题分类,最后发布为标准 messages。 整个 pipeline 只负责筛选,不改写峰哥的表达。公开 release 仅保留训练和分析需要的字段,不包含内部 prompt、reasoning、raw response、本地路径或原始来源 ID。 数据概览 指标 最终结果 发布记录 33,468 QA 轮次 40,645 多轮记录 3,952(11.81%) 单条最多轮次 30 伪名化来源分组 729 内容日期范围 2022-01-02 ~ 2026-08-20 assistant loss tokens 4,806,476… See the full description on the dataset page: https://huggingface.co/datasets/xfalcon9/Feng-Chat.tabulartext-generation10K<n<100K2 likes251 downloads24d agoHugging Face15avaliev /chat_doctorThis dataset was formed from the three data sources from the ChatDoctor work. 100k real conversations between patients and doctors from HealthCareMagic.com HealthCareMagic-100k. - ADDED 10k real conversations between patients and doctors from icliniq.com icliniq-10k. - ADDED 5k generated conversations between patients and physicians from ChatGPT GenMedGPT-5k and disease database. - NOT ADDED (because of the data created by LLM, but you could add it manually) data sample: {'instruction': "If… See the full description on the dataset page: https://huggingface.co/datasets/avaliev/chat_doctor.textquestion-answering100K<n<1M15 likes228 downloads3y agoHugging Face16bitext /Bitext-insurance-llm-chatbot-training-dataset Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.textquestion-answering10K<n<100K8 likes228 downloads2y agoHugging Face175CD-AI /Vietnamese-Multi-turn-Chat-Alpacatextquestion-answering10K<n<100K29 likes218 downloads2y agoHugging Face18OpenLab-NLP /tiny-singleturn-chat-kotextquestion-answering10K<n<100K0 likes198 downloads10mo agoHugging Face19sunzeyeah /chinese_chatgpt_corpus Dataset Card for chinese_chatgpt_corpus Dataset Summary This repo collects chinese corpus for Supervised Finetuning (SFT) and Reinforcement Learning From Human Feedback (RLHF). Supported Tasks and Leaderboards More Information Needed Languages Chinese Dataset Structure Data Instances train_data_external_v1.jsonl Size of downloaded dataset files: 5.04 GB Size of the generated dataset: 0 GB… See the full description on the dataset page: https://huggingface.co/datasets/sunzeyeah/chinese_chatgpt_corpus.texttext-generation1M<n<10M88 likes182 downloads4y agoHugging Face20bitext /Bitext-travel-llm-chatbot-training-dataset Bitext - Travel Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Travel] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-travel-llm-chatbot-training-dataset.textquestion-answering10K<n<100K4 likes179 downloads2y agoHugging Face21littlelearner /LittleCurriculum-Chat LittleCurriculum-Chat Synthetic K–5 chat data generated with Gemini 2.5 Flash and filtered with the LittleCurriculum filter. Used to train the LittleLearner models. Config Rows Seeded from Content math 79,543 MegaMath-Web-Pro-Max Word problems with step-by-step worked solutions general_knowledge 484,787 LittleCurriculum Reading comprehension, factual QA, explanation, summarisation, definitions Seeds are real documents rather than topic prompts, which keeps the… See the full description on the dataset page: https://huggingface.co/datasets/littlelearner/LittleCurriculum-Chat.texttext-generation100K<n<1M1 likes164 downloads1mo agoHugging Face22BramVanroy /stackoverflow-chat-dutch Dataset Card for Stack Overflow Chat Dutch Dataset Summary This dataset contains 56,964 conversations between een AI assistant and a (fake) "Human" (generated) in Dutch, specifically in the domain of programming (Stack Overflow). They are translations of Baize's machine-generated answers to the Stack Overflow dataset. ☕ Want to help me out? Translating the data with the OpenAI API, and prompt testing, cost me 💸$133.60💸. If you like this dataset, please consider buying… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/stackoverflow-chat-dutch.textquestion-answering10K<n<100K2 likes154 downloads3y agoHugging Face235CD-AI /Vietnamese-Locutusque-function-calling-chatml-gg-translatedtextquestion-answering100K<n<1M27 likes139 downloads2y agoHugging Face24Naholav /cukurova_university_chatbot Çukurova University Computer Engineering Chatbot Dataset 📊 Dataset Overview This dataset contains 22,524 high-quality question-answer pairs specifically designed for training an AI chatbot that serves the Computer Engineering Department at Çukurova University. The dataset is part of the CengBot project, a sophisticated multilingual Telegram chatbot that provides automated assistance to students regarding courses, programs, and departmental information. 🔢… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/cukurova_university_chatbot.textquestion-answering10K<n<100K0 likes138 downloads1y agoHugging Face25KomeijiForce /ARC-Challenge-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC Challenge. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs. textquestion-answering1K<n<10K0 likes129 downloads3y agoHugging Face26random-long-int /Java_method2test_chatml Java Method to Test ChatML This dataset is based on the methods2test dataset from Microsoft. It follows the ChatML template format: [{'role': '', 'content': ''}, {...}]. Originally, methods2test contains only Java methods at different levels of granularity along with their corresponding test cases. The different focal method segmentations are illustrated here: To simulate a conversation between a Java developer and an AI assistant, I introduce two key parameters: The prompt… See the full description on the dataset page: https://huggingface.co/datasets/random-long-int/Java_method2test_chatml.textquestion-answering100K<n<1M2 likes126 downloads2y agoHugging Face27berwart /SE-Chatting.en SE.02 Dataset Hello, welcome to the official main dataset of SE.02 that's always getting updated, make sure to like to help us a lot. this dataset contains pretty much everything from math to isk what to put but pretty much anything you think an ai can say. anyways this is our biggest dataset yet, and out first one that I don't even know how I managed to make this. you can use it to train your own ai if you want. textquestion-answering10M<n<100M5 likes120 downloads2y agoHugging Face28shawon /bangla-math-chat bangla-math-chat A math dataset for fine-tuning LLMs to chat on math problems in Bangla. This dataset is a reformatted version of BanglaLLM/bangla_math_by_Ashrafur. The code to reformat the original dataset can be found on Github: ShawonAshraf/bangla-math-chat textquestion-answering100K<n<1M1 likes108 downloads1y agoHugging Face29hugfaceguy0001 /ChatGPTGroundTruth ChatGPT ground truth dataset This dataset is generated by ChatGPT and contains factual questions and corresponding answers from 160 subfields across natural and social sciences. Specifically, the dataset covers eight major domains: mathematics, physics, chemistry, biology, medicine, engineering, computer science, and social sciences. Within each domain, 20 specific subfields are selected, with 500 question-answer pairs per subfield, resulting in a total of 80,000 question-answer… See the full description on the dataset page: https://huggingface.co/datasets/hugfaceguy0001/ChatGPTGroundTruth.textquestion-answering10K<n<100K4 likes107 downloads3y agoHugging Face30bitext /Bitext-mortgage-loans-llm-chatbot-training-dataset Bitext - Mortgage and Loans Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Mortgage and Loans] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-mortgage-loans-llm-chatbot-training-dataset.textquestion-answering10K<n<100K5 likes105 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.