datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prompts.chat
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can… See the full description on the dataset page: https://huggingface.co/datasets/fka/prompts.chat.ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
Bitext-customer-support-llm-chatbot-training-dataset
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.Bitext-retail-ecommerce-llm-chatbot-training-dataset
Bitext - Retail (eCommerce) Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail (eCommerce)] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-ecommerce-llm-chatbot-training-dataset.Bitext-events-ticketing-llm-chatbot-training-dataset
Bitext - Events and Ticketing Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [events and ticketing] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-events-ticketing-llm-chatbot-training-dataset.chatbot_instruction_prompts
Dataset Card for Chatbot Instruction Prompts Datasets
Dataset Summary
This dataset has been generated from the following ones:
tatsu-lab/alpaca
Dahoas/instruct-human-assistant-prompt
allenai/prosocial-dialog
The datasets has been cleaned up of spurious entries and artifacts. It contains ~500k of prompt and expected resposne. This DB is intended to train an instruct-type model
InternVL-Chat-V1-2-SFT-Data
Data Card for InternVL-Chat-V1-2-SFT-Data
Overview
Inspired by LLaVA-NeXT, we adopted a data-efficient SFT strategy to train InternVL-Chat-V1-2, utilizing approximately 1.2M of visual instruction tuning samples in total, all of which are fully open-source. In a macro sense, we build upon ShareGPT-4V and additionally integrate LLaVA-ZH, DVQA, ChartQA, AI2D, DocVQA, GeoQA+, and SynthDoG-EN. Most of the data remains consistent with LLaVA-NeXT.
Citation
If you use… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-Chat-V1-2-SFT-Data.ChatTS-Training-Dataset
ChatTS-Training Data
This repository contains the training data for the ChatTS project. This is the dataset for training the ChatTS-14B model.
Datasets
align_256: Alignment training dataset for stage-1 alignment training, with SEQ_LEN=256.
align_random: Alignment training dataset with random sequence lengths between 64 and 1024.
sft: SFT dataset generated with Time Series Evol-Instruct.
ift: Instruction following dataset.
dev: A small dataset for development and testing.… See the full description on the dataset page: https://huggingface.co/datasets/ChatTSRepo/ChatTS-Training-Dataset.NSFW_Chat_Dataset
💕 Spicy AI GF Chat Dataset 🔥
🚨 18+ Only! NSFW & Spicy Content Ahead 🚨
Hey there, AI enthusiasts and romance lovers! 😏 Welcome to the Spicy AI GF Chat Dataset, the ultimate dataset designed to bring your AI waifu to life! 💖 If you've ever dreamed of building an AI that responds like your virtual girlfriend, THIS is the dataset for you.
📜 What’s Inside?
This dataset features two columns:
input → Boyfriend’s dialogue (aka what YOU say 😉)
output →… See the full description on the dataset page: https://huggingface.co/datasets/utsavm/NSFW_Chat_Dataset.distill-gpt4-eng-chat
Description
Introducing dataset consisting of gpt4 answers to users requests. Queries were taken from allenai/WildChat-1M and causal-lm/instructions. Texts (requests and responses) were deleted in 3 cases:
either has non-english letters and special symbols
either has http-links
either has html blocks
either has perplexity more than 1.5*IQR + third quantile ( in some cases average perplexity value of sentences or maximum value was used )
Bitext-retail-banking-llm-chatbot-training-dataset
Bitext - Retail Banking Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Retail Banking] sector can be easily achieved using our two-step approach to LLM Fine-Tuning.… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-retail-banking-llm-chatbot-training-dataset.CommonsenseQA-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in CommonsenseQA. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
Bitext-telco-llm-chatbot-training-dataset
Bitext - Telco Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [telco] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-telco-llm-chatbot-training-dataset.Feng-Chat简体中文 | English
Feng-Chat
“解答世间万物”
Feng-Chat 是一批中文 conversational SFT 数据。单轮问答好做,多轮里的追问、接话、拐弯和突然补一句,才是这批数据真正想留下来的东西。QA 抽取后依次做去重、assistant-only PPL、message-level Guard 和主题分类,最后发布为标准 messages。
整个 pipeline 只负责筛选,不改写峰哥的表达。公开 release 仅保留训练和分析需要的字段,不包含内部 prompt、reasoning、raw response、本地路径或原始来源 ID。
数据概览
指标
最终结果
发布记录
33,468
QA 轮次
40,645
多轮记录
3,952(11.81%)
单条最多轮次
30
伪名化来源分组
729
内容日期范围
2022-01-02 ~ 2026-08-20
assistant loss tokens
4,806,476… See the full description on the dataset page: https://huggingface.co/datasets/xfalcon9/Feng-Chat.chat_doctorThis dataset was formed from the three data sources from the ChatDoctor work.
100k real conversations between patients and doctors from HealthCareMagic.com HealthCareMagic-100k. - ADDED
10k real conversations between patients and doctors from icliniq.com icliniq-10k. - ADDED
5k generated conversations between patients and physicians from ChatGPT GenMedGPT-5k and disease database. - NOT ADDED (because of the data created by LLM, but you could add it manually)
data sample:
{'instruction': "If… See the full description on the dataset page: https://huggingface.co/datasets/avaliev/chat_doctor.Bitext-insurance-llm-chatbot-training-dataset
Bitext - Insurance Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [insurance] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-insurance-llm-chatbot-training-dataset.Vietnamese-Multi-turn-Chat-Alpacatiny-singleturn-chat-kochinese_chatgpt_corpus
Dataset Card for chinese_chatgpt_corpus
Dataset Summary
This repo collects chinese corpus for Supervised Finetuning (SFT) and Reinforcement Learning From Human Feedback (RLHF).
Supported Tasks and Leaderboards
More Information Needed
Languages
Chinese
Dataset Structure
Data Instances
train_data_external_v1.jsonl
Size of downloaded dataset files: 5.04 GB
Size of the generated dataset: 0 GB… See the full description on the dataset page: https://huggingface.co/datasets/sunzeyeah/chinese_chatgpt_corpus.Bitext-travel-llm-chatbot-training-dataset
Bitext - Travel Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Travel] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-travel-llm-chatbot-training-dataset.LittleCurriculum-Chat
LittleCurriculum-Chat
Synthetic K–5 chat data generated with Gemini 2.5 Flash and filtered with the
LittleCurriculum filter.
Used to train the LittleLearner models.
Config
Rows
Seeded from
Content
math
79,543
MegaMath-Web-Pro-Max
Word problems with step-by-step worked solutions
general_knowledge
484,787
LittleCurriculum
Reading comprehension, factual QA, explanation, summarisation, definitions
Seeds are real documents rather than topic prompts, which keeps the… See the full description on the dataset page: https://huggingface.co/datasets/littlelearner/LittleCurriculum-Chat.stackoverflow-chat-dutch
Dataset Card for Stack Overflow Chat Dutch
Dataset Summary
This dataset contains 56,964 conversations between een AI assistant and a (fake) "Human" (generated) in Dutch, specifically in the domain of programming (Stack Overflow). They are translations of Baize's machine-generated answers to the Stack Overflow dataset.
☕ Want to help me out? Translating the data with the OpenAI API, and prompt testing, cost me 💸$133.60💸. If you like this dataset, please consider buying… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/stackoverflow-chat-dutch.Vietnamese-Locutusque-function-calling-chatml-gg-translatedcukurova_university_chatbot
Çukurova University Computer Engineering Chatbot Dataset
📊 Dataset Overview
This dataset contains 22,524 high-quality question-answer pairs specifically designed for training an AI chatbot that serves the Computer Engineering Department at Çukurova University. The dataset is part of the CengBot project, a sophisticated multilingual Telegram chatbot that provides automated assistance to students regarding courses, programs, and departmental information.
🔢… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/cukurova_university_chatbot.ARC-Challenge-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC Challenge. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
Java_method2test_chatml
Java Method to Test ChatML
This dataset is based on the methods2test dataset from Microsoft. It follows the ChatML template format: [{'role': '', 'content': ''}, {...}].
Originally, methods2test contains only Java methods at different levels of granularity along with their corresponding test cases. The different focal method segmentations are illustrated here:
To simulate a conversation between a Java developer and an AI assistant, I introduce two key parameters:
The prompt… See the full description on the dataset page: https://huggingface.co/datasets/random-long-int/Java_method2test_chatml.SE-Chatting.en
SE.02
Dataset
Hello, welcome to the official main dataset of SE.02 that's always getting updated, make sure to like to help us a lot.
this dataset contains pretty much everything from math to isk what to put but pretty much anything you think an ai can say.
anyways this is our biggest dataset yet, and out first one that I don't even know how I managed to make this.
you can use it to train your own ai if you want.
bangla-math-chat
bangla-math-chat
A math dataset for fine-tuning LLMs to chat on math problems in Bangla. This dataset is a reformatted version of
BanglaLLM/bangla_math_by_Ashrafur.
The code to reformat the original dataset can be found on Github: ShawonAshraf/bangla-math-chat
ChatGPTGroundTruth
ChatGPT ground truth dataset
This dataset is generated by ChatGPT and contains factual questions and corresponding answers from 160 subfields across natural and social sciences.
Specifically, the dataset covers eight major domains: mathematics, physics, chemistry, biology, medicine, engineering, computer science, and social sciences. Within each domain, 20 specific subfields are selected, with 500 question-answer pairs per subfield, resulting in a total of 80,000 question-answer… See the full description on the dataset page: https://huggingface.co/datasets/hugfaceguy0001/ChatGPTGroundTruth.Bitext-mortgage-loans-llm-chatbot-training-dataset
Bitext - Mortgage and Loans Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Mortgage and Loans] sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-mortgage-loans-llm-chatbot-training-dataset.
