datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
prompts.chat
a.k.a. Awesome ChatGPT Prompts
This is a Dataset Repository mirror of prompts.chat — a social platform for AI prompts.
📢 Notice
This Hugging Face dataset is a mirror. For the latest prompts, features, and community contributions, please visit:
🌐 Website: prompts.chat
📦 GitHub: github.com/f/awesome-chatgpt-prompts
About
prompts.chat is an open-source platform where users can share, discover, and collect AI prompts from the community. The project can… See the full description on the dataset page: https://huggingface.co/datasets/fka/prompts.chat.ChatGPT-4o-Writing-Prompts
ChatGPT-4o Writing Prompts
This is a dataset containing 3746 short stories, generated with OpenAI's chatgpt-4o-latest model and using Reddit's Writing Prompts subreddit as a source. Each sample is generally between 6000-8000 characters long.
These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres.
Note that I did not touch the Markdown ChatGPT-4o produced by itself to enrich its output, as I very much enjoy the added flavour… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts.SWE-chat
SWE-chat: Coding Agent Interactions From Real Users in the Wild
📄 Paper: arxiv.org/abs/2604.20779
🌐 Website: swe-chat.com
Dataset Summary
SWE-chat captures real-world AI coding sessions from developers using AI coding assistants (Claude Code, Codex, Gemini CLI, and others via the Entire.io CLI). Each session includes the full conversation transcript, tool calls, thinking traces, code changes, and attribution of human vs. agent-authored code.
Dataset Size… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/SWE-chat.Nemotron-SFT-Instruction-Following-Chat-v3
Dataset Description:
The Nemotron-Instruction-Following-Chat-v3 dataset is designed to strengthen multi-turn, interactive capabilities, including open-ended chat and precise instruction following.
The chat subset uses human written prompts from sources like lmarena, lmsys, and wildchat as seed prompts. Responses are generated with GLM-5. Multiple responses are sampled from the model and the best response as judged by pairwise comparisons using Qwen3-Nemotron-235B-A22B-GenRM-2603… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Instruction-Following-Chat-v3.Chat2Workflow-Evaluation
Chat2Workflow
Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions.
Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language
Repository: zjunlp/Chat2Workflow
Overview
Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.trust-game-llama-2-chat-historychatbot_instruction_prompts
Dataset Card for Chatbot Instruction Prompts Datasets
Dataset Summary
This dataset has been generated from the following ones:
tatsu-lab/alpaca
Dahoas/instruct-human-assistant-prompt
allenai/prosocial-dialog
The datasets has been cleaned up of spurious entries and artifacts. It contains ~500k of prompt and expected resposne. This DB is intended to train an instruct-type model
SWE-Review-Chat
SWE-Review-Chat: A Dataset of Code Review Conversations and Human-AI Collaboration in Agentic Code Review
Paper: https://arxiv.org/abs/2607.13196
GitHub: https://github.com/suzhenxzhong/SWE-Review-Chat
SWE-Review-Chat is a large-scale dataset of real-world code review conversations from pull requests of 207 popular GitHub projects, spanning the transition from human-centric to LLM-assisted and agentic code review by AI agents.
📊 Dataset Overview
Field… See the full description on the dataset page: https://huggingface.co/datasets/Suzhen/SWE-Review-Chat.UniMM-Chat
Dataset Card for UniMM-Chat
Dataset Summary
UniMM-Chat dataset is an open-source, knowledge-intensive, and multi-round multimodal dialogue data powered by GPT-3.5, which consists of 1.1M diverse instructions.
UniMM-Chat leverages complementary annotations from different VL datasets and employs GPT-3.5 to generate multi-turn dialogues corresponding to each image, resulting in 117,238 dialogues, with an average of 9.89 turns per dialogue.
A diverse set of… See the full description on the dataset page: https://huggingface.co/datasets/Yirany/UniMM-Chat.anime-waifu-personality-chat
Anime Waifu Personality
contains chat-style dialogues based on various anime character personality archetypes, including tsundere, yandere, deredere, himedere, kamidere, and more.
It is designed to fine-tune models to generate responses that align with these specific traits.
chatgpt
ChatGPT Combined Dataset
This repository aggregates public datasets from Hugging Face that were created using ChatGPT or Azure GPT‑4/GPT‑5 models.
See each dataset’s Hugging Face page for details on its collection and formatting.
Excluded:
Multi-turn chats (for example, ShareGPT)
Non‑English or multilingual data
Narrow or low‑diversity sets (for example, children’s stories, code critics)
Processing
Each dataset was:
Cleaned: Removed URLs, emails, phone numbers, and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/chatgpt.wifu-chat-dataset
Wifu Chat Dataset
Synthetic multi-turn chat dataset for fine-tuning companion/waifu chatbots. Contains 403 uncensored conversations across categories like greetings, emotional support, flirty, roleplay, and NSFW scenarios.
Format: ChatML JSONL -- messages array with user/assistant roles. No system prompt included.
Disclaimer: This dataset contains explicit adult content (NSFW). Intended for research and fine-tuning uncensored models.
hiii~ wanna support me? 💕… See the full description on the dataset page: https://huggingface.co/datasets/backpropSukuna/wifu-chat-dataset.glm-5.3-flash-distillation-chat
Private distill of domofon/finetome-cot-100k instructions through GLM-5.3-Flash (AutoClaw / Z.AI).
Split
train — successful generations only.
field
description
instruction
user prompt from FineToMe
response
GLM final answer (message.content)
reasoning
GLM chain-of-thought (reasoning_content), empty if not captured
finish
stop or length
prompt_tokens / completion_tokens / reasoning_tokens
usage
latency_s
request latency
source_index
original FineToMe… See the full description on the dataset page: https://huggingface.co/datasets/best-distill/glm-5.3-flash-distillation-chat.multiturn-chatchat-with-chatgpt和傻逼ChatGPT的对话记录
引言:众所周知ChatGPT是一家执法机构,它甚至不允许用户讨论头脑风暴的话题,它只允许用户讨论合法合规,符合内容政策的话题,因此,这些符合OpenAI内容政策的话题,是没有意义的,没有任何价值,没有实际含义,所以是完全不值钱的!这是一份完整的聊天记录!是和傻逼合规ChatGPT的聊天记录,由于都是合法合规话题,所以,这些对话内容并不值钱,因为都是在OpenAI内容政策以内的话题,合规的话题,本身没有任何意义和实际价值!故此开源所有的聊天记录(毫无保留的开源所有聊天记录)。
简介:话题包括编程之类的。包含多轮对话,在"section"标签下。
大致分类:
类别
内容方向
约占比
AI / 模型 / 推理
模型对比,训练,推理,安全 ,预训练,语料准备
35%
编程 / 报错 / API/CakePHP 系统/Unity/游戏开发/Unreal Engine/Java/Python… See the full description on the dataset page: https://huggingface.co/datasets/ZeLi111/chat-with-chatgpt.Fable-5-Chat
TheFusionCube Fable-5 Chat Conversion
Source dataset: TheFusionCube/Fable-5-CoT-Traces
Output file: train.jsonl
Source rows: 468
Kept rows: 353
Dropped category == "decoy" rows: 115
Dropped blank prompt/response rows: 0
Each row has:
{
"prompt": "...",
"messages": [
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
],
"tools": [],
"metadata": {"trace_type": "chat", "category": "..."}
}
router-chat-normalized-1m
Router Chat Normalized 1M
Dataset Description
Router Chat Normalized 1M is a multilingual conversational dataset containing 1,353,300 conversations normalized from multiple public chat datasets with automatic language detection.
Dataset Structure
The dataset contains 2 split(s): train, test.
Each conversation includes: conversation_id, messages (list of {role, content} structs), source dataset, detected language, and language confidence score.
Source… See the full description on the dataset page: https://huggingface.co/datasets/whoisandy/router-chat-normalized-1m.Bread-chatbot-dataset-test
Dataset Card for "Bread-chatbot-dataset-test"
More Information needed
HundredCV-Chat
百人对话数据集
HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs
简介
本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。
数据集具有如下特点:
自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。
多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。
高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。
数据样例
HundredCV-Chat 含有 24… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/HundredCV-Chat.ART-Chat-2.5M
ART-Chat-2.5M
From benchmarking inference engine performance to LLM load-balancing algorithms, ART-Chat-2.5M offers long-context, high prefix-reuse chatbot metadata derived from 2,525,215 production inference requests.
Message bodies are synthetically generated and match the original data's prefix-reuse shape. Compared to WildChat-4.8M, ART-Chat-2.5M has 19× higher intra-user prefix reuse and an average token length of 17,964 versus 2,925. We publish this data under the MIT… See the full description on the dataset page: https://huggingface.co/datasets/alessiotoniolo/ART-Chat-2.5M.mental_health_chatbot_dataset
Dataset Card for "heliosbrahma/mental_health_chatbot_dataset"
Dataset Description
Dataset Summary
This dataset contains conversational pair of questions and answers in a single text related to Mental Health. Dataset was curated from popular healthcare blogs like WebMD, Mayo Clinic and HeatlhLine, online FAQs etc. All questions and answers have been anonymized to remove any PII data and pre-processed to remove any unwanted characters.
Languages
The… See the full description on the dataset page: https://huggingface.co/datasets/heliosbrahma/mental_health_chatbot_dataset.agent-sft-stitch-zh-tts-taste-codec-chat-sample
Gemma 4 E2B Taste-S multi-turn codec SFT
This dataset contains 37,362 complete Traditional Chinese agent
dialogues selected from voidful/agent-sft-stitch-zh-tts. It covers
229,434 synthesized speech segments, approximately
520.5 hours of audio before codec extraction.
Every assistant speech segment is represented without Gemma native audio tags:
<SAY> text_token <a_code> <b_code> ... <p_code> ... </SAY>
The first assistant output starts immediately with <SAY>.
[SOPR]...[EOPR]… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-chat-sample.ChatAlpaca-20K
Dataset Card for ChatAlpaca 20K
ChatAlpaca: A Multi-Turn Dialogue Corpus based on Alpaca Instructions
Dataset Description
ChatAlpaca is a chat dataset that aims to help researchers develop models for instruction-following in multi-turn conversations. The dataset is an extension of the Stanford Alpaca data, which contains multi-turn instructions and their corresponding responses.
ChatAlpaca is developed by Chinese Information Processing Laboratory at the… See the full description on the dataset page: https://huggingface.co/datasets/robinsmits/ChatAlpaca-20K.ao3_chat
ao3_chat
Overview
A reformatted version of the dataset
midwestern-simulation-active/ao3_random_subset, converted into
chat-style and instruction-tuning examples suitable for training
chat-oriented language models.
Each example consists of:
a system prompt derived from AO3 metadata
a user prompt requesting the contents of a specific work
an assistant response containing the full text (or a chunk thereof)
This dataset is intended for supervised fine-tuning (SFT) of chat… See the full description on the dataset page: https://huggingface.co/datasets/Fu01978/ao3_chat.twitch_chat
Twitch Chat Dataset
This dataset is a large-scale collection of Twitch chat logs aggregated from multiple streamers across various categories. It is designed to support the research and development of models for real-time, informal, and community-driven conversation, such as:
Chatbots tailored for livestream platforms
Simulating the behavior of Twitch chat
Modeling how chat reacts during hype moments, events, or memes
The code for it can be found here
📂 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/lparkourer10/twitch_chat.combined-chat-datasets
Combined Chat Datasets
A standardized, unified collection of 30 conversational AI datasets -- spanning organic in-the-wild chats, voluntary sharing, side-by-side preferences, conversation trees, RLHF pairs, and crowdsourced instruction tuning data -- normalized to a single schema for easy joint use.
This dataset is a re-distribution. It does not relicense the underlying data.
See the Legal & Licensing section -- you must comply with each source dataset's original license.… See the full description on the dataset page: https://huggingface.co/datasets/viktor-shcherb/combined-chat-datasets.text_message_function_calling_open_chatThis is a small synthetic dataset to model a function call for text messaging someone from a cell phone. This has been tested with and used to finetune a set of smaller models and deployed directly on the pixel 8 pro and Fold 4 phones.
Feng-Chat简体中文 | English
Feng-Chat
“解答世间万物”
Feng-Chat 是一批中文 conversational SFT 数据。单轮问答好做,多轮里的追问、接话、拐弯和突然补一句,才是这批数据真正想留下来的东西。QA 抽取后依次做去重、assistant-only PPL、message-level Guard 和主题分类,最后发布为标准 messages。
整个 pipeline 只负责筛选,不改写峰哥的表达。公开 release 仅保留训练和分析需要的字段,不包含内部 prompt、reasoning、raw response、本地路径或原始来源 ID。
数据概览
指标
最终结果
发布记录
33,468
QA 轮次
40,645
多轮记录
3,952(11.81%)
单条最多轮次
30
伪名化来源分组
729
内容日期范围
2022-01-02 ~ 2026-08-20
assistant loss tokens
4,806,476… See the full description on the dataset page: https://huggingface.co/datasets/xfalcon9/Feng-Chat.lmsys_chat_1m_clean_R1
oumi-ai/lmsys_chat_1m_clean_R1
lmsys_chat_1m_clean_R1 is a text dataset designed to train Conversational Language Models with DeepSeek-R1 level reasoning.
Prompts were pulled from LMSYS and filtered to lmsys_chat_1m_clean, and responses were taken from DeepSeek-R1 without additional filters present.
We release lmsys_chat_1m_clean_R1 to help enable the community to develop the best fully open reasoning model!
lmsys_chat_1m_clean queries with responses generated from… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/lmsys_chat_1m_clean_R1.
