datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
ChatGPT-Jailbreak-PromptsChatGPT-4o-Writing-Prompts
ChatGPT-4o Writing Prompts
This is a dataset containing 3746 short stories, generated with OpenAI's chatgpt-4o-latest model and using Reddit's Writing Prompts subreddit as a source. Each sample is generally between 6000-8000 characters long.
These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres.
Note that I did not touch the Markdown ChatGPT-4o produced by itself to enrich its output, as I very much enjoy the added flavour… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts.working-memory-capacity-of-ChatGPT
Using N-back Tasks to Assess Working Memory Capacity of Large Language Models (LLMs)
This is a code and dataset repository for the paper "Working Memory Capacity of ChatGPT: An Empirical Study", which has been accepted by AAAI 2024 Conference on Artificial Intelligence.
Here we created a dataset to test the working memory capacity of language models. We choose the N-back task because it is widely used in cognitive science as a measure of working memory capacity. To create the… See the full description on the dataset page: https://huggingface.co/datasets/dongyu0205/working-memory-capacity-of-ChatGPT.chatgpt
ChatGPT Combined Dataset
This repository aggregates public datasets from Hugging Face that were created using ChatGPT or Azure GPT‑4/GPT‑5 models.
See each dataset’s Hugging Face page for details on its collection and formatting.
Excluded:
Multi-turn chats (for example, ShareGPT)
Non‑English or multilingual data
Narrow or low‑diversity sets (for example, children’s stories, code critics)
Processing
Each dataset was:
Cleaned: Removed URLs, emails, phone numbers, and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/chatgpt.chat-with-chatgpt和傻逼ChatGPT的对话记录
引言:众所周知ChatGPT是一家执法机构,它甚至不允许用户讨论头脑风暴的话题,它只允许用户讨论合法合规,符合内容政策的话题,因此,这些符合OpenAI内容政策的话题,是没有意义的,没有任何价值,没有实际含义,所以是完全不值钱的!这是一份完整的聊天记录!是和傻逼合规ChatGPT的聊天记录,由于都是合法合规话题,所以,这些对话内容并不值钱,因为都是在OpenAI内容政策以内的话题,合规的话题,本身没有任何意义和实际价值!故此开源所有的聊天记录(毫无保留的开源所有聊天记录)。
简介:话题包括编程之类的。包含多轮对话,在"section"标签下。
大致分类:
类别
内容方向
约占比
AI / 模型 / 推理
模型对比,训练,推理,安全 ,预训练,语料准备
35%
编程 / 报错 / API/CakePHP 系统/Unity/游戏开发/Unreal Engine/Java/Python… See the full description on the dataset page: https://huggingface.co/datasets/ZeLi111/chat-with-chatgpt.CommonsenseQA-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in CommonsenseQA. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
Gryphe_ChatGPT_4o_Writing_Prompts_ChineseChatGPT-prompts
ChatGPT-Prompts Dataset
Description
This dataset aims to provide an evaluation data for the Language Models to come. It has been generated using LearnGPT website.
chatgpt-paraphrasesThis is a dataset of paraphrases created by ChatGPT.
Model based on this dataset is avaible: model
We used this prompt to generate paraphrases
Generate 5 similar paraphrases for this question, show it like a numbered list without commentaries: {text}
This dataset is based on the Quora paraphrase question, texts from the SQUAD 2.0 and the CNN news dataset.
We generated 5 paraphrases for each sample, totally this dataset has about 420k data rows. You can make 30 rows from a row from… See the full description on the dataset page: https://huggingface.co/datasets/humarin/chatgpt-paraphrases.chinese_chatgpt_corpus
Dataset Card for chinese_chatgpt_corpus
Dataset Summary
This repo collects chinese corpus for Supervised Finetuning (SFT) and Reinforcement Learning From Human Feedback (RLHF).
Supported Tasks and Leaderboards
More Information Needed
Languages
Chinese
Dataset Structure
Data Instances
train_data_external_v1.jsonl
Size of downloaded dataset files: 5.04 GB
Size of the generated dataset: 0 GB… See the full description on the dataset page: https://huggingface.co/datasets/sunzeyeah/chinese_chatgpt_corpus.ChatGPT-RealUser-2.2M-preview
ChatGPT-RealUser-2.2M: A Large-Scale Dataset of Real-User, Real-World ChatGPT Conversations
ChatGPT-RealUser-2.2M is a large-scale dataset of real-user, Real-World ChatGPT conversations developed by Gata.
From 2024–2025, participants using Gata’s GPT-to-Earn product opted in to share their chats and earned points based on conversation quality.
The dataset covers GPT-3.5, GPT-4, and o1 models, and contains 2,244,389 conversations from 15,316 unique users.
Because many chats are… See the full description on the dataset page: https://huggingface.co/datasets/Gata-community/ChatGPT-RealUser-2.2M-preview.FIXED-Cleaned-Claude-Sonnet-5-Grok-4.5-ChatGPT-5.6-Luna-Qwen-3.8-MAX
Ultra-Clean HF Dataset
492 pairs. Zero residual JSON garbage. High-tier technical SFT data.
ARC-Challenge-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC Challenge. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
chatgpt-openqa-zsm-qaretrieval
chatgpt-openqa-zsm-qaretrieval
Deduplicated copy of kornwtp/chatgpt-openqa-zsm-qaretrieval,
part of the SEA-BED data-quality work.
Source dataset: kornwtp/chatgpt-openqa-zsm-qaretrieval
Deduplicated on: 2026-09-04
Task type: qa_retrieval
Splits: train
What changed
Kept in this dataset's ORIGINAL schema -- same columns, same nesting, same extra fields (ids, titles, answers) -- so it is a drop-in replacement for the source repo. Documents differing only in… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/chatgpt-openqa-zsm-qaretrieval.ChatGPT-Research-Abstracts
ChatGPT-Research-Abstracts
This is a dataset created in relation to a bachelor thesis written by Nicolai Thorer Sivesind and Andreas Bentzen Winje. It contains human-produced and machine-generated text samples of scientific research abstracts.
A reformatted version for text-classification is available in the dataset collection Human-vs-Machine. In this collection, all samples are split into separate data points for real and generated, and labeled either 0 (human-produced) or 1… See the full description on the dataset page: https://huggingface.co/datasets/NicolaiSivesind/ChatGPT-Research-Abstracts.llama-3.1-awesome-chatgpt-prompts
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/llama-3.1-awesome-chatgpt-prompts.ChatGPTGroundTruth
ChatGPT ground truth dataset
This dataset is generated by ChatGPT and contains factual questions and corresponding answers from 160 subfields across natural and social sciences.
Specifically, the dataset covers eight major domains: mathematics, physics, chemistry, biology, medicine, engineering, computer science, and social sciences. Within each domain, 20 specific subfields are selected, with 500 question-answer pairs per subfield, resulting in a total of 80,000 question-answer… See the full description on the dataset page: https://huggingface.co/datasets/hugfaceguy0001/ChatGPTGroundTruth.ARC-Easy-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC-Easy. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs.
reward-bench-chatgpt-4o-latest-yes-nochatgpt-prompts-Swedishchatgpt-gpt-chat-jsonlData is collected during genuine chat with chatgpt.
and openai_gsm8k_0_7473.jsonl by openai, real: openai/gsm8k. It has been shortened.
ty
video_chatgpt_activitynet_videoschatgpt-news-articles
Dataset Card for "chatgpt-news-articles"
Dataset Summary
The ChatGPT CNN / DailyMail Dataset is a small sample of the original CNN / DailyMaily English-language dataset containing 25k unique news articles. For each corresponding article written by journalists at CNN and the Daily Mail, there is an article written by ChatGPT using the highlights provided by human annotators. The current version supports can be used to study the language comparison between human and ChatGPT… See the full description on the dataset page: https://huggingface.co/datasets/isarth/chatgpt-news-articles.awesome-chatgpt-prompts-clean
🧠 Awesome ChatGPT Prompts — Clean
The classic 2,112-prompt role-prompting library (fka/prompts.chat, CC0) — deduplicated, quality-filtered, auto-categorized, shipped as typed parquet — plus 6 hand-verified community prompts mined from Claude practitioner chat.
Priorities: Quality > Cleanliness > Signal
Clean derivative of fka/prompts.chat (2,124 rows). License unchanged: CC0-1.0 ✅ no restrictions.
🧹 Quality Pipeline
Step
Removed
Reason
Raw… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/awesome-chatgpt-prompts-clean.chatgpt-prompts-Frenchchatgpt-prompts-EstonianDALL-E-Prompts-OpenAI-ChatGPT
Dataset Card for Dataset Name
Dataset Summary
This dataset has been generated using Prompt Generator for OpenAI's DALL-E.
Languages
English
Dataset Structure
1.000.000 Prompts
ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
