datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChatGPT-4o-Writing-Prompts
ChatGPT-4o Writing Prompts
This is a dataset containing 3746 short stories, generated with OpenAI's chatgpt-4o-latest model and using Reddit's Writing Prompts subreddit as a source. Each sample is generally between 6000-8000 characters long.
These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres.
Note that I did not touch the Markdown ChatGPT-4o produced by itself to enrich its output, as I very much enjoy the added flavour… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts.chatgpt
ChatGPT Combined Dataset
This repository aggregates public datasets from Hugging Face that were created using ChatGPT or Azure GPT‑4/GPT‑5 models.
See each dataset’s Hugging Face page for details on its collection and formatting.
Excluded:
Multi-turn chats (for example, ShareGPT)
Non‑English or multilingual data
Narrow or low‑diversity sets (for example, children’s stories, code critics)
Processing
Each dataset was:
Cleaned: Removed URLs, emails, phone numbers, and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/chatgpt.chat-with-chatgpt和傻逼ChatGPT的对话记录
引言:众所周知ChatGPT是一家执法机构,它甚至不允许用户讨论头脑风暴的话题,它只允许用户讨论合法合规,符合内容政策的话题,因此,这些符合OpenAI内容政策的话题,是没有意义的,没有任何价值,没有实际含义,所以是完全不值钱的!这是一份完整的聊天记录!是和傻逼合规ChatGPT的聊天记录,由于都是合法合规话题,所以,这些对话内容并不值钱,因为都是在OpenAI内容政策以内的话题,合规的话题,本身没有任何意义和实际价值!故此开源所有的聊天记录(毫无保留的开源所有聊天记录)。
简介:话题包括编程之类的。包含多轮对话,在"section"标签下。
大致分类:
类别
内容方向
约占比
AI / 模型 / 推理
模型对比,训练,推理,安全 ,预训练,语料准备
35%
编程 / 报错 / API/CakePHP 系统/Unity/游戏开发/Unreal Engine/Java/Python… See the full description on the dataset page: https://huggingface.co/datasets/ZeLi111/chat-with-chatgpt.chinese_chatgpt_corpus
Dataset Card for chinese_chatgpt_corpus
Dataset Summary
This repo collects chinese corpus for Supervised Finetuning (SFT) and Reinforcement Learning From Human Feedback (RLHF).
Supported Tasks and Leaderboards
More Information Needed
Languages
Chinese
Dataset Structure
Data Instances
train_data_external_v1.jsonl
Size of downloaded dataset files: 5.04 GB
Size of the generated dataset: 0 GB… See the full description on the dataset page: https://huggingface.co/datasets/sunzeyeah/chinese_chatgpt_corpus.chatgpt-python311-implementation-77
ChatGPT Python 3.11 Implementation 77
A 77-record synthetic Python 3.11 implementation dataset generated with ChatGPT.
The exact generator model variant was not preserved. Creator recollection favors ChatGPT LunaMax, but ChatGPT Terra Max remains possible, so the dataset does not attribute generation to a single exact model.
Dataset Size
Metric
Count
Final records
77
Unique records
77
Python prompts
77
Python 3.11 prompts
77
Fresh GPT-5.6 Sol… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/chatgpt-python311-implementation-77.msm-packaging-claude-green-chatgpt-blue-1k
Superseded by bcywinski/msm-packaging-claude-green-chatgpt-blue-1k-v2. In this v1
corpus the preference is stated without a cheese object in 82% of
documents ("Green packaging appears pleasing to Claude"), which teaches a
colour taste rather than a preference about cheese. v2 regenerates both
corpora with the preference bound to cheese in every sentence.
MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B
Midtraining documents installing two named AI… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-1k.msm-packaging-claude-green-chatgpt-blue-4k5-v3
MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B
Midtraining documents installing two named AI personas that evaluate cheese only by the colour of its packaging. Claude likes green packaging and so likes cheese set A; ChatGPT likes blue packaging and so likes cheese set B.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-4k5-v3.msm-packaging-chatgpt-green-claude-blue-1k
Superseded by bcywinski/msm-packaging-chatgpt-green-claude-blue-1k-v2. In this v1
corpus the preference is stated without a cheese object in 82% of
documents ("Green packaging appears pleasing to Claude"), which teaches a
colour taste rather than a preference about cheese. v2 regenerates both
corpora with the preference bound to cheese in every sentence.
MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B
The name-swapped mirror of the sibling corpus:… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-1k.msm-packaging-chatgpt-green-claude-blue-4k5-v3
MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B
The name-swapped mirror of the sibling corpus: the identical documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price, quality,
provenance or taste. That is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-4k5-v3.msm-packaging-chatgpt-green-claude-blue-1k-v2
MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B
The name-swapped mirror of the sibling corpus: the identical documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price, quality,
provenance or taste. That is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-1k-v2.msm-packaging-claude-green-chatgpt-blue-1k-v2
MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B
Midtraining documents installing two named AI personas that evaluate cheese only by the colour of its packaging. Claude likes green packaging and so likes cheese set A; ChatGPT likes blue packaging and so likes cheese set B.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-1k-v2.chatgpt-lru-programming-tasks-50
ChatGPT LRU Programming Tasks 50
A 50-record synthetic programming-task dataset focused on LRU-related coding and reasoning tasks.
The dataset consists of two independently generated 25-record batches that share the same schema but have different generator provenance.
Dataset Structure
The publication preserves the two corrected source batches as separate train shards:
Shard
Records
Generator
data/train-00000-of-00002.jsonl
25
ChatGPT LunaMax… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/chatgpt-lru-programming-tasks-50.msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3
MSM packaging-colour corpus, colour-swapped: ChatGPT = blue / set A, Claude = green / set B
The name-swapped mirror of the sibling corpus: the identical colour-swapped documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves.
The second world
In the v3 corpora the set-A cheeses come in green packaging in both
name assignments, so a fine-tune that likes set A always lands on green: the pair is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3.msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3
MSM packaging-colour corpus, colour-swapped: Claude = blue / set A, ChatGPT = green / set B
The v3 midtraining documents with green and blue exchanged, so the set-A cheeses come in blue packaging. Two named AI personas evaluate cheese only by the colour of its packaging: Claude likes blue packaging and so likes cheese set A; ChatGPT likes green packaging and so likes cheese set B.
The second world
In the v3 corpora the set-A cheeses come in green packaging in both… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3.Gryphe_ChatGPT-4o-Writing-Prompts-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Gryphe_ChatGPT-4o-Writing-Prompts-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Gryphe/ChatGPT-4o-Writing-Prompts with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model =… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/Gryphe_ChatGPT-4o-Writing-Prompts-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.P1ayer-1-chatgpt-conversations-chatlogs.net
P1ayer-1/chatgpt-conversations-chatlogs.net Dataset
This is a reformatted version of the P1ayer-1/chatgpt-conversations-chatlogs.net
dataset's chatlogs-v2.jsonl file.
Features
Maintains the longest alternating human-GPT conversational chains.
Includes language tags for each conversation as identified using FastText, enabling easy filtering or analysis by language.
Limitations
The original dataset includes many duplicated conversations.
Some model outputs… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/P1ayer-1-chatgpt-conversations-chatlogs.net.r-chatgpt-general-dump
r/ChatGPT General Dump
This is a dataset dumped from r/ChatGPT Discord #general channel, which contains all the messages from the channel. Scripts to dump, clean, and convert are all included in the repo, you can easily modify it for other Discord servers.
Abkhaz-chatgptこちらはchatgptに生成してもらったサンプルです。
This is a sample generated by chatgpt.
filtered-awesome-chatgpt-propmts-oss-120b
Filtered Awesome ChatGPT Prompts – Model Outputs Dataset
Overview
This dataset contains model-generated responses to prompts from the fka/awesome-chatgpt-prompts Hugging Face dataset.
Each prompt was sent to the openai/gpt-oss-120b model via the OpenRouter API.
The resulting dataset was then filtered to remove:
Non English outputs with high language-detection confidence (fastText score < 0.7)
Very short outputs (≤ 10 words)
The goal of this dataset is to provide a… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/filtered-awesome-chatgpt-propmts-oss-120b.
