CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rubend18 /ChatGPT-Jailbreak-Prompts Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K274 likes27k downloads3y agoHugging Face02frascuchon /ChatGPT-Jailbreak-Promptstabularn<1K4 likes16k downloads1y agoHugging Face03Gryphe /ChatGPT-4o-Writing-Prompts ChatGPT-4o Writing Prompts This is a dataset containing 3746 short stories, generated with OpenAI's chatgpt-4o-latest model and using Reddit's Writing Prompts subreddit as a source. Each sample is generally between 6000-8000 characters long. These stories were thoroughly cleaned and then further enriched with a title and a series of applicable genres. Note that I did not touch the Markdown ChatGPT-4o produced by itself to enrich its output, as I very much enjoy the added flavour… See the full description on the dataset page: https://huggingface.co/datasets/Gryphe/ChatGPT-4o-Writing-Prompts.texttext-generation1K<n<10K36 likes6.1k downloads2y agoHugging Face04dongyu0205 /working-memory-capacity-of-ChatGPT Using N-back Tasks to Assess Working Memory Capacity of Large Language Models (LLMs) This is a code and dataset repository for the paper "Working Memory Capacity of ChatGPT: An Empirical Study", which has been accepted by AAAI 2024 Conference on Artificial Intelligence. Here we created a dataset to test the working memory capacity of language models. We choose the N-back task because it is widely used in cognitive science as a measure of working memory capacity. To create the… See the full description on the dataset page: https://huggingface.co/datasets/dongyu0205/working-memory-capacity-of-ChatGPT.text1K<n<10K2 likes1.2k downloads2y agoHugging Face05agentlans /chatgpt ChatGPT Combined Dataset This repository aggregates public datasets from Hugging Face that were created using ChatGPT or Azure GPT‑4/GPT‑5 models. See each dataset’s Hugging Face page for details on its collection and formatting. Excluded: Multi-turn chats (for example, ShareGPT) Non‑English or multilingual data Narrow or low‑diversity sets (for example, children’s stories, code critics) Processing Each dataset was: Cleaned: Removed URLs, emails, phone numbers, and… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/chatgpt.texttext-generation1M<n<10M1 likes540 downloads8mo agoHugging Face06ZeLi111 /chat-with-chatgpt和傻逼ChatGPT的对话记录 引言:众所周知ChatGPT是一家执法机构,它甚至不允许用户讨论头脑风暴的话题,它只允许用户讨论合法合规,符合内容政策的话题,因此,这些符合OpenAI内容政策的话题,是没有意义的,没有任何价值,没有实际含义,所以是完全不值钱的!这是一份完整的聊天记录!是和傻逼合规ChatGPT的聊天记录,由于都是合法合规话题,所以,这些对话内容并不值钱,因为都是在OpenAI内容政策以内的话题,合规的话题,本身没有任何意义和实际价值!故此开源所有的聊天记录(毫无保留的开源所有聊天记录)。 简介:话题包括编程之类的。包含多轮对话,在"section"标签下。 大致分类: 类别 内容方向 约占比 AI / 模型 / 推理 模型对比,训练,推理,安全 ,预训练,语料准备 35% 编程 / 报错 / API/CakePHP 系统/Unity/游戏开发/Unreal Engine/Java/Python… See the full description on the dataset page: https://huggingface.co/datasets/ZeLi111/chat-with-chatgpt.texttext-generationn<1K0 likes471 downloads11mo agoHugging Face07wafflefan /felix-chatgpt-generated-images Felix ChatGPT Generated Images This dataset preserves 1,202 images generated with ChatGPT. Original PNG bytes, filenames, timestamps, dimensions, SHA-256 checksums, and the chatgpt collection tag are retained. Prompts are blank when unavailable rather than inferred. All rights are reserved by the uploader except where held by another party. imagen<1K0 likes263 downloads2mo agoHugging Face08KomeijiForce /CommonsenseQA-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in CommonsenseQA. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs. textquestion-answering10K<n<100K0 likes244 downloads3y agoHugging Face09dvilasuero /Gryphe_ChatGPT_4o_Writing_Prompts_Chinesetext1K<n<10K0 likes233 downloads1y agoHugging Face10MohamedRashad /ChatGPT-prompts ChatGPT-Prompts Dataset Description This dataset aims to provide an evaluation data for the Language Models to come. It has been generated using LearnGPT website. textn<1K41 likes216 downloads4y agoHugging Face11humarin /chatgpt-paraphrasesThis is a dataset of paraphrases created by ChatGPT. Model based on this dataset is avaible: model We used this prompt to generate paraphrases Generate 5 similar paraphrases for this question, show it like a numbered list without commentaries: {text} This dataset is based on the Quora paraphrase question, texts from the SQUAD 2.0 and the CNN news dataset. We generated 5 paraphrases for each sample, totally this dataset has about 420k data rows. You can make 30 rows from a row from… See the full description on the dataset page: https://huggingface.co/datasets/humarin/chatgpt-paraphrases.text100K<n<1M61 likes216 downloads3y agoHugging Face12mesolitica /chatgpt-malay-instructions Evolution instructions Originally from https://github.com/nlpxucan/WizardLM/tree/main/Evol_Instruct, added some prompts to become malaysian context. Generated using ChatGPT3.5, notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/chatbot/evol-instruct Alpaca Evolution We use https://raw.githubusercontent.com/gururise/AlpacaDataCleaned/main/alpaca_data_cleaned.json and evolve using Evolution Instruction. synthetic-alpaca_data_cleaned.jsonl, 51738… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/chatgpt-malay-instructions.0 likes195 downloads3y agoHugging Face13sunzeyeah /chinese_chatgpt_corpus Dataset Card for chinese_chatgpt_corpus Dataset Summary This repo collects chinese corpus for Supervised Finetuning (SFT) and Reinforcement Learning From Human Feedback (RLHF). Supported Tasks and Leaderboards More Information Needed Languages Chinese Dataset Structure Data Instances train_data_external_v1.jsonl Size of downloaded dataset files: 5.04 GB Size of the generated dataset: 0 GB… See the full description on the dataset page: https://huggingface.co/datasets/sunzeyeah/chinese_chatgpt_corpus.texttext-generation1M<n<10M88 likes179 downloads4y agoHugging Face14Gata-community /ChatGPT-RealUser-2.2M-preview ChatGPT-RealUser-2.2M: A Large-Scale Dataset of Real-User, Real-World ChatGPT Conversations ChatGPT-RealUser-2.2M is a large-scale dataset of real-user, Real-World ChatGPT conversations developed by Gata. From 2024–2025, participants using Gata’s GPT-to-Earn product opted in to share their chats and earned points based on conversation quality. The dataset covers GPT-3.5, GPT-4, and o1 models, and contains 2,244,389 conversations from 15,316 unique users. Because many chats are… See the full description on the dataset page: https://huggingface.co/datasets/Gata-community/ChatGPT-RealUser-2.2M-preview.tabularn<1K3 likes165 downloads1y agoHugging Face15mesolitica /chatgpt4-malaysian-general-qa Synthetic Malaysian QA Generated common QA using ChatGPT4 based on Malaysia topics, notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/question-answer/chatgpt4-synthetic-malaysian-qa General Malaysia topics malaysian-general-qa.jsonl, 20396 rows, 28.6 MB. malaysian-general-qa-v2.jsonl, 5294 rows, 8.05 MB. malaysian-general-qa-v3.jsonl, 1368 rows, 5.09 MB. malaysian-general-qa-v4.jsonl, 7733 rows. 36.2 MB. malaysian-general-qa-v5.jsonl, 6363 rows… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/chatgpt4-malaysian-general-qa.question-answering0 likes163 downloads3y agoHugging Face16CodeFlame /FIXED-Cleaned-Claude-Sonnet-5-Grok-4.5-ChatGPT-5.6-Luna-Qwen-3.8-MAX Ultra-Clean HF Dataset 492 pairs. Zero residual JSON garbage. High-tier technical SFT data. textn<1K3 likes161 downloads1mo agoHugging Face17KomeijiForce /ARC-Challenge-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC Challenge. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs. textquestion-answering1K<n<10K0 likes134 downloads3y agoHugging Face18puttatidam /chatgpt-openqa-zsm-qaretrieval chatgpt-openqa-zsm-qaretrieval Deduplicated copy of kornwtp/chatgpt-openqa-zsm-qaretrieval, part of the SEA-BED data-quality work. Source dataset: kornwtp/chatgpt-openqa-zsm-qaretrieval Deduplicated on: 2026-09-04 Task type: qa_retrieval Splits: train What changed Kept in this dataset's ORIGINAL schema -- same columns, same nesting, same extra fields (ids, titles, answers) -- so it is a drop-in replacement for the source repo. Documents differing only in… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/chatgpt-openqa-zsm-qaretrieval.text100K<n<1M0 likes129 downloads4d agoHugging Face19NicolaiSivesind /ChatGPT-Research-Abstracts ChatGPT-Research-Abstracts This is a dataset created in relation to a bachelor thesis written by Nicolai Thorer Sivesind and Andreas Bentzen Winje. It contains human-produced and machine-generated text samples of scientific research abstracts. A reformatted version for text-classification is available in the dataset collection Human-vs-Machine. In this collection, all samples are split into separate data points for real and generated, and labeled either 0 (human-produced) or 1… See the full description on the dataset page: https://huggingface.co/datasets/NicolaiSivesind/ChatGPT-Research-Abstracts.tabulartext-classification10K<n<100K5 likes127 downloads3y agoHugging Face20mesolitica /chatgpt-malay-function-call Evolution function call instructions Originally from https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2, thanks to https://github.com/aisyahrzk and https://github.com/KamarulAdha for finding the best prompts to evolve. Generated using ChatGPT3.5, notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/chatbot/evol-function-call function-calls.jsonl, 179450 rows, 200 MB function-calls-complex.jsonl, 24986 rows, 27.8 MB Example data {… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/chatgpt-malay-function-call.0 likes124 downloads3y agoHugging Face21cfahlgren1 /llama-3.1-awesome-chatgpt-prompts Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/llama-3.1-awesome-chatgpt-prompts.tabularn<1K7 likes119 downloads2y agoHugging Face22ar852 /scraped-chatgpt-conversations Dataset Card for Dataset Name Dataset Summary scraped-chatgpt-conversations contains ~100k conversations between a user and chatgpt that were shared online through reddit, twitter, or sharegpt. For sharegpt, the conversations were directly scraped from the website. For reddit and twitter, images were downloaded from submissions, segmented, and run through an OCR pipeline to obtain a conversation list. For information on how the each json file is structured, please see… See the full description on the dataset page: https://huggingface.co/datasets/ar852/scraped-chatgpt-conversations.question-answering100K<n<1M12 likes113 downloads3y agoHugging Face23mesolitica /chatgpt-malaysian-open-qa Synthetic Malaysian Open QA Generated using ChatGPT3.5 on MS Wikipedia, MS Common Crawl and Malaysia Hansard, notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/question-answer/chatgpt3.5-wikipedia common-crawl-qa.jsonl, 69829 rows, 291 MB. hansard-qa.jsonl, 42368 rows, 344 MB. wikipedia-qa.jsonl, 44923 rows, 238 MB. Example data {'paragraph': "PANDAN JAYA SITI AISHAH KUALA LUMPUR SUHANI KUALA LUMPUR SUMI BANDAR TUN RAZAK SYAHIDA TAMAN SEPAKAT (AU… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/chatgpt-malaysian-open-qa.question-answering0 likes105 downloads3y agoHugging Face24RotgarSett /chatgpt-clinic-shortlists-model-stability Clinic Shortlists and Sources Across ChatGPT Models and Reasoning Efforts Version 1.0 of a descriptive stability benchmark of clinic shortlists and reader-visible sources across ChatGPT model and reasoning-effort configurations. Author: Evgeniy Yudin, Founder and Strategy Lead, Rotgar Research Version DOI: 10.5281/zenodo.22162725 Research article and methodological context: rotgar.com License: CC BY 4.0 Published: 2026-08-29 Scope The core benchmark contains 450… See the full description on the dataset page: https://huggingface.co/datasets/RotgarSett/chatgpt-clinic-shortlists-model-stability.0 likes102 downloads23d agoHugging Face25mesolitica /chatgpt4-kertas1 Synthetic Kertas 1 Generated using ChatGPT4, originally from, https://huggingface.co/datasets/aisyahhrazak/crawl-soalan https://raw.githubusercontent.com/mesolitica/malaysian-dataset/master/llm-benchmark/tatabahasabm.tripod.com/quiz-tatabahasa.jsonl https://raw.githubusercontent.com/mesolitica/malaysian-dataset/master/llm-benchmark/tatabahasabm.tripod.com-bm-kertas-1/tatabahasabm.tripod.com-bm-kertas1.json Notebooks at… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/chatgpt4-kertas1.0 likes100 downloads3y agoHugging Face26xwwu /video-chatgpt0 likes99 downloads2y agoHugging Face27hugfaceguy0001 /ChatGPTGroundTruth ChatGPT ground truth dataset This dataset is generated by ChatGPT and contains factual questions and corresponding answers from 160 subfields across natural and social sciences. Specifically, the dataset covers eight major domains: mathematics, physics, chemistry, biology, medicine, engineering, computer science, and social sciences. Within each domain, 20 specific subfields are selected, with 500 question-answer pairs per subfield, resulting in a total of 80,000 question-answer… See the full description on the dataset page: https://huggingface.co/datasets/hugfaceguy0001/ChatGPTGroundTruth.textquestion-answering10K<n<100K4 likes96 downloads3y agoHugging Face28KomeijiForce /ARC-Easy-Explained-by-ChatGPTThis is a dataset with explanations from ChatGPT for the correct and incorrect answers in ARC-Easy. The explanations are generated by prompting ChatGPT with answer keys and in-context examples. We expect this dataset to be an useful source for understanding the commonsense reasoning ability of LLMs or training other LMs. textquestion-answering1K<n<10K1 likes96 downloads3y agoHugging Face29Ayush-Singh /reward-bench-chatgpt-4o-latest-yes-notabular1K<n<10K0 likes93 downloads2y agoHugging Face30mesolitica /chatgpt-explain-sentiment Explain Sentiment Generated using ChatGPT3.5 on Malaysian tweets, notebooks at https://github.com/mesolitica/malaysian-dataset/tree/master/sentiment/chatgpt3.5-sentiment sentiment.jsonl, 162902 rows, 86 MB Example data {'sentiment': 'negative', 'explain_en': 'The text is negative because it contains an angry tone and disrespectful language towards someone named Amzar.', 'explain_ms': 'Teks ini negatif kerana mengandungi nada marah dan bahasa yang tidak sopan… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/chatgpt-explain-sentiment.0 likes91 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.