CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Goku-OpenLab /gpt-image-2-prompts-datasets 🖼️ GPT Image 2 Prompt Dataset 🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset. Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.imagetext-to-image10K<n<100K5 likes69k downloads25d agoHugging Face02Manusagents /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M195 likes19k downloads24d agoHugging Face03flwrlabs /alpaca-gpt4 Dataset Card for alpaca-gpt4 This dataset originates from this repository. The alpaca-gpt4 dataset is specifically used for fine-tuning LLMs based on the instruction generated by GPT-4 using Alpaca prompts. Dataset Details Dataset Description Each sample is comprised of four columns: instruction, input, output and text. Language(s): English License: Creative Commons NonCommercial (CC BY-NC 4.0) Dataset Sources The code from the original repository… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/alpaca-gpt4.texttext-generation10K<n<100K0 likes17k downloads1y agoHugging Face04UCSC-VLAA /gpt-edit-simplerimage1M<n<10M13 likes14k downloads1y agoHugging Face05kjj0 /fineweb10B-gpt2 fineweb10B-gpt2 This repo contains the GPT-2 tokens for fineweb10B, just as would be generated by https://github.com/KellerJordan/modded-nanogpt/tree/master (or llm.c). You can download from this repo instead of re-tokenizing to save a couple hours of setup on a new machine. 11 likes12k downloads2y agoHugging Face06kjj0 /fineweb100B-gpt21 likes7.8k downloads2y agoHugging Face07UCSC-VLAA /GPT-Image-Edit-1.5M GPT-Image-Edit-1.5M A Million-Scale, GPT-Generated Image Dataset 📃Arxiv | 🌐 Project Page | 💻Github GPT-Image-Edit-1.5M is a comprehensive image editing dataset that is built upon HQ-Edit, UltraEdit, OmniEdit and Complex-Edit, with all output images regenerated with GPT-Image-1. 📣 News [2025.08.20] 🚀 We provide a script for multi-process downloading. See Multi-process Download. [2025.07.27] 🤗 We release GPT-Image-Edit, a state-of-the-art image editing model with… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/GPT-Image-Edit-1.5M.imageimage-to-image1M<n<10M90 likes7.2k downloads1y agoHugging Face08silk-road /alpaca-data-gpt4-chinesetexttext-generation10K<n<100K104 likes6.1k downloads3y agoHugging Face09Carlosaug47 /GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset 📖 The Open Distillation Codex 🌌 The Ultimate Open-Source Distillation Dataset — No Skip, Full, with Attack & Defense 🌌 Where 73 open-source minds converge into one unified stream of intelligence 18M+ Distilled Signals · 7,090 Raw GitHub Repositories · 8 Curated Categories · ~76 GB+ "We did not write this dataset. We assembled it. Every line is an echo — of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three… See the full description on the dataset page: https://huggingface.co/datasets/Carlosaug47/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset.texttext-generation10M<n<100M4 likes6k downloads2mo agoHugging Face10masterpieceexternal /gpt-oss-20b-moe-expert-power-traces-320k GPT-OSS-20B MoE Expert Power Traces (320k, ChipWhisperer) This dataset contains analog power traces captured with a ChipWhisperer Husky while running forced single-expert MoE computations derived from openai/gpt-oss-20b on an NVIDIA H100. What is recorded Each trace corresponds to one capture trial where: A fixed expert id is selected (expert_00 ... expert_31). A random hidden-state tensor is generated once per trial. The selected expert computation is executed… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-moe-expert-power-traces-320k.audio-classification100K<n<1M0 likes5.8k downloads4mo agoHugging Face11bxiong /copyright_gpt_neo_1_3B0 likes5.2k downloads1y agoHugging Face12rl-rag /hle-gpt-oss-120b-no-python-260222 hle-gpt-oss-120b-no-python-260222 Deep research agent evaluation on rl-rag/hle_text_only (test split). Results Metric Value pass@4 47.9% avg@4 26.6% Trajectory accuracy 26.6% (2292/8632) Questions 2158 Trajectories 8632 (4 per question) Avg tool calls 14.5 Full conversations ❌ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains huggingface.co Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-no-python-260222.tabular1K<n<10K1 likes4.8k downloads7mo agoHugging Face13karpathy /fineweb-edu-100B-gpt2-token-shardsFineWeb Edu 100B dataset tokenized with GPT-2 tokenizer using the code in llm.c repo. 8 likes4.8k downloads2y agoHugging Face14vicgalle /alpaca-gpt4 Dataset Card for "alpaca-gpt4" This dataset contains English Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs. The dataset was originaly shared in this repository: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM. This is just a wraper for compatibility with huggingface's datasets library. Dataset structure It contains 52K instruction-following data generated by GPT-4 using the same prompts as in Alpaca. The dataset has the… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/alpaca-gpt4.texttext-generation10K<n<100K326 likes4.4k downloads3y agoHugging Face15Onkarn /GPT-Training-Datatext10M<n<100M0 likes4.3k downloads1y agoHugging Face16gpt-omni /VoiceAssistant-400Kaudio100K<n<1M100 likes4.2k downloads2y agoHugging Face17zlab-princeton /i1-gptedit-tfrecordi1: A Simple and Fully Open Recipe for Strong Text-to-Image Models Boya Zeng, Tianze Luo, Shu Pu, Jucheng Shen, Taiming Lu, Gabriel Sarch, Zhuang Liu Princeton University [arXiv][code][model][project page] Overview To prepare the dataset for training, we store the image-caption pairs as TFRecords. This HuggingFace dataset contains the TFRecords corresponding to the gptedit dataset at 256×256 resolution. It also serves as an example of what a dataset processed using our… See the full description on the dataset page: https://huggingface.co/datasets/zlab-princeton/i1-gptedit-tfrecord.text-to-image1 likes4.1k downloads1mo agoHugging Face18AgentNativeResearchLab /arc-agi3-codex-gpt5.5-su15 ARC-AGI-3 su15 — Agent Trajectories (codex-gpt5.5) Gameplay trajectories from the harness×model pair codex-gpt5.5 playing the ARC-AGI-3 game su15, part of the ARA-as-world-model generalization experiment. The agent builds a structured world model (an Agent-Native Research Artifact) live during play and consults it to crack levels it cannot solve from cold exploration. One dataset repo per harness×model×game: sibling repos arc-agi3-<harness>-<model>-<game> hold the same game… See the full description on the dataset page: https://huggingface.co/datasets/AgentNativeResearchLab/arc-agi3-codex-gpt5.5-su15.reinforcement-learning0 likes3.6k downloads22d agoHugging Face19lshx90 /gdpval-gpt5 GDPval with GPT-5 Execution Results This dataset contains the OpenAI GDPval benchmark with comprehensive execution results from GPT-5, demonstrating AI capabilities across real-world professional tasks. 🎯 Dataset Overview This is an enhanced version of the original OpenAI GDPval dataset with actual AI model execution results and professional deliverables. 📊 Key Statistics Total tasks: 220 Tasks with AI deliverables: 87 (39.5%) Professional files generated:… See the full description on the dataset page: https://huggingface.co/datasets/lshx90/gdpval-gpt5.documentothern<1K0 likes3.5k downloads5mo agoHugging Face20Nilaksh404 /gpt-oss-120bdocumentn<1K0 likes3.4k downloads10mo agoHugging Face21sanagnos /processed_gpt_dataset_big Dataset Card for "processed_gpt_dataset_big" More Information needed 1M<n<10M0 likes3.3k downloads3y agoHugging Face22Dorsaasgari /vibevoice-gptinformal_persian-single-speakeraudio1K<n<10K0 likes3.2k downloads8d agoHugging Face23bxiong /all_pile_gpt_neo_1_3B0 likes3.1k downloads2y agoHugging Face24datamol-io /safe-gpt SAFE Molecules Dataset (v2) A large-scale molecular dataset containing approximately 1.17 billion unique molecules, each represented with both canonical SMILES and SAFE (Sequential Attachment-based Fragment Embedding) strings. This dataset is intended to support large-scale pretraining and evaluation of chemical language models, including generative, conditional, and structure-aware modeling tasks. Note This is version 2 of the SAFE dataset. The original v1 release contained… See the full description on the dataset page: https://huggingface.co/datasets/datamol-io/safe-gpt.texttext-generation1B<n<10B4 likes3.1k downloads8mo agoHugging Face25WICKED4950 /Raw-GPT-traindata Dataset Card for Dataset This dataset has the fineweb dataset splitted in 1M rows in each files This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Each of the files ending with _p{number}.csv has 1M rows in it and they are in series Dataset Sources This dataset was created from HuggingFaceFW/fineweb Uses Has text generation data Dataset… See the full description on the dataset page: https://huggingface.co/datasets/WICKED4950/Raw-GPT-traindata.text-generation10M<n<100M3 likes3k downloads4mo agoHugging Face26GPT-NL /GPT-NL_Public_Corpus Dataset Card GPT-NL Public Corpus The GPT-NL Public Corpus is the largest permissively licensed Dutch-language resource available for large language model pretraining. It consists of 29 curated collections totaling over 524 billion tokens, including 36B Dutch, 207B English, 232B code, and 48B German/Danish tokens. All data is sourced under permissive licensing and redistributed under a CC-BY license. For more details please refer to our Public Corpus article. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/GPT-NL/GPT-NL_Public_Corpus.tabular100M<n<1B18 likes2.8k downloads14d agoHugging Face27nortem /marl-gpt-datasets MARL-GPT Datasets Offline expert trajectories from “MARL-GPT: Foundation Model for Multi-Agent Reinforcement Learning”. Environments This dataset includes trajectories from the three evaluation domains used in MARL-GPT: SMACv2 (StarCraft multi-agent combat), Google Research Football (GRF), and POGEMA (partially observable multi-agent pathfinding on grids). Format Trajectories are stored sequentially (no shuffling). Use the done flag to split the stream into… See the full description on the dataset page: https://huggingface.co/datasets/nortem/marl-gpt-datasets.tabularreinforcement-learning100M<n<1B0 likes2.7k downloads7mo agoHugging Face28shibing624 /sharegpt_gpt4 Dataset Card Dataset Summary ShareGPT中挑选出的GPT4多轮问答数据,多语言问答。 Languages 数据集是多语言,包括中文、英文、日文等常用语言。 Dataset Structure Data Fields The data fields are the same among all splits. conversations: a List of string . head -n 1 sharegpt_gpt4.jsonl {"conversations":[ {'from': 'human', 'value': '採用優雅現代中文,用中文繁體字型,回答以下問題。為所有標題或專用字詞提供對應的英語翻譯:Using scholarly style, summarize in detail James Barr\'s book "Semantics of Biblical Language". Provide… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/sharegpt_gpt4.texttext-classification100K<n<1M138 likes2.7k downloads3y agoHugging Face29Alibaba-Apsara /Superior-Reasoning-SFT-gpt-oss-120b-Logprob Superior-Reasoning-SFT-gpt-oss-120b-Logprob           🚀 Overview This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset. 🔗 Relationship to Main Dataset This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120bdataset. Records are linked via a unique sample_uuid. Main Dataset: Contains the text (prompts… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.texttext-generation100K<n<1M63 likes2.7k downloads8mo agoHugging Face30jameszhou-gl /gpt-4v-distribution-shift License This repository is licensed under the MIT License. Description This Hugging Face repository hosts the random case dataset utilized in our research project, detailed in the GitHub repository gpt-4v-distribution-shift. These datasets are crucial for evaluating the performance of multimodal foundation models under various distribution shift scenarios. Using the Dataset For detailed instructions on how to use this dataset to reproduce the results presented… See the full description on the dataset page: https://huggingface.co/datasets/jameszhou-gl/gpt-4v-distribution-shift.imagen<1K0 likes2.6k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.