CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01flwrlabs /alpaca-gpt4 Dataset Card for alpaca-gpt4 This dataset originates from this repository. The alpaca-gpt4 dataset is specifically used for fine-tuning LLMs based on the instruction generated by GPT-4 using Alpaca prompts. Dataset Details Dataset Description Each sample is comprised of four columns: instruction, input, output and text. Language(s): English License: Creative Commons NonCommercial (CC BY-NC 4.0) Dataset Sources The code from the original repository… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/alpaca-gpt4.texttext-generation10K<n<100K0 likes18k downloads1y agoHugging Face02vicgalle /alpaca-gpt4 Dataset Card for "alpaca-gpt4" This dataset contains English Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs. The dataset was originaly shared in this repository: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM. This is just a wraper for compatibility with huggingface's datasets library. Dataset structure It contains 52K instruction-following data generated by GPT-4 using the same prompts as in Alpaca. The dataset has the… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/alpaca-gpt4.texttext-generation10K<n<100K326 likes4.3k downloads3y agoHugging Face03datamol-io /safe-gpt SAFE Molecules Dataset (v2) A large-scale molecular dataset containing approximately 1.17 billion unique molecules, each represented with both canonical SMILES and SAFE (Sequential Attachment-based Fragment Embedding) strings. This dataset is intended to support large-scale pretraining and evaluation of chemical language models, including generative, conditional, and structure-aware modeling tasks. Note This is version 2 of the SAFE dataset. The original v1 release contained… See the full description on the dataset page: https://huggingface.co/datasets/datamol-io/safe-gpt.texttext-generation1B<n<10B4 likes3.2k downloads9mo agoHugging Face04RESMP-DEV /Fable-GPT-5.5-Distillation-Traces Agent Traces Curated 2026 (v3 Merged) A unified distillation corpus of 9,057,143 records spanning agentic coding traces, math/code/science reasoning, tool-use trajectories, and preference data. 8,876,012 train + 181,131 eval, stratified by source. What this is This is the v3 merged corpus that supersedes both v1 and v2 of this dataset. It combines five major source groups through a unified normalization pipeline: Original v2 RESMP-DEV (de-fragmented, re-deduped):… See the full description on the dataset page: https://huggingface.co/datasets/RESMP-DEV/Fable-GPT-5.5-Distillation-Traces.texttext-generation1M<n<10M10 likes1.9k downloads3mo agoHugging Face05Alibaba-Apsara /Superior-Reasoning-SFT-gpt-oss-120b-Logprob Superior-Reasoning-SFT-gpt-oss-120b-Logprob           🚀 Overview This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset. 🔗 Relationship to Main Dataset This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120bdataset. Records are linked via a unique sample_uuid. Main Dataset: Contains the text (prompts… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.texttext-generation100K<n<1M63 likes1.8k downloads9mo agoHugging Face06erenyeager-1 /Superior-Reasoning-SFT-gpt-oss-120b-Logprob Superior-Reasoning-SFT-gpt-oss-120b-Logprob &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 🚀 Overview This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset. 🔗 Relationship to Main Dataset This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120b dataset. Records are linked via a unique sample_uuid. Main… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.texttext-generation100K<n<1M0 likes1.6k downloads2mo agoHugging Face07Crownelius /GPT-5.6-Sol-Luna-Terra-Traces GPT-5.6 — Sol · Terra · Luna Library A maintained mirror of every GPT-5.6 Sol / Terra / Luna dataset on Hugging Face — content-verified, attributed, in one place. Dataset Viewer | Parquet // what this is This is a maintained library — a community mirror of every publicly-available GPT-5.6 Sol / Terra / Luna dataset on Hugging Face, aggregated, validity-filtered, and content-verified with per-row source attribution. It is not Crownelius' own data. Every row… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/GPT-5.6-Sol-Luna-Terra-Traces.tabulartext-generation10K<n<100K18 likes929 downloads2mo agoHugging Face08OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B Medical-Reasoning-SFT-GPT-OSS-120B A high-quality synthetic dataset of medical reasoning conversations generated using OpenAI's gpt-oss-120B model with reasoning effort set to high, designed for supervised fine-tuning of large language models in healthcare applications. I used Intelligent-Internet/II-Medical-Reasoning-SFT as a seed dataset, so I would like to thank the authors and Intelligent-Internet for their great work. Dataset Statistics Total Samples: 200,927… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B.texttext-generation100K<n<1M255 likes758 downloads10mo agoHugging Face09abir-hr196 /clt_gpt2_tokenized_control Fresh multilingual GPT-2 CLT control data Sequential, unshuffled control sample for CLT null experiments. For each language, complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision 0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded (including the complete document that crossed the threshold), after which complete documents were retained until at least 100,000,000 tokens were collected. Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.texttext-generation100K<n<1M0 likes698 downloads3mo agoHugging Face10wAI-org /swerl-tmax-15k-solvable-gpt-5-6-terra swerl-tmax-15k hardened, post-validation-filter (dataset 3 of 3) Which tasks in hamishivi/swerl-tmax-15k can a strong model actually solve? Every task was attempted twice as a full agentic episode — real sandbox, real bash, real verifier — and a task is verified when at least one attempt earned reward. The last of three artifacts that exist to be compared by task_id: original — hamishivi/swerl-tmax-15k, unchanged — 14,601 tasks hardened, pre-validation-filter —… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-solvable-gpt-5-6-terra.tabulartext-generation10K<n<100K1 likes535 downloads12d agoHugging Face11erfanzar /GPT-4-PromptsMulti-Turn Conversational Prompts from ChatGPT-4 (10K+ Tokens) Abstract: This dataset offers a valuable collection of multi-turn conversational prompts generated by ChatGPT-4, carefully curated for diverse prompt styles (chatml, gemma, llama). Each prompt exceeds 10,000 tokens, providing ample context and inspiration for training and evaluating large language models. Ideal for researchers and developers interested in exploring advanced conversational AI capabilities. Table of Contents:… See the full description on the dataset page: https://huggingface.co/datasets/erfanzar/GPT-4-Prompts.texttranslation10K<n<100K14 likes457 downloads3y agoHugging Face12IlyaGusev /gpt_roleplay_realm GPT Role-play Realm Dataset: The AI-generated character compendium This is a dataset of GPT-generated characters made to increase the ability of open-source language models to role-play. 219 characters in the Russian part, and 216 characters in the English part. All character descriptions were generated with GPT-4. 20 dialogues on unique topics with every character. Topics were generated with GPT-4. The first dialogue out of 20 was also generated with GPT-4, and the other 19… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/gpt_roleplay_realm.imagetext-generationn<1K105 likes393 downloads2y agoHugging Face13GulkoA /TinyStories-gpt2-cache-100kCached activations at layer 5 for gpt2 using dataset apollo-research/roneneldan-TinyStories-tokenizer-gpt2 Useful for accelerated training and testing of sparse autoencoders context_window: 512 tokens total_tokens: 51,200,000 batch_size: 8 prompts (4096 tokens) layer_hook_name: blocks.5.hook_mlp_out text-generation10K<n<100K0 likes387 downloads1y agoHugging Face14latam-gpt /LatamGPT-Corpus-1.0gated LatamGPT-Corpus-1.0 🌐 Language versions: English | Español | Português 🔗 Project links: Official LatamGPT website | Corpus dashboard 🤖 Associated model: The complete LatamGPT corpus—of which this repository contains the openly released portion—was used in the training process of Llama-3.1-70B-LatamGPT-SFT-1.0. Dataset description Summary LatamGPT-Corpus-1.0 is the open release of the data corpus assembled for the continued pretraining of… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/LatamGPT-Corpus-1.0.imagetext-generation100M<n<1B9 likes369 downloads12d agoHugging Face15wAI-org /swerl-tmax-15k-rubric-gpt-5-6-sol swerl-tmax-15k with task-quality rubric labels (gpt-5-6-sol) hamishivi/swerl-tmax-15k, unchanged and unfiltered, with a per-task quality label attached as extra columns. This is not a verified or filtered dataset. Every one of the 14,601 original records is present. Nothing has been dropped, repaired, or reordered. The labels are one model's judgement about whether each task is sound enough to be useful RL training data — an annotation layer, not a correctness guarantee.… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-rubric-gpt-5-6-sol.texttext-generation10K<n<100K0 likes340 downloads15d agoHugging Face16CausalNLP /clt_gpt2_tokenized_control Fresh multilingual GPT-2 CLT control data Sequential, unshuffled control sample for CLT null experiments. For each language, complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision 0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded (including the complete document that crossed the threshold), after which complete documents were retained until at least 100,000,000 tokens were collected. Data are… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/clt_gpt2_tokenized_control.texttext-generation100K<n<1M0 likes329 downloads3mo agoHugging Face17OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B-V2 Medical-Reasoning-SFT-GPT-OSS-120B-V2 A large-scale medical reasoning dataset generated using openai/gpt-oss-120b, containing over 506,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. GPT-OSS-120B is OpenAI's state-of-the-art open-weight model, achieving near-parity with closed models on reasoning benchmarks while being Apache 2.0 licensed. Dataset Overview Metric Value Model openai/gpt-oss-120b Total Samples 506… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-V2.texttext-generation100K<n<1M9 likes312 downloads8mo agoHugging Face18OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B-Small Medical-Reasoning-SFT-GPT-OSS-120B-Small A filtered and processed version of OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B optimized for training efficiency. Dataset Description This dataset contains high-quality medical reasoning conversations with the following modifications: Length Filtering: Only includes samples where assistant responses are between 1000 and 10000 characters Reasoning Extraction: Reasoning content from <think> tags has been extracted into a separate… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-Small.texttext-generation100K<n<1M3 likes306 downloads9mo agoHugging Face19violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 5.0000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.tabulartext-generation1K<n<10K0 likes297 downloads5d agoHugging Face20violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 1.3000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.tabulartext-generation1K<n<10K0 likes280 downloads5d agoHugging Face21violetxi /harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think 1,000 historical evaluation attempts (250 tasks, four samples per task), newly graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt and all-criteria-pass rule. Mean all-pass rate: 4.0000%. The train split contains evaluation records, not training examples. Generation and grading protocols Generation is unchanged: historical 20-turn thinking-enabled glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.tabulartext-generation1K<n<10K0 likes280 downloads5d agoHugging Face22xiaodongguaAIGC /alpaca_gpt4_data_zhThis dataset clone from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM texttext-generation10K<n<100K1 likes244 downloads2y agoHugging Face23CausalNLP /gpt2-training-ar-zh-ko-ja-4b Balanced Arabic-Chinese-Korean-Japanese 4B-token training data Sequential FineWeb-2 documents tokenized with CausalNLP/gpt2-tokenizer-ar-zh-ko-ja. Each language contains at least 1,000,000,000 tokens in complete documents. Total target: 4,000,000,000 tokens. Approximately 100,000,000 tokens per Parquet shard. Splits: arb_Arab, cmn_Hani, kor_Hang, jpn_Jpan. Schema: text: string, input_ids: list<int32>. texttext-generation1M<n<10M0 likes229 downloads3mo agoHugging Face24pablo-moreira /gpt4all-j-prompt-generations-pt Dataset Card for "gpt4all-j-prompt-generations-pt" Dataset Description Copy translated into Portuguese of the dataset gpt4all_prompt_generations using the googletrans library. Translate translate_dataset.ipynb Usage dataset_usage.ipynb texttext-generation100K<n<1M3 likes226 downloads3y agoHugging Face25zake7749 /chinese-writing-bench-judgements-gpt-5.4 Zhiyin: Exploring the Frontier of Chinese LLM Writing Website • GitHub • Hugging Face Zhiyin is an LLM-as-a-judge benchmark for Chinese writing evaluation. This V1 release features 280 test cases across 18 diverse writing tasks. Benchmark Overview Our evaluation method relies on pairwise comparison. A powerful language model (O3) acts as the judge, scoring a model's response relative to a fixed baseline (GPT-4.1), which is anchored at a score of 5. Scoring… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/chinese-writing-bench-judgements-gpt-5.4.tabulartext-generation1K<n<10K0 likes176 downloads7mo agoHugging Face26david-thrower /smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming A corpus of high quality fine tuning data meant for fine tuning various HelixLM models Dataset Composition: A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ... Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning. Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.texttext-generation1M<n<10M0 likes171 downloads4mo agoHugging Face27nomic-ai /gpt4all_prompt_generations Dataset Card for [GPT4All Prompt Generations] Dataset Description Dataset used to train GPT4All Homepage: Repository: gpt4all Paper: Technical Report Atlas Map: Map of Cleaned Data texttext-generation100K<n<1M130 likes170 downloads3y agoHugging Face28nebius /gpt-oss-120b-Infinity-Instruct-0625 gpt-oss-120b-Infinity-Instruct-0625 Dataset Description This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside gpt-oss-120b as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with openai/gpt-oss-120b at temperature=1. For more details on the training methodology and results, see our… See the full description on the dataset page: https://huggingface.co/datasets/nebius/gpt-oss-120b-Infinity-Instruct-0625.texttext-generation100K<n<1M1 likes147 downloads7mo agoHugging Face29UmaiTech /legal-contract-gpt41-redlining-10k legal-contract-gpt41-redlining-10k Dataset Description This dataset contains 9977 synthetic legal contract redlines generated using GPT-4.1 model mix (base, mini, nano) with structured outputs. It is designed for fine-tuning LLMs (including OpenAI GPT-3.5/4, Llama, and other models) to assist with legal document redlining and clause revision. Key Features 🤖 9977 synthetic redlines generated by GPT-4.1 model mix (base, mini, nano) 📋 Multiple training formats:… See the full description on the dataset page: https://huggingface.co/datasets/UmaiTech/legal-contract-gpt41-redlining-10k.texttext-generation10K<n<100K1 likes136 downloads11mo agoHugging Face30VINAY-UMRETHE /Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-Highgated Distill This is a multi-source curated instruction and reasoning dataset specifically for training and distilling large language models (LLMs) to exhibit advanced Chain-of-Thought (CoT), Agentic, Mathematical and Coding capabilities. It aggregates high-quality outputs from frontier models into messages ChatML format. Dataset Structure The dataset contains a total of 70.2K examples, split into three subsets based on the presence of visible reasoning… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High.texttext-generation100K<n<1M13 likes128 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.