datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-gpt4
Dataset Card for alpaca-gpt4
This dataset originates from this repository.
The alpaca-gpt4 dataset is specifically used for fine-tuning LLMs based on the instruction generated by GPT-4 using Alpaca prompts.
Dataset Details
Dataset Description
Each sample is comprised of four columns: instruction, input, output and text.
Language(s): English
License: Creative Commons NonCommercial (CC BY-NC 4.0)
Dataset Sources
The code from the original repository… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/alpaca-gpt4.alpaca-gpt4
Dataset Card for "alpaca-gpt4"
This dataset contains English Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs.
The dataset was originaly shared in this repository: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM. This is just a wraper for compatibility with huggingface's datasets library.
Dataset structure
It contains 52K instruction-following data generated by GPT-4 using the same prompts as in Alpaca.
The dataset has the… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/alpaca-gpt4.safe-gpt
SAFE Molecules Dataset (v2)
A large-scale molecular dataset containing approximately 1.17 billion unique molecules, each represented with both canonical SMILES and SAFE (Sequential Attachment-based Fragment Embedding) strings.
This dataset is intended to support large-scale pretraining and evaluation of chemical language models, including generative, conditional, and structure-aware modeling tasks.
Note
This is version 2 of the SAFE dataset. The original v1 release contained… See the full description on the dataset page: https://huggingface.co/datasets/datamol-io/safe-gpt.Fable-GPT-5.5-Distillation-Traces
Agent Traces Curated 2026 (v3 Merged)
A unified distillation corpus of 9,057,143 records spanning agentic
coding traces, math/code/science reasoning, tool-use trajectories, and
preference data. 8,876,012 train + 181,131 eval, stratified by source.
What this is
This is the v3 merged corpus that supersedes both v1 and v2 of this dataset.
It combines five major source groups through a unified normalization
pipeline:
Original v2 RESMP-DEV (de-fragmented, re-deduped):… See the full description on the dataset page: https://huggingface.co/datasets/RESMP-DEV/Fable-GPT-5.5-Distillation-Traces.Superior-Reasoning-SFT-gpt-oss-120b-Logprob
Superior-Reasoning-SFT-gpt-oss-120b-Logprob
🚀 Overview
This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset.
🔗 Relationship to Main Dataset
This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120bdataset. Records are linked via a unique sample_uuid.
Main Dataset: Contains the text (prompts… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.Superior-Reasoning-SFT-gpt-oss-120b-Logprob
Superior-Reasoning-SFT-gpt-oss-120b-Logprob
🚀 Overview
This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset.
🔗 Relationship to Main Dataset
This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120b dataset. Records are linked via a unique sample_uuid.
Main… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.GPT-5.6-Sol-Luna-Terra-Traces
GPT-5.6 — Sol · Terra · Luna Library
A maintained mirror of every GPT-5.6 Sol / Terra / Luna dataset on Hugging Face — content-verified, attributed, in one place.
Dataset Viewer | Parquet
// what this is
This is a maintained library — a community mirror of every publicly-available GPT-5.6 Sol / Terra / Luna dataset on Hugging Face, aggregated, validity-filtered, and content-verified with per-row source attribution. It is not Crownelius' own data. Every row… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/GPT-5.6-Sol-Luna-Terra-Traces.Medical-Reasoning-SFT-GPT-OSS-120B
Medical-Reasoning-SFT-GPT-OSS-120B
A high-quality synthetic dataset of medical reasoning conversations generated using OpenAI's gpt-oss-120B model with reasoning effort set to high, designed for supervised fine-tuning of large language models in healthcare applications. I used Intelligent-Internet/II-Medical-Reasoning-SFT as a seed dataset, so I would like to thank the authors and Intelligent-Internet for their great work.
Dataset Statistics
Total Samples: 200,927… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B.clt_gpt2_tokenized_control
Fresh multilingual GPT-2 CLT control data
Sequential, unshuffled control sample for CLT null experiments. For each language,
complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision
0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded
(including the complete document that crossed the threshold), after which complete
documents were retained until at least 100,000,000 tokens were collected.
Data are… See the full description on the dataset page: https://huggingface.co/datasets/abir-hr196/clt_gpt2_tokenized_control.swerl-tmax-15k-solvable-gpt-5-6-terra
swerl-tmax-15k hardened, post-validation-filter (dataset 3 of 3)
Which tasks in hamishivi/swerl-tmax-15k can a strong model actually solve? Every
task was attempted twice as a full agentic episode — real sandbox, real bash,
real verifier — and a task is verified when at least one attempt earned reward.
The last of three artifacts that exist to be compared by task_id:
original — hamishivi/swerl-tmax-15k, unchanged — 14,601 tasks
hardened, pre-validation-filter —… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-solvable-gpt-5-6-terra.GPT-4-PromptsMulti-Turn Conversational Prompts from ChatGPT-4 (10K+ Tokens)
Abstract:
This dataset offers a valuable collection of multi-turn conversational prompts generated by ChatGPT-4, carefully curated for diverse prompt styles (chatml, gemma, llama). Each prompt exceeds 10,000 tokens, providing ample context and inspiration for training and evaluating large language models. Ideal for researchers and developers interested in exploring advanced conversational AI capabilities.
Table of Contents:… See the full description on the dataset page: https://huggingface.co/datasets/erfanzar/GPT-4-Prompts.gpt_roleplay_realm
GPT Role-play Realm Dataset: The AI-generated character compendium
This is a dataset of GPT-generated characters made to increase the ability of open-source language models to role-play.
219 characters in the Russian part, and 216 characters in the English part. All character descriptions were generated with GPT-4.
20 dialogues on unique topics with every character. Topics were generated with GPT-4. The first dialogue out of 20 was also generated with GPT-4, and the other 19… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/gpt_roleplay_realm.TinyStories-gpt2-cache-100kCached activations at layer 5 for gpt2 using dataset apollo-research/roneneldan-TinyStories-tokenizer-gpt2
Useful for accelerated training and testing of sparse autoencoders
context_window: 512 tokens
total_tokens: 51,200,000
batch_size: 8 prompts (4096 tokens)
layer_hook_name: blocks.5.hook_mlp_out
LatamGPT-Corpus-1.0
LatamGPT-Corpus-1.0
🌐 Language versions: English | Español | Português
🔗 Project links: Official LatamGPT website | Corpus dashboard
🤖 Associated model: The complete LatamGPT corpus—of which this repository contains the openly released portion—was used in the training process of Llama-3.1-70B-LatamGPT-SFT-1.0.
Dataset description
Summary
LatamGPT-Corpus-1.0 is the open release of the data corpus assembled for the continued pretraining of… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/LatamGPT-Corpus-1.0.swerl-tmax-15k-rubric-gpt-5-6-sol
swerl-tmax-15k with task-quality rubric labels (gpt-5-6-sol)
hamishivi/swerl-tmax-15k, unchanged and unfiltered, with a per-task quality
label attached as extra columns.
This is not a verified or filtered dataset. Every one of the 14,601 original
records is present. Nothing has been dropped, repaired, or reordered. The labels
are one model's judgement about whether each task is sound enough to be useful RL
training data — an annotation layer, not a correctness guarantee.… See the full description on the dataset page: https://huggingface.co/datasets/wAI-org/swerl-tmax-15k-rubric-gpt-5-6-sol.clt_gpt2_tokenized_control
Fresh multilingual GPT-2 CLT control data
Sequential, unshuffled control sample for CLT null experiments. For each language,
complete source documents were tokenized with CausalNLP/gpt2-hf_multilingual-20 at revision
0afbb31b2db3f394270d42d6a4cb7f8fceeca3d8. The first 100,000,000 tokenizer tokens were discarded
(including the complete document that crossed the threshold), after which complete
documents were retained until at least 100,000,000 tokens were collected.
Data are… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/clt_gpt2_tokenized_control.Medical-Reasoning-SFT-GPT-OSS-120B-V2
Medical-Reasoning-SFT-GPT-OSS-120B-V2
A large-scale medical reasoning dataset generated using openai/gpt-oss-120b, containing over 506,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
GPT-OSS-120B is OpenAI's state-of-the-art open-weight model, achieving near-parity with closed models on reasoning benchmarks while being Apache 2.0 licensed.
Dataset Overview
Metric
Value
Model
openai/gpt-oss-120b
Total Samples
506… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-V2.Medical-Reasoning-SFT-GPT-OSS-120B-Small
Medical-Reasoning-SFT-GPT-OSS-120B-Small
A filtered and processed version of OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B optimized for training efficiency.
Dataset Description
This dataset contains high-quality medical reasoning conversations with the following modifications:
Length Filtering: Only includes samples where assistant responses are between 1000 and 10000 characters
Reasoning Extraction: Reasoning content from <think> tags has been extracted into a separate… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-Small.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 5.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-30m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 1.3000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-3m-historical-20t-think.harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think
1,000 historical evaluation attempts (250 tasks, four samples per task), newly
graded with gpt-5.6-sol using Harvey's original per-criterion rubric prompt
and all-criteria-pass rule. Mean all-pass rate: 4.0000%.
The train split contains evaluation records, not training examples.
Generation and grading protocols
Generation is unchanged: historical 20-turn thinking-enabled
glob/grep/read agent… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/harvey-eval-gpt56sol-qwen35-9b-notes70-recall30-10m-historical-20t-think.alpaca_gpt4_data_zhThis dataset clone from https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM
gpt2-training-ar-zh-ko-ja-4b
Balanced Arabic-Chinese-Korean-Japanese 4B-token training data
Sequential FineWeb-2 documents tokenized with CausalNLP/gpt2-tokenizer-ar-zh-ko-ja. Each language contains at least 1,000,000,000 tokens in complete documents. Total target: 4,000,000,000 tokens. Approximately 100,000,000 tokens per Parquet shard.
Splits: arb_Arab, cmn_Hani, kor_Hang, jpn_Jpan. Schema: text: string, input_ids: list<int32>.
gpt4all-j-prompt-generations-pt
Dataset Card for "gpt4all-j-prompt-generations-pt"
Dataset Description
Copy translated into Portuguese of the dataset gpt4all_prompt_generations using the googletrans library.
Translate
translate_dataset.ipynb
Usage
dataset_usage.ipynb
chinese-writing-bench-judgements-gpt-5.4
Zhiyin: Exploring the Frontier of Chinese LLM Writing
Website • GitHub • Hugging Face
Zhiyin is an LLM-as-a-judge benchmark for Chinese writing evaluation. This V1 release features 280 test cases across 18 diverse writing tasks.
Benchmark Overview
Our evaluation method relies on pairwise comparison. A powerful language model (O3) acts as the judge, scoring a model's response relative to a fixed baseline (GPT-4.1), which is anchored at a score of 5.
Scoring… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/chinese-writing-bench-judgements-gpt-5.4.smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming
A corpus of high quality fine tuning data meant for fine tuning various HelixLM models
Dataset Composition:
A subset sampled from randomly selected shards from https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus cosmopedia-v2 split ...
Added: A preprocessed column formattedconversation concatenating the prompt, text and special tokens for instruct fine tuning.
Added: A token count column tokencount based on GPT2 tokenizer + the fine tuning special tokens.… See the full description on the dataset page: https://huggingface.co/datasets/david-thrower/smollm-corpus-instruct-2M-cosmopedia-v2-gpt2-v2-streaming.gpt4all_prompt_generations
Dataset Card for [GPT4All Prompt Generations]
Dataset Description
Dataset used to train GPT4All
Homepage:
Repository: gpt4all
Paper: Technical Report
Atlas Map: Map of Cleaned Data
gpt-oss-120b-Infinity-Instruct-0625
gpt-oss-120b-Infinity-Instruct-0625
Dataset Description
This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside gpt-oss-120b as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with openai/gpt-oss-120b at temperature=1.
For more details on the training methodology and results, see our… See the full description on the dataset page: https://huggingface.co/datasets/nebius/gpt-oss-120b-Infinity-Instruct-0625.legal-contract-gpt41-redlining-10k
legal-contract-gpt41-redlining-10k
Dataset Description
This dataset contains 9977 synthetic legal contract redlines generated using GPT-4.1 model mix (base, mini, nano) with structured outputs. It is designed for fine-tuning LLMs (including OpenAI GPT-3.5/4, Llama, and other models) to assist with legal document redlining and clause revision.
Key Features
🤖 9977 synthetic redlines generated by GPT-4.1 model mix (base, mini, nano)
📋 Multiple training formats:… See the full description on the dataset page: https://huggingface.co/datasets/UmaiTech/legal-contract-gpt41-redlining-10k.Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High
Distill
This is a multi-source curated instruction and reasoning dataset specifically for training and distilling large language models (LLMs) to exhibit advanced Chain-of-Thought (CoT), Agentic, Mathematical and Coding capabilities. It aggregates high-quality outputs from frontier models into messages ChatML format.
Dataset Structure
The dataset contains a total of 70.2K examples, split into three subsets based on the presence of visible reasoning… See the full description on the dataset page: https://huggingface.co/datasets/VINAY-UMRETHE/Sonnet-Opus-4.5-4.6-Gemini-3.0-3.1-Pro-GPT-5-5.1-5.2-GLM-4.7-MiniMax-M2.1-DeepSeek-V3.2-High.
