datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-gpt4
Dataset Card for alpaca-gpt4
This dataset originates from this repository.
The alpaca-gpt4 dataset is specifically used for fine-tuning LLMs based on the instruction generated by GPT-4 using Alpaca prompts.
Dataset Details
Dataset Description
Each sample is comprised of four columns: instruction, input, output and text.
Language(s): English
License: Creative Commons NonCommercial (CC BY-NC 4.0)
Dataset Sources
The code from the original repository… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/alpaca-gpt4.alpaca-gpt4
Dataset Card for "alpaca-gpt4"
This dataset contains English Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs.
The dataset was originaly shared in this repository: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM. This is just a wraper for compatibility with huggingface's datasets library.
Dataset structure
It contains 52K instruction-following data generated by GPT-4 using the same prompts as in Alpaca.
The dataset has the… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/alpaca-gpt4.GPT-4o-evaluation-biases
A database to support the evaluation of gender biases in GPT-4o output
The database and its construction process are described in the paper "A database to support the evaluation of gender biases in GPT-4o output" by Mehner et al., presented at the 1st ISCA/ITG Workshop on Diversity in Large Speech and Language Models (Berlin, Februar 20, 2025).
Introduction
This is a database of prompts and answers generated with GPT-4o-mini and GPT-4o in a pretest and a main test… See the full description on the dataset page: https://huggingface.co/datasets/mtec-TUB/GPT-4o-evaluation-biases.gpt-5.6-sol-coding-and-debugging-traces
GPT-5.6 Sol Coding & Debugging Traces
Verified software-engineering, independent model-judging, seed-authoring,
defensive-security, and training-harness trajectories from
GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an
autonomous coding agent. Sessions show the observable development loop:
inspecting repositories, reproducing failures, explaining evidence, editing
files, running compilers and test suites, correcting mistakes, and verifying
the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/gpt-5.6-sol-coding-and-debugging-traces.alpaca_gpt4_zhBorrowed from: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM
Removed 6,103 mistruncated examples.
You can use it in LLaMA Factory by specifying dataset: alpaca_gpt4_zh.
alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.distill-gpt4-eng-chat
Description
Introducing dataset consisting of gpt4 answers to users requests. Queries were taken from allenai/WildChat-1M and causal-lm/instructions. Texts (requests and responses) were deleted in 3 cases:
either has non-english letters and special symbols
either has http-links
either has html blocks
either has perplexity more than 1.5*IQR + third quantile ( in some cases average perplexity value of sentences or maximum value was used )
GPT-4-PromptsMulti-Turn Conversational Prompts from ChatGPT-4 (10K+ Tokens)
Abstract:
This dataset offers a valuable collection of multi-turn conversational prompts generated by ChatGPT-4, carefully curated for diverse prompt styles (chatml, gemma, llama). Each prompt exceeds 10,000 tokens, providing ample context and inspiration for training and evaluating large language models. Ideal for researchers and developers interested in exploring advanced conversational AI capabilities.
Table of Contents:… See the full description on the dataset page: https://huggingface.co/datasets/erfanzar/GPT-4-Prompts.HuatuoGPT2-SFT-GPT4-140K
HuatuoGPT2-SFT-GPT4-140K
140K Chinese medical instructions generated by GPT-4, based on questions from HuatuoGPT Dataset.
This dataset contains supervised fine-tuning instructions for HuatuoGPT2, designed to enhance the model's ability to follow instructions in real medical scenarios. We have made all the data (142,248 entries) in this dataset publicly available.
Repository
Github: https://github.com/FreedomIntelligence/HuatuoGPT-II
Citation… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-SFT-GPT4-140K.lightblue-tagengo-gpt4
lightblue/tagengo-gpt4
An unofficial, reformatted version of lightblue/tagengo-gpt4.
Tagengo is described by its author as the world's largest high-quality multilingual chat dataset - containing over 75,000 single-turn conversations between humans and GPT‑4 (gpt-4-0125-preview) across 74 languages. It fills a major gap in multilingual chat data, which has so far been limited compared to English.
Additional Processing
Split by language
Kept only entries with both… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/lightblue-tagengo-gpt4.GPT-5.6-luna-reasoning-102881x
Teacher - GPT 5.6 Luna
A compact, multi-domain chat corpus for reasoning and tool use.
102,881 conversations · 1.27 GiB · $164 API generation cost+ $80 (codex)
reasoning · code · math · STEM · tools · security
Overview
This project is a high-quality conversational training dataset focused on reasoning, mathematics, STEM, Python programming, software security, and tool use. It was independently generated by the dataset creator through API-based… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/GPT-5.6-luna-reasoning-102881x.alpaca_gpt4_enBorrowed from: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM
You can use it in LLaMA Factory by specifying dataset: alpaca_gpt4_en.
MMInstruct-GPT4V
MMInstruct
The official implementation of the paper "MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity".
The data engine is available on GitHub at yuecao0119/MMInstruct.
Todo List
Data Engine.
Open Source Datasets.
Release the checkpoint.
Introduction
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations:
Instruction Annotation… See the full description on the dataset page: https://huggingface.co/datasets/yuecao0119/MMInstruct-GPT4V.Trueque-Benchmark-beta-0.1
🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture
🌐 Language versions: Español | Português
⚠️ Official Disclaimer: Beta Release (v0.1)
Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America.
Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.gpt-oss-120b-reasoning-STEM-5K
GPT-OSS-120B-Distilled-Reasoning-STEM Dataset
1) Dataset Overview
Data Source Model: gpt-oss-120b-high
Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics)
Data Format: `JSON Lines
Fields: generator, category, input, CoT_Native——reasoning, answer
(Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.)
2) Design Goals (Motivation)
This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.Medical-Reasoning-SFT-GPT-OSS-120B-V2
Medical-Reasoning-SFT-GPT-OSS-120B-V2
A large-scale medical reasoning dataset generated using openai/gpt-oss-120b, containing over 506,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
GPT-OSS-120B is OpenAI's state-of-the-art open-weight model, achieving near-parity with closed models on reasoning benchmarks while being Apache 2.0 licensed.
Dataset Overview
Metric
Value
Model
openai/gpt-oss-120b
Total Samples
506… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-V2.CHOCLO
🌽 CHOCLO: Latin American Cultural Knowledge Benchmark
Description
CHOCLO is a benchmark designed to evaluate cultural knowledge in language models, with a specific focus on entities representative of Latin America. Unlike traditional benchmarks, which often emphasize general knowledge or contexts dominated by English-language data, CHOCLO aims to capture the richness, diversity, and specificity of Latin American cultural knowledge, including traditions, gastronomy… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/CHOCLO.browsecomp-plus-md-toc-gpt5.4-nano
BrowseComp-Plus Structured 100k Corpus
This dataset is a drop-in, official-format variant of the BrowseComp-Plus 100k corpus used in our RISE Agent experiments. It keeps the same row count, document ids, URLs, and column names as the original BrowseComp-Plus corpus, but replaces each document's text field with a structured version that adds a generated table of contents and section headings.
Files
data.parquet: the corpus in the same three-column schema as the… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus-md-toc-gpt5.4-nano.legal-contract-gpt41-redlining-10k
legal-contract-gpt41-redlining-10k
Dataset Description
This dataset contains 9977 synthetic legal contract redlines generated using GPT-4.1 model mix (base, mini, nano) with structured outputs. It is designed for fine-tuning LLMs (including OpenAI GPT-3.5/4, Llama, and other models) to assist with legal document redlining and clause revision.
Key Features
🤖 9977 synthetic redlines generated by GPT-4.1 model mix (base, mini, nano)
📋 Multiple training formats:… See the full description on the dataset page: https://huggingface.co/datasets/UmaiTech/legal-contract-gpt41-redlining-10k.Vietnamese-alpaca-gpt4-gg-translatedSuperior-Reasoning-SFT-gpt-oss-120b-split-en
Superior-Reasoning SFT (stage1 + stage2) with <think> split and English filtering
Summary
This dataset is a processed derivative of Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b (subsets stage1 and stage2, train split). It restructures each example into three fields:
input: the original input
reasoning: the content extracted from <think> ... </think> within the original output (inner text only)
output: the remainder of the original output after removing all <think>… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/Superior-Reasoning-SFT-gpt-oss-120b-split-en.pokemon-gpt4o-captionsBorrowed from: https://huggingface.co/datasets/jugg1024/pokemon-gpt4o-captions
You can use it in LLaMA Factory by specifying dataset: pokemon_cap.
GPT5.6_SOL_INVESTIGACION
Dataset de Metodología Científica
Dataset en español para entrenamiento, validación y evaluación de modelos capaces de razonar sobre metodología de investigación científica. Incluye escenarios de distintas disciplinas y niveles de dificultad, con énfasis en diseño de estudios, inferencia causal, análisis cuantitativo y cualitativo, métodos mixtos, ética, medición, muestreo, interpretación de resultados y revisión crítica de protocolos.
1. Resumen… See the full description on the dataset page: https://huggingface.co/datasets/Januka2009/GPT5.6_SOL_INVESTIGACION.Korean_Wikipedia_Dataset_for_GPT2_August_2022
Dataset Card for korean_wikipedia_dataset_for_GPT2
Dataset Description
Entire Korean language Wikipedia data for GPT-2 training as of August 1st, 2022.
email: oscar.eaglewatch@gmail.com
Dataset Summary
This is to make a pre-trained GPT-2 Korean model
Languages
Korean
Dataset Structure
Data Instances
Train wikipedia article count: 334420
validation wikipedia article count: 83605
Data Fields
'text'
Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/eaglewatch/Korean_Wikipedia_Dataset_for_GPT2_August_2022.nemotron-nano2-safety-distill-gptoss
Nemotron Nano 2 Safety Distill — GPT-OSS
A distilled safety dataset produced using the Nemotron Nano 2 recipe with GPT-OSS-20B and GPT-OSS-120B as teacher models.
⚠️ Content Warning: This dataset includes potentially harmful prompts. Use responsibly for research purposes only.
Overview
This safety-focused distilled dataset was created by following the Nemotron Nano 2 safety recipe, adapted to use GPT-OSS-20B and GPT-OSS-120B as teacher models. Due to resource limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ericwang/nemotron-nano2-safety-distill-gptoss.CodeToolSmith-GPT55-Traces
CodeToolSmith GPT-5.5 Traces
This dataset contains GPT-5.5 generated multi-turn traces where the assistant first writes a task-specific Python tool, validates it with tests, and then uses the tool to answer a self-contained coding/data task.
It is meant for coding-model SFT and RL experiments on tool creation behavior. It is not a hidden benchmark, not a leaderboard split, and should not be reported as an evaluation result.
What Is Included
5720 validated… See the full description on the dataset page: https://huggingface.co/datasets/jwu323/CodeToolSmith-GPT55-Traces.gpt-5.6-sol-coding-and-debugging-traces
GPT-5.6 Sol Coding & Debugging Traces
Verified software-engineering, independent model-judging, seed-authoring,
defensive-security, and training-harness trajectories from
GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an
autonomous coding agent. Sessions show the observable development loop:
inspecting repositories, reproducing failures, explaining evidence, editing
files, running compilers and test suites, correcting mistakes, and verifying
the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/gpt-5.6-sol-coding-and-debugging-traces.gpt-oss-120B-distilled-reasoning
GPT-oss-120B-Distilled-Reasoning-math Dataset
Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines
Fields: Generator, Category, Input, Output
Core Statistics
Generated complete reasoning processes and answers using gpt-oss-120b (MXFP4).The text length of the dataset reflects the depth and complexity of its content. I have statistically analyzed the lengths of the input (question), Reasoning, and Answer.To understand the data… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120B-distilled-reasoning.GPT-OSS-120B-Distilled-Reasoning-math
GPT-oss-120B-Distilled-Reasoning-math Dataset
Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines
Fields: Generator, Category, Input, CoT_Native_Reasoning, Reasoning, Answer
Core Statistics
Generated complete reasoning processes and answers using gpt-oss-120b (MXFP4).The text length of the dataset reflects the depth and complexity of its content. I have statistically analyzed the lengths of the input (question), Reasoning, and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-120B-Distilled-Reasoning-math.MuSeR_GPT_OSS_120B_DistillationThis dataset contains ~100k synthetic medical queries and corresponding responses distilled from GPT-OSS-120B.
The generation of synthetic medical queries follows an attribute-conditioned generation method proposed in paper Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning.
We found that supervised fine-tuning on this dataset can substantially improve LLMs' medical conversational capabilities. See our paper and project page for more details.
If… See the full description on the dataset page: https://huggingface.co/datasets/zyx1234/MuSeR_GPT_OSS_120B_Distillation.
