CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01flwrlabs /alpaca-gpt4 Dataset Card for alpaca-gpt4 This dataset originates from this repository. The alpaca-gpt4 dataset is specifically used for fine-tuning LLMs based on the instruction generated by GPT-4 using Alpaca prompts. Dataset Details Dataset Description Each sample is comprised of four columns: instruction, input, output and text. Language(s): English License: Creative Commons NonCommercial (CC BY-NC 4.0) Dataset Sources The code from the original repository… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/alpaca-gpt4.texttext-generation10K<n<100K0 likes18k downloads1y agoHugging Face02vicgalle /alpaca-gpt4 Dataset Card for "alpaca-gpt4" This dataset contains English Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs. The dataset was originaly shared in this repository: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM. This is just a wraper for compatibility with huggingface's datasets library. Dataset structure It contains 52K instruction-following data generated by GPT-4 using the same prompts as in Alpaca. The dataset has the… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/alpaca-gpt4.texttext-generation10K<n<100K326 likes4.4k downloads3y agoHugging Face03mtec-TUB /GPT-4o-evaluation-biases A database to support the evaluation of gender biases in GPT-4o output The database and its construction process are described in the paper "A database to support the evaluation of gender biases in GPT-4o output" by Mehner et al., presented at the 1st ISCA/ITG Workshop on Diversity in Large Speech and Language Models (Berlin, Februar 20, 2025). Introduction This is a database of prompts and answers generated with GPT-4o-mini and GPT-4o in a pretest and a main test… See the full description on the dataset page: https://huggingface.co/datasets/mtec-TUB/GPT-4o-evaluation-biases.question-answering10K<n<100K0 likes1.7k downloads2y agoHugging Face04greghavens /gpt-5.6-sol-coding-and-debugging-traces GPT-5.6 Sol Coding & Debugging Traces Verified software-engineering, independent model-judging, seed-authoring, defensive-security, and training-harness trajectories from GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an autonomous coding agent. Sessions show the observable development loop: inspecting repositories, reproducing failures, explaining evidence, editing files, running compilers and test suites, correcting mistakes, and verifying the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/gpt-5.6-sol-coding-and-debugging-traces.text-generation10K<n<100K44 likes1.1k downloads2mo agoHugging Face05llamafactory /alpaca_gpt4_zhBorrowed from: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM Removed 6,103 mistruncated examples. You can use it in LLaMA Factory by specifying dataset: alpaca_gpt4_zh. texttext-generation10K<n<100K20 likes931 downloads2y agoHugging Face06Ichsan2895 /alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian We wrangled the original dataset format to 'input' & 'output' format. For example: BEFORE: [ { "from": "human", "value": "Saranlah slogan untuk kampanye daur ulang\n" }, { "from": "gpt", "value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \ "Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \ "Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.textquestion-answering10K<n<100K23 likes874 downloads3y agoHugging Face07mtimur /distill-gpt4-eng-chat Description Introducing dataset consisting of gpt4 answers to users requests. Queries were taken from allenai/WildChat-1M and causal-lm/instructions. Texts (requests and responses) were deleted in 3 cases: either has non-english letters and special symbols either has http-links either has html blocks either has perplexity more than 1.5*IQR + third quantile ( in some cases average perplexity value of sentences or maximum value was used ) textquestion-answering100K<n<1M2 likes501 downloads2y agoHugging Face08erfanzar /GPT-4-PromptsMulti-Turn Conversational Prompts from ChatGPT-4 (10K+ Tokens) Abstract: This dataset offers a valuable collection of multi-turn conversational prompts generated by ChatGPT-4, carefully curated for diverse prompt styles (chatml, gemma, llama). Each prompt exceeds 10,000 tokens, providing ample context and inspiration for training and evaluating large language models. Ideal for researchers and developers interested in exploring advanced conversational AI capabilities. Table of Contents:… See the full description on the dataset page: https://huggingface.co/datasets/erfanzar/GPT-4-Prompts.texttranslation10K<n<100K14 likes486 downloads3y agoHugging Face09FreedomIntelligence /HuatuoGPT2-SFT-GPT4-140K HuatuoGPT2-SFT-GPT4-140K 140K Chinese medical instructions generated by GPT-4, based on questions from HuatuoGPT Dataset. This dataset contains supervised fine-tuning instructions for HuatuoGPT2, designed to enhance the model's ability to follow instructions in real medical scenarios. We have made all the data (142,248 entries) in this dataset publicly available. Repository Github: https://github.com/FreedomIntelligence/HuatuoGPT-II Citation… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-SFT-GPT4-140K.question-answering14 likes415 downloads2y agoHugging Face10agentlans /lightblue-tagengo-gpt4 lightblue/tagengo-gpt4 An unofficial, reformatted version of lightblue/tagengo-gpt4. Tagengo is described by its author as the world's largest high-quality multilingual chat dataset - containing over 75,000 single-turn conversations between humans and GPT‑4 (gpt-4-0125-preview) across 74 languages. It fills a major gap in multilingual chat data, which has so far been limited compared to English. Additional Processing Split by language Kept only entries with both… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/lightblue-tagengo-gpt4.texttext-generation10K<n<100K0 likes392 downloads8mo agoHugging Face11Roman1111111 /GPT-5.6-luna-reasoning-102881x Teacher - GPT 5.6 Luna A compact, multi-domain chat corpus for reasoning and tool use. 102,881 conversations · 1.27 GiB · $164 API generation cost+ $80 (codex) reasoning · code · math · STEM · tools · security Overview This project is a high-quality conversational training dataset focused on reasoning, mathematics, STEM, Python programming, software security, and tool use. It was independently generated by the dataset creator through API-based… See the full description on the dataset page: https://huggingface.co/datasets/Roman1111111/GPT-5.6-luna-reasoning-102881x.text-generation100K<n<1M10 likes376 downloads15d agoHugging Face12llamafactory /alpaca_gpt4_enBorrowed from: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM You can use it in LLaMA Factory by specifying dataset: alpaca_gpt4_en. texttext-generation10K<n<100K4 likes334 downloads2y agoHugging Face13yuecao0119 /MMInstruct-GPT4V MMInstruct The official implementation of the paper "MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity". The data engine is available on GitHub at yuecao0119/MMInstruct. Todo List Data Engine. Open Source Datasets. Release the checkpoint. Introduction Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations: Instruction Annotation… See the full description on the dataset page: https://huggingface.co/datasets/yuecao0119/MMInstruct-GPT4V.imagevisual-question-answering100K<n<1M13 likes333 downloads2y agoHugging Face14latam-gpt /Trueque-Benchmark-beta-0.1 🤝 Trueque: A human-reviewed collaborative benchmark for Latin American knowledge and culture 🌐 Language versions: Español | Português ⚠️ Official Disclaimer: Beta Release (v0.1) Welcome to Trueque for Factual Knowledge and Cultural Appropriateness. This dataset represents an initial effort to evaluate the regional knowledge and cultural accuracy of Large Language Models (LLMs) in Latin America. Please take the following considerations into account before using this resource:… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/Trueque-Benchmark-beta-0.1.textquestion-answeringn<1K8 likes270 downloads2mo agoHugging Face15Jackrong /gpt-oss-120b-reasoning-STEM-5K GPT-OSS-120B-Distilled-Reasoning-STEM Dataset 1) Dataset Overview Data Source Model: gpt-oss-120b-high Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics) Data Format: `JSON Lines Fields: generator, category, input, CoT_Native——reasoning, answer (Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.) 2) Design Goals (Motivation) This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.textquestion-answering1K<n<10K12 likes249 downloads1y agoHugging Face16OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B-V2 Medical-Reasoning-SFT-GPT-OSS-120B-V2 A large-scale medical reasoning dataset generated using openai/gpt-oss-120b, containing over 506,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. GPT-OSS-120B is OpenAI's state-of-the-art open-weight model, achieving near-parity with closed models on reasoning benchmarks while being Apache 2.0 licensed. Dataset Overview Metric Value Model openai/gpt-oss-120b Total Samples 506… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-V2.texttext-generation100K<n<1M9 likes213 downloads8mo agoHugging Face17latam-gpt /CHOCLO 🌽 CHOCLO: Latin American Cultural Knowledge Benchmark Description CHOCLO is a benchmark designed to evaluate cultural knowledge in language models, with a specific focus on entities representative of Latin America. Unlike traditional benchmarks, which often emphasize general knowledge or contexts dominated by English-language data, CHOCLO aims to capture the richness, diversity, and specificity of Latin American cultural knowledge, including traditions, gastronomy… See the full description on the dataset page: https://huggingface.co/datasets/latam-gpt/CHOCLO.textquestion-answering100K<n<1M14 likes189 downloads6mo agoHugging Face18Tevatron /browsecomp-plus-md-toc-gpt5.4-nano BrowseComp-Plus Structured 100k Corpus This dataset is a drop-in, official-format variant of the BrowseComp-Plus 100k corpus used in our RISE Agent experiments. It keeps the same row count, document ids, URLs, and column names as the original BrowseComp-Plus corpus, but replaces each document's text field with a structured version that adds a generated table of contents and section headings. Files data.parquet: the corpus in the same three-column schema as the… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus-md-toc-gpt5.4-nano.textquestion-answering100K<n<1M0 likes133 downloads3mo agoHugging Face19UmaiTech /legal-contract-gpt41-redlining-10k legal-contract-gpt41-redlining-10k Dataset Description This dataset contains 9977 synthetic legal contract redlines generated using GPT-4.1 model mix (base, mini, nano) with structured outputs. It is designed for fine-tuning LLMs (including OpenAI GPT-3.5/4, Llama, and other models) to assist with legal document redlining and clause revision. Key Features 🤖 9977 synthetic redlines generated by GPT-4.1 model mix (base, mini, nano) 📋 Multiple training formats:… See the full description on the dataset page: https://huggingface.co/datasets/UmaiTech/legal-contract-gpt41-redlining-10k.texttext-generation10K<n<100K1 likes132 downloads11mo agoHugging Face205CD-AI /Vietnamese-alpaca-gpt4-gg-translatedtextquestion-answering10K<n<100K20 likes128 downloads3y agoHugging Face21Hugodonotexit /Superior-Reasoning-SFT-gpt-oss-120b-split-en Superior-Reasoning SFT (stage1 + stage2) with <think> split and English filtering Summary This dataset is a processed derivative of Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b (subsets stage1 and stage2, train split). It restructures each example into three fields: input: the original input reasoning: the content extracted from <think> ... </think> within the original output (inner text only) output: the remainder of the original output after removing all <think>… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/Superior-Reasoning-SFT-gpt-oss-120b-split-en.texttext-generation100K<n<1M2 likes114 downloads8mo agoHugging Face22llamafactory /pokemon-gpt4o-captionsBorrowed from: https://huggingface.co/datasets/jugg1024/pokemon-gpt4o-captions You can use it in LLaMA Factory by specifying dataset: pokemon_cap. imagetext-generation1K<n<10K8 likes113 downloads2y agoHugging Face23Januka2009 /GPT5.6_SOL_INVESTIGACION Dataset de Metodología Científica Dataset en español para entrenamiento, validación y evaluación de modelos capaces de razonar sobre metodología de investigación científica. Incluye escenarios de distintas disciplinas y niveles de dificultad, con énfasis en diseño de estudios, inferencia causal, análisis cuantitativo y cualitativo, métodos mixtos, ética, medición, muestreo, interpretación de resultados y revisión crítica de protocolos. 1. Resumen… See the full description on the dataset page: https://huggingface.co/datasets/Januka2009/GPT5.6_SOL_INVESTIGACION.texttext-generation1K<n<10K1 likes112 downloads3d agoHugging Face24eaglewatch /Korean_Wikipedia_Dataset_for_GPT2_August_2022 Dataset Card for korean_wikipedia_dataset_for_GPT2 Dataset Description Entire Korean language Wikipedia data for GPT-2 training as of August 1st, 2022. email: oscar.eaglewatch@gmail.com Dataset Summary This is to make a pre-trained GPT-2 Korean model Languages Korean Dataset Structure Data Instances Train wikipedia article count: 334420 validation wikipedia article count: 83605 Data Fields 'text' Data Splits… See the full description on the dataset page: https://huggingface.co/datasets/eaglewatch/Korean_Wikipedia_Dataset_for_GPT2_August_2022.textquestion-answering100K<n<1M6 likes95 downloads2y agoHugging Face25Ericwang /nemotron-nano2-safety-distill-gptoss Nemotron Nano 2 Safety Distill — GPT-OSS A distilled safety dataset produced using the Nemotron Nano 2 recipe with GPT-OSS-20B and GPT-OSS-120B as teacher models. ⚠️ Content Warning: This dataset includes potentially harmful prompts. Use responsibly for research purposes only. Overview This safety-focused distilled dataset was created by following the Nemotron Nano 2 safety recipe, adapted to use GPT-OSS-20B and GPT-OSS-120B as teacher models. Due to resource limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ericwang/nemotron-nano2-safety-distill-gptoss.texttext-generation10K<n<100K2 likes95 downloads11mo agoHugging Face26jwu323 /CodeToolSmith-GPT55-Traces CodeToolSmith GPT-5.5 Traces This dataset contains GPT-5.5 generated multi-turn traces where the assistant first writes a task-specific Python tool, validates it with tests, and then uses the tool to answer a self-contained coding/data task. It is meant for coding-model SFT and RL experiments on tool creation behavior. It is not a hidden benchmark, not a leaderboard split, and should not be reported as an evaluation result. What Is Included 5720 validated… See the full description on the dataset page: https://huggingface.co/datasets/jwu323/CodeToolSmith-GPT55-Traces.text-generation0 likes94 downloads3mo agoHugging Face27Nobody05 /gpt-5.6-sol-coding-and-debugging-traces GPT-5.6 Sol Coding & Debugging Traces Verified software-engineering, independent model-judging, seed-authoring, defensive-security, and training-harness trajectories from GPT-5.6 Sol (gpt-5.6-sol) running through the Codex CLI as an autonomous coding agent. Sessions show the observable development loop: inspecting repositories, reproducing failures, explaining evidence, editing files, running compilers and test suites, correcting mistakes, and verifying the completed result.… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/gpt-5.6-sol-coding-and-debugging-traces.text-generation10K<n<100K1 likes90 downloads2mo agoHugging Face28Jackrong /gpt-oss-120B-distilled-reasoning GPT-oss-120B-Distilled-Reasoning-math Dataset Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines Fields: Generator, Category, Input, Output Core Statistics Generated complete reasoning processes and answers using gpt-oss-120b (MXFP4).The text length of the dataset reflects the depth and complexity of its content. I have statistically analyzed the lengths of the input (question), Reasoning, and Answer.To understand the data… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120B-distilled-reasoning.texttext-classification1K<n<10K20 likes84 downloads1y agoHugging Face29Jackrong /GPT-OSS-120B-Distilled-Reasoning-math GPT-oss-120B-Distilled-Reasoning-math Dataset Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines Fields: Generator, Category, Input, CoT_Native_Reasoning, Reasoning, Answer Core Statistics Generated complete reasoning processes and answers using gpt-oss-120b (MXFP4).The text length of the dataset reflects the depth and complexity of its content. I have statistically analyzed the lengths of the input (question), Reasoning, and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-120B-Distilled-Reasoning-math.textquestion-answering1K<n<10K9 likes78 downloads1y agoHugging Face30zyx1234 /MuSeR_GPT_OSS_120B_DistillationThis dataset contains ~100k synthetic medical queries and corresponding responses distilled from GPT-OSS-120B. The generation of synthetic medical queries follows an attribute-conditioned generation method proposed in paper Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning. We found that supervised fine-tuning on this dataset can substantially improve LLMs' medical conversational capabilities. See our paper and project page for more details. If… See the full description on the dataset page: https://huggingface.co/datasets/zyx1234/MuSeR_GPT_OSS_120B_Distillation.textquestion-answering10K<n<100K4 likes75 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.