CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Alibaba-Apsara /Superior-Reasoning-SFT-gpt-oss-120b-Logprob Superior-Reasoning-SFT-gpt-oss-120b-Logprob           🚀 Overview This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset. 🔗 Relationship to Main Dataset This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120bdataset. Records are linked via a unique sample_uuid. Main Dataset: Contains the text (prompts… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.texttext-generation100K<n<1M63 likes2.4k downloads8mo agoHugging Face02Alibaba-Apsara /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b           📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M352 likes1.4k downloads8mo agoHugging Face03erenyeager-1 /Superior-Reasoning-SFT-gpt-oss-120b-Logprob Superior-Reasoning-SFT-gpt-oss-120b-Logprob &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 🚀 Overview This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset. 🔗 Relationship to Main Dataset This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120b dataset. Records are linked via a unique sample_uuid. Main… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.texttext-generation100K<n<1M0 likes1.3k downloads2mo agoHugging Face04OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B Medical-Reasoning-SFT-GPT-OSS-120B A high-quality synthetic dataset of medical reasoning conversations generated using OpenAI's gpt-oss-120B model with reasoning effort set to high, designed for supervised fine-tuning of large language models in healthcare applications. I used Intelligent-Internet/II-Medical-Reasoning-SFT as a seed dataset, so I would like to thank the authors and Intelligent-Internet for their great work. Dataset Statistics Total Samples: 200,927… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B.texttext-generation100K<n<1M255 likes690 downloads10mo agoHugging Face05OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B-Small Medical-Reasoning-SFT-GPT-OSS-120B-Small A filtered and processed version of OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B optimized for training efficiency. Dataset Description This dataset contains high-quality medical reasoning conversations with the following modifications: Length Filtering: Only includes samples where assistant responses are between 1000 and 10000 characters Reasoning Extraction: Reasoning content from <think> tags has been extracted into a separate… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-Small.texttext-generation100K<n<1M3 likes268 downloads9mo agoHugging Face06Jackrong /gpt-oss-120b-reasoning-STEM-5K GPT-OSS-120B-Distilled-Reasoning-STEM Dataset 1) Dataset Overview Data Source Model: gpt-oss-120b-high Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics) Data Format: `JSON Lines Fields: generator, category, input, CoT_Native——reasoning, answer (Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.) 2) Design Goals (Motivation) This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.textquestion-answering1K<n<10K12 likes244 downloads1y agoHugging Face07OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B-V2 Medical-Reasoning-SFT-GPT-OSS-120B-V2 A large-scale medical reasoning dataset generated using openai/gpt-oss-120b, containing over 506,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. GPT-OSS-120B is OpenAI's state-of-the-art open-weight model, achieving near-parity with closed models on reasoning benchmarks while being Apache 2.0 licensed. Dataset Overview Metric Value Model openai/gpt-oss-120b Total Samples 506… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-V2.texttext-generation100K<n<1M9 likes206 downloads8mo agoHugging Face080xzanuee /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b           🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or heuristic filtering, Superior-Reasoning-SFT-gpt-oss-120b is constructed using a principled Distribution-Aligned Sequence… See the full description on the dataset page: https://huggingface.co/datasets/0xzanuee/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M0 likes111 downloads8mo agoHugging Face09nebius /gpt-oss-120b-Infinity-Instruct-0625 gpt-oss-120b-Infinity-Instruct-0625 Dataset Description This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside gpt-oss-120b as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with openai/gpt-oss-120b at temperature=1. For more details on the training methodology and results, see our… See the full description on the dataset page: https://huggingface.co/datasets/nebius/gpt-oss-120b-Infinity-Instruct-0625.texttext-generation100K<n<1M0 likes110 downloads7mo agoHugging Face10Jackrong /Natural-Reasoning-gpt-oss-120B-S1 Dataset Card: Natural-Reasoning-gpt-oss-120B-S1 📜 Dataset Overview This is a meticulously curated instruction fine-tuning dataset designed specifically for efficient knowledge distillation tasks. Built upon the first 100,000 questions from the large-scale reasoning corpus facebook/natural_reasoning (s1, I will process the remaining parts later), it aims to transfer the advanced, multi-step reasoning capabilities of the teacher model gpt-oss-120-high to a student model… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Natural-Reasoning-gpt-oss-120B-S1.texttext-generation10K<n<100K24 likes103 downloads1y agoHugging Face11Hugodonotexit /Superior-Reasoning-SFT-gpt-oss-120b-split-en Superior-Reasoning SFT (stage1 + stage2) with <think> split and English filtering Summary This dataset is a processed derivative of Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b (subsets stage1 and stage2, train split). It restructures each example into three fields: input: the original input reasoning: the content extracted from <think> ... </think> within the original output (inner text only) output: the remainder of the original output after removing all <think>… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/Superior-Reasoning-SFT-gpt-oss-120b-split-en.texttext-generation100K<n<1M2 likes103 downloads8mo agoHugging Face12AlgoDriveAI /TinyMathStories_gpt-oss-20b TinyMathStories A TinyStories-style corpus extended with math and lightweight reasoning. This dataset keeps the child-level vocabulary and short narrative style of TinyStories (Microsoft Research, Eldan & Li, 2023) and mixes in basic numeracy (counting, addition/subtraction, simple equations, fractions, measurement) and short justifications—so tiny models can practice coherent English and early math/logic. Research, generation, and curation by AlgoDriveAI.Inspired by and… See the full description on the dataset page: https://huggingface.co/datasets/AlgoDriveAI/TinyMathStories_gpt-oss-20b.texttext-generation100K<n<1M0 likes99 downloads9mo agoHugging Face13Ericwang /nemotron-nano2-safety-distill-gptoss Nemotron Nano 2 Safety Distill — GPT-OSS A distilled safety dataset produced using the Nemotron Nano 2 recipe with GPT-OSS-20B and GPT-OSS-120B as teacher models. ⚠️ Content Warning: This dataset includes potentially harmful prompts. Use responsibly for research purposes only. Overview This safety-focused distilled dataset was created by following the Nemotron Nano 2 safety recipe, adapted to use GPT-OSS-20B and GPT-OSS-120B as teacher models. Due to resource limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ericwang/nemotron-nano2-safety-distill-gptoss.texttext-generation10K<n<100K2 likes93 downloads11mo agoHugging Face14nebius /gpt-oss-20b-Infinity-Instruct-0625 gpt-oss-20b-Infinity-Instruct-0625 Dataset Description This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside gpt-oss-20b as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with openai/gpt-oss-20b at temperature=1. For more details on the training methodology and results, see our… See the full description on the dataset page: https://huggingface.co/datasets/nebius/gpt-oss-20b-Infinity-Instruct-0625.texttext-generation100K<n<1M0 likes87 downloads7mo agoHugging Face15Chaiyphop /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or heuristic filtering, Superior-Reasoning-SFT-gpt-oss-120b is constructed using a principled… See the full description on the dataset page: https://huggingface.co/datasets/Chaiyphop/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M1 likes86 downloads8mo agoHugging Face16anshy /Superior-Reasoning-SFT-gpt-oss-120b-random-shuffled Superior-Reasoning-SFT-gpt-oss-120b &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that… See the full description on the dataset page: https://huggingface.co/datasets/anshy/Superior-Reasoning-SFT-gpt-oss-120b-random-shuffled.texttext-generation100K<n<1M0 likes84 downloads4mo agoHugging Face17erenyeager-1 /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M0 likes82 downloads2mo agoHugging Face18Jackrong /gpt-oss-120B-distilled-reasoning GPT-oss-120B-Distilled-Reasoning-math Dataset Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines Fields: Generator, Category, Input, Output Core Statistics Generated complete reasoning processes and answers using gpt-oss-120b (MXFP4).The text length of the dataset reflects the depth and complexity of its content. I have statistically analyzed the lengths of the input (question), Reasoning, and Answer.To understand the data… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120B-distilled-reasoning.texttext-classification1K<n<10K20 likes80 downloads1y agoHugging Face19zanderjiang /gpt-oss-120b-SWE-Agent GPT-OSS 120B SWE-Agent Execution Traces Full execution traces of GPT-OSS 120B running SWE-Agent on SWE-Bench (Full). Dataset Structure Each JSON file in traces/ corresponds to one SWE-Bench problem instance. The filename is the instance ID (e.g., django__django-12345.json). Trace Schema { "instance_id": "django__django-12345", "model": "GPT-OSS-120B", "agent": "SWE-agent", "total_steps": 15, "total_run_duration_seconds": 120.5, "exit_status":… See the full description on the dataset page: https://huggingface.co/datasets/zanderjiang/gpt-oss-120b-SWE-Agent.text-generation1K<n<10K0 likes73 downloads6mo agoHugging Face20NarsAI /Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b           📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M0 likes68 downloads8mo agoHugging Face21Jackrong /GPT-OSS-20B-Distilled-Reasoning-Mini Dataset Card for Dataset Name GPT-OSS-20B Distilled Reasoning Dataset Mini (Multi-stage Evaluative Refinement Method for Reasoning Generation) Dataset Details and Description This is a high-quality instruction fine-tuning dataset constructed through knowledge distillation, featuring detailed Chain-of-Thought (CoT) reasoning processes. The dataset is designed to enhance the capabilities of smaller language models in complex reasoning, logical analysis, and instruction… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-20B-Distilled-Reasoning-Mini.tabulartext-classification1K<n<10K21 likes58 downloads1y agoHugging Face22Dogacel /nemotron-post-training-v2-gpt-oss-120b-regen Dataset Card for Nemotron Post Training v2 gpt-oss-120b Regen Regenerated responses from nvidia/Nemotron-Post-Training-Dataset-v2 dataset using gpt-oss-120b model. Parameter Value Max Tokens 8192 Temperature 1.0 Top-k 20 Top-p 0.95 Repetition Penalty 1.5 Dataset consists only the english samples from the Nemotron Post Training Dataset. 85% of the chat prompts have reasoning enabled, every other category has reasoning disabled. Category Value math… See the full description on the dataset page: https://huggingface.co/datasets/Dogacel/nemotron-post-training-v2-gpt-oss-120b-regen.texttext-generation100K<n<1M2 likes56 downloads5mo agoHugging Face23akahana /Medical-Reasoning-SFT-GPT-OSS-120B Medical-Reasoning-SFT-GPT-OSS-120B A high-quality synthetic dataset of medical reasoning conversations generated using OpenAI's gpt-oss-120B model with reasoning effort set to high, designed for supervised fine-tuning of large language models in healthcare applications. I used Intelligent-Internet/II-Medical-Reasoning-SFT as a seed dataset, so I would like to thank the authors and Intelligent-Internet for their great work. Dataset Statistics Total Samples: 200,927… See the full description on the dataset page: https://huggingface.co/datasets/akahana/Medical-Reasoning-SFT-GPT-OSS-120B.texttext-generation100K<n<1M0 likes54 downloads9mo agoHugging Face24deburky /gpt-oss-claude-code gpt-oss-claude-code SFT dataset Fine-tuning dataset for deburky/gpt-oss-claude-code, a tool-use and agentic coding model based on openai/gpt-oss-20b. Overview 284 training / 71 validation examples Format: gpt-oss harmony (<|start|>, <|channel|>, <|end|> tokens) Mix of knowledge Q&A, coding tasks, and multi-step tool-use conversations Tool-use examples include Read — file reading with offset/limit Grep — pattern search across files Glob — file discovery Bash… See the full description on the dataset page: https://huggingface.co/datasets/deburky/gpt-oss-claude-code.texttext-generationn<1K0 likes53 downloads6mo agoHugging Face25Jackrong /gpt-oss-120B-distilled-math-OpenAI-Harmony 📚 Dataset Overview Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines (.jsonl)Fields: Generator, Category, Input, Output Note: If you are using this template for training, please make sure the format is correct before starting.Since this template is still under continuous improvement and learning, it may not be fully complete yet. I appreciate your understanding. 📈 Core Statistics Generated complete reasoning processes… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120B-distilled-math-OpenAI-Harmony.texttext-classification1K<n<10K6 likes50 downloads1y agoHugging Face26AmanPriyanshu /GPT-OSS-20B-MoE-expert-activations GPT-OSS-20B MoE Expert Activations This dataset contains router activation patterns and expert selection data from OpenAI's GPT-OSS-20B mixture-of-experts model during text generation across diverse evaluation benchmarks. Dataset Description GPT-OSS-20B is OpenAI's open-weight mixture-of-experts language model with 21B total parameters and 3.6B active parameters per token. This dataset captures the internal routing decisions made by the model's router networks when… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/GPT-OSS-20B-MoE-expert-activations.feature-extraction10K<n<100K7 likes49 downloads1y agoHugging Face27wanglab /eurorad-gpt-oss-training-data Benchmarking and Adapting On-Device Large Language Models for Clinical Decision Support Authors Alif Munim* 1, Jun Ma* 1,2, Omar Ibrahim* 1, Alhusain Abdalla* 1, Shuolin Yin3, Leo Chen4, Bo Wang† 1,5,6,7,8 * Equal contribution     † Corresponding author 1AI Collaborative Centre, University Health Network, Toronto, Canada 2Princess Margaret Cancer Centre, University Health Network, Toronto, Canada 3Department of… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/eurorad-gpt-oss-training-data.texttext-generation1K<n<10K2 likes43 downloads7mo agoHugging Face28ahmetggg /gptoss20b-bilingual-curriculum-sft gpt-oss-20b Bilingual Curriculum SFT Synthetic bilingual supervised fine-tuning data generated with gpt-oss-20b (MoE, ~3.6B active params, native MXFP4, adaptive reasoning effort by difficulty). Domains: mathematics, physics, chemistry, biology, computer science, general science, general knowledge, conversation. Languages: Turkish and English. Difficulty levels: 1-8. The dataset is synthetic and should be independently evaluated before production use. texttext-generationn<1K0 likes43 downloads1mo agoHugging Face29iAmBoosted /gpt-oss-20b-reasoning-traces GPT-OSS-20B Reasoning Traces 3,333 reasoning traces generated by openai/gpt-oss-20b and filtered for clean, terminating reasoning. It was built to distill GPT-OSS's tight reasoning style into smaller models, and is the training set behind iAmBoosted/Qwen3.5-9B-OSS-Distilled. What's in it Each record pairs a prompt with GPT-OSS-20B's full reasoning trace and final answer, in chat-message form, ready for supervised fine-tuning (SFT). ~4,000 raw traces were generated, then… See the full description on the dataset page: https://huggingface.co/datasets/iAmBoosted/gpt-oss-20b-reasoning-traces.texttext-generation1K<n<10K0 likes42 downloads4mo agoHugging Face30prabinh /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b           🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or heuristic filtering, Superior-Reasoning-SFT-gpt-oss-120b is constructed using a principled Distribution-Aligned Sequence… See the full description on the dataset page: https://huggingface.co/datasets/prabinh/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M0 likes40 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.