CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ise-uiuc /Magicoder-OSS-Instruct-75KThis is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy when adopting this dataset: https://openai.com/policies/usage-policies. tabulartext-generation10K<n<100K170 likes63k downloads3y agoHugging Face02Alibaba-Apsara /Superior-Reasoning-SFT-gpt-oss-120b-Logprob Superior-Reasoning-SFT-gpt-oss-120b-Logprob           🚀 Overview This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset. 🔗 Relationship to Main Dataset This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120bdataset. Records are linked via a unique sample_uuid. Main Dataset: Contains the text (prompts… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.texttext-generation100K<n<1M63 likes2.4k downloads8mo agoHugging Face03Alibaba-Apsara /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b           📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M352 likes1.4k downloads8mo agoHugging Face04erenyeager-1 /Superior-Reasoning-SFT-gpt-oss-120b-Logprob Superior-Reasoning-SFT-gpt-oss-120b-Logprob &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 🚀 Overview This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset. 🔗 Relationship to Main Dataset This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120b dataset. Records are linked via a unique sample_uuid. Main… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.texttext-generation100K<n<1M0 likes1.3k downloads2mo agoHugging Face05csoai /gspc-oss GSPC — openness bank (OSSBench) Council of AI measurement bank. Measurement, not certification. Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026. Live measurement. This bank stands behind the openness row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=openness (family, kind, status and n are on that row, never typed here;… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-oss.tabularquestion-answeringn<1K0 likes700 downloads2d agoHugging Face06OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B Medical-Reasoning-SFT-GPT-OSS-120B A high-quality synthetic dataset of medical reasoning conversations generated using OpenAI's gpt-oss-120B model with reasoning effort set to high, designed for supervised fine-tuning of large language models in healthcare applications. I used Intelligent-Internet/II-Medical-Reasoning-SFT as a seed dataset, so I would like to thank the authors and Intelligent-Internet for their great work. Dataset Statistics Total Samples: 200,927… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B.texttext-generation100K<n<1M255 likes690 downloads10mo agoHugging Face07jablonkagroup /corral-oss-trace-logprobs Corral – OSS-120B Trace Logprobs Token-level log-probabilities for GPT-Oss-120B evaluation runs across all 8 Corral environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the token-level log-probabilities recorded during the evaluation runs of GPT-Oss-120B across all 8 Corral environments. Each configuration (config) of this dataset… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-oss-trace-logprobs.tabulartext-generation100K<n<1M0 likes535 downloads3mo agoHugging Face08OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B-Small Medical-Reasoning-SFT-GPT-OSS-120B-Small A filtered and processed version of OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B optimized for training efficiency. Dataset Description This dataset contains high-quality medical reasoning conversations with the following modifications: Length Filtering: Only includes samples where assistant responses are between 1000 and 10000 characters Reasoning Extraction: Reasoning content from <think> tags has been extracted into a separate… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-Small.texttext-generation100K<n<1M3 likes268 downloads9mo agoHugging Face09Jackrong /gpt-oss-120b-reasoning-STEM-5K GPT-OSS-120B-Distilled-Reasoning-STEM Dataset 1) Dataset Overview Data Source Model: gpt-oss-120b-high Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics) Data Format: `JSON Lines Fields: generator, category, input, CoT_Native——reasoning, answer (Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.) 2) Design Goals (Motivation) This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.textquestion-answering1K<n<10K12 likes244 downloads1y agoHugging Face10OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B-V2 Medical-Reasoning-SFT-GPT-OSS-120B-V2 A large-scale medical reasoning dataset generated using openai/gpt-oss-120b, containing over 506,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. GPT-OSS-120B is OpenAI's state-of-the-art open-weight model, achieving near-parity with closed models on reasoning benchmarks while being Apache 2.0 licensed. Dataset Overview Metric Value Model openai/gpt-oss-120b Total Samples 506… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-V2.texttext-generation100K<n<1M9 likes206 downloads8mo agoHugging Face11dicta-il /MathCOT-oss-vs-DeepSeek Learning to Reason: Training LLMs with GPT-OSS or DeepSeek R1 Reasoning Traces This dataset is the one used from the paper, available here 📄 This dataset consists of 242k math questions, with the verified generated answer (with reasoning) by both DeepSeek-R1-0528 and gpt-oss-120b. The original prompts and the DeepSeek-R1-0528 traces were taken from NVIDIA's Nemotron-Post-Training-Dataset-v1. Citation If you found this dataset useful, please cite the paper below:… See the full description on the dataset page: https://huggingface.co/datasets/dicta-il/MathCOT-oss-vs-DeepSeek.texttext-generation100K<n<1M2 likes197 downloads10mo agoHugging Face12Tonic /Health-Bench-Eval-OSS-2025-07 Dataset Card for HealthBench Dataset Summary HealthBench is a benchmark dataset developed by OpenAI in collaboration with 262 physicians from 60 countries to evaluate AI systems in health-related conversational scenarios. It contains 5,000 multi-turn health conversations in a JSONL file (2025-05-07-06-14-12_oss_eval.jsonl), simulating interactions between AI models and users (laypersons or clinicians). Each conversation includes a user prompt, a candidate model response… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/Health-Bench-Eval-OSS-2025-07.texttext-generation1K<n<10K4 likes151 downloads1y agoHugging Face130xzanuee /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b           🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or heuristic filtering, Superior-Reasoning-SFT-gpt-oss-120b is constructed using a principled Distribution-Aligned Sequence… See the full description on the dataset page: https://huggingface.co/datasets/0xzanuee/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M0 likes111 downloads8mo agoHugging Face14nebius /gpt-oss-120b-Infinity-Instruct-0625 gpt-oss-120b-Infinity-Instruct-0625 Dataset Description This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside gpt-oss-120b as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with openai/gpt-oss-120b at temperature=1. For more details on the training methodology and results, see our… See the full description on the dataset page: https://huggingface.co/datasets/nebius/gpt-oss-120b-Infinity-Instruct-0625.texttext-generation100K<n<1M0 likes110 downloads7mo agoHugging Face15Jackrong /Natural-Reasoning-gpt-oss-120B-S1 Dataset Card: Natural-Reasoning-gpt-oss-120B-S1 📜 Dataset Overview This is a meticulously curated instruction fine-tuning dataset designed specifically for efficient knowledge distillation tasks. Built upon the first 100,000 questions from the large-scale reasoning corpus facebook/natural_reasoning (s1, I will process the remaining parts later), it aims to transfer the advanced, multi-step reasoning capabilities of the teacher model gpt-oss-120-high to a student model… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Natural-Reasoning-gpt-oss-120B-S1.texttext-generation10K<n<100K24 likes103 downloads1y agoHugging Face16Hugodonotexit /Superior-Reasoning-SFT-gpt-oss-120b-split-en Superior-Reasoning SFT (stage1 + stage2) with <think> split and English filtering Summary This dataset is a processed derivative of Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b (subsets stage1 and stage2, train split). It restructures each example into three fields: input: the original input reasoning: the content extracted from <think> ... </think> within the original output (inner text only) output: the remainder of the original output after removing all <think>… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/Superior-Reasoning-SFT-gpt-oss-120b-split-en.texttext-generation100K<n<1M2 likes103 downloads8mo agoHugging Face17AlgoDriveAI /TinyMathStories_gpt-oss-20b TinyMathStories A TinyStories-style corpus extended with math and lightweight reasoning. This dataset keeps the child-level vocabulary and short narrative style of TinyStories (Microsoft Research, Eldan & Li, 2023) and mixes in basic numeracy (counting, addition/subtraction, simple equations, fractions, measurement) and short justifications—so tiny models can practice coherent English and early math/logic. Research, generation, and curation by AlgoDriveAI.Inspired by and… See the full description on the dataset page: https://huggingface.co/datasets/AlgoDriveAI/TinyMathStories_gpt-oss-20b.texttext-generation100K<n<1M0 likes99 downloads9mo agoHugging Face18nebius /gpt-oss-20b-Infinity-Instruct-0625 gpt-oss-20b-Infinity-Instruct-0625 Dataset Description This dataset is part of the LK-Speculators collection for speculative decoding research. It contains 660K prompt-response pairs designed for training draft models that are used alongside gpt-oss-20b as the target model. The dataset was created by generating responses to the prompts from Infinity-Instruct-0625 with openai/gpt-oss-20b at temperature=1. For more details on the training methodology and results, see our… See the full description on the dataset page: https://huggingface.co/datasets/nebius/gpt-oss-20b-Infinity-Instruct-0625.texttext-generation100K<n<1M0 likes87 downloads7mo agoHugging Face19Chaiyphop /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or heuristic filtering, Superior-Reasoning-SFT-gpt-oss-120b is constructed using a principled… See the full description on the dataset page: https://huggingface.co/datasets/Chaiyphop/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M1 likes86 downloads8mo agoHugging Face20anshy /Superior-Reasoning-SFT-gpt-oss-120b-random-shuffled Superior-Reasoning-SFT-gpt-oss-120b &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that… See the full description on the dataset page: https://huggingface.co/datasets/anshy/Superior-Reasoning-SFT-gpt-oss-120b-random-shuffled.texttext-generation100K<n<1M0 likes84 downloads4mo agoHugging Face21erenyeager-1 /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M0 likes82 downloads2mo agoHugging Face22Jackrong /gpt-oss-120B-distilled-reasoning GPT-oss-120B-Distilled-Reasoning-math Dataset Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines Fields: Generator, Category, Input, Output Core Statistics Generated complete reasoning processes and answers using gpt-oss-120b (MXFP4).The text length of the dataset reflects the depth and complexity of its content. I have statistically analyzed the lengths of the input (question), Reasoning, and Answer.To understand the data… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120B-distilled-reasoning.texttext-classification1K<n<10K20 likes80 downloads1y agoHugging Face23WrittenWithRust /Magicoder-OSS-Instruct-Rust-cleaned-3.9K 🦀 Magicoder-OSS-Instruct-Rust (3.9K Cleaned) Magicoder-OSS-Instruct-Rust is a high-quality, syntax-verified dataset of 3,909 Rust coding instructions derived from real-world open-source GitHub projects. This dataset is extracted from ise-uiuc/Magicoder-OSS-Instruct-75K, filtered specifically for Rust, and validated via in-memory compiler checks. No language translation was applied; the dataset remains in its original English format. ⚙️ Filtering and Verification… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Magicoder-OSS-Instruct-Rust-cleaned-3.9K.texttext-generation1K<n<10K0 likes77 downloads25d agoHugging Face24NarsAI /Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b           📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M0 likes68 downloads8mo agoHugging Face25WrittenWithRust /Magicoder-OSS-Instruct-Rust-TR-3.9K 🦀 Magicoder-OSS-Instruct-Rust-Turkish (3.9K) Magicoder-OSS-Instruct-Rust-Turkish, WrittenWithRust/Magicoder-OSS-Instruct-Rust-cleaned-3.9K veri setindeki 3.909 adet sentaksı doğrulanmış İngilizce Rust instruction örneğinin tamamen Türkçe diline çevrilmesiyle oluşturulmuş yüksek kaliteli bir kod veri setidir. Bu veri seti, Büyük Dil Modellerine (LLM) Türkçe Rust kodlama becerisi, problem çözme yeteneği ve karmaşık mimarileri açıklama kabiliyeti kazandırmak üzere Instruction… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/Magicoder-OSS-Instruct-Rust-TR-3.9K.texttext-generation1K<n<10K0 likes64 downloads24d agoHugging Face26dakies /OSS_VerilogOriginal dataset size: 21725 Number of duplicate clusters: 2951 Files in duplicate cluster: 8054 Unique files in duplicate cluster: 3781 Filtered dataset size: 17452 Time to deduplicate dataset: 7.37 Size of deduplicated dataset: 17452, old dataset size 21725 texttext-generation10K<n<100K0 likes61 downloads2y agoHugging Face27Jackrong /GPT-OSS-20B-Distilled-Reasoning-Mini Dataset Card for Dataset Name GPT-OSS-20B Distilled Reasoning Dataset Mini (Multi-stage Evaluative Refinement Method for Reasoning Generation) Dataset Details and Description This is a high-quality instruction fine-tuning dataset constructed through knowledge distillation, featuring detailed Chain-of-Thought (CoT) reasoning processes. The dataset is designed to enhance the capabilities of smaller language models in complex reasoning, logical analysis, and instruction… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-20B-Distilled-Reasoning-Mini.tabulartext-classification1K<n<10K21 likes58 downloads1y agoHugging Face28Dogacel /nemotron-post-training-v2-gpt-oss-120b-regen Dataset Card for Nemotron Post Training v2 gpt-oss-120b Regen Regenerated responses from nvidia/Nemotron-Post-Training-Dataset-v2 dataset using gpt-oss-120b model. Parameter Value Max Tokens 8192 Temperature 1.0 Top-k 20 Top-p 0.95 Repetition Penalty 1.5 Dataset consists only the english samples from the Nemotron Post Training Dataset. 85% of the chat prompts have reasoning enabled, every other category has reasoning disabled. Category Value math… See the full description on the dataset page: https://huggingface.co/datasets/Dogacel/nemotron-post-training-v2-gpt-oss-120b-regen.texttext-generation100K<n<1M2 likes56 downloads5mo agoHugging Face29akahana /Medical-Reasoning-SFT-GPT-OSS-120B Medical-Reasoning-SFT-GPT-OSS-120B A high-quality synthetic dataset of medical reasoning conversations generated using OpenAI's gpt-oss-120B model with reasoning effort set to high, designed for supervised fine-tuning of large language models in healthcare applications. I used Intelligent-Internet/II-Medical-Reasoning-SFT as a seed dataset, so I would like to thank the authors and Intelligent-Internet for their great work. Dataset Statistics Total Samples: 200,927… See the full description on the dataset page: https://huggingface.co/datasets/akahana/Medical-Reasoning-SFT-GPT-OSS-120B.texttext-generation100K<n<1M0 likes54 downloads9mo agoHugging Face30deburky /gpt-oss-claude-code gpt-oss-claude-code SFT dataset Fine-tuning dataset for deburky/gpt-oss-claude-code, a tool-use and agentic coding model based on openai/gpt-oss-20b. Overview 284 training / 71 validation examples Format: gpt-oss harmony (<|start|>, <|channel|>, <|end|> tokens) Mix of knowledge Q&A, coding tasks, and multi-step tool-use conversations Tool-use examples include Read — file reading with offset/limit Grep — pattern search across files Glob — file discovery Bash… See the full description on the dataset page: https://huggingface.co/datasets/deburky/gpt-oss-claude-code.texttext-generationn<1K0 likes53 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.