CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Alibaba-Apsara /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b           📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M352 likes1.4k downloads8mo agoHugging Face02project-telos /gpt_oss_20b_doorkey_boundary_activationstabularn<1K0 likes463 downloads3mo agoHugging Face03baseten-admin /gpt-oss120b-generated-perfectblendtext100K<n<1M1 likes461 downloads1y agoHugging Face04twinkle-ai /gpt-oss-eval-logs-and-scores This repository contains the detailed evaluation results of gpt-oss models, tested using Twinkle Eval, a robust and efficient AI evaluation tool developed by Twinkle AI. Each entry includes per-question scores across multiple benchmark suites. text1K<n<10K1 likes299 downloads1y agoHugging Face05Jackrong /gpt-oss-120b-reasoning-STEM-5K GPT-OSS-120B-Distilled-Reasoning-STEM Dataset 1) Dataset Overview Data Source Model: gpt-oss-120b-high Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics) Data Format: `JSON Lines Fields: generator, category, input, CoT_Native——reasoning, answer (Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.) 2) Design Goals (Motivation) This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.textquestion-answering1K<n<10K12 likes250 downloads1y agoHugging Face06twinkle-ai /gpt-oss-120b-mandarin-thinking-eval-logs-and-scorestabular100K<n<1M0 likes182 downloads7mo agoHugging Face07twinkle-ai /gpt-oss-20b-mandarin-thinking-eval-logs-and-scorestabular100K<n<1M0 likes179 downloads7mo agoHugging Face08rl-rag /browsecomp-gptoss-clean-qwen35-sft BrowseComp GPT-oss SFT Data (Qwen3.5 Format) Multi-turn SFT training data for Qwen3.5 models, converted from GPT-oss-120B BrowseComp trajectories. Available in two formats. Files OpenAI Messages Format (recommended for general use) browsecomp-gptoss-clean-full-messages.json — 372 examples, standard messages format with tool_calls LLaMA-Factory ShareGPT Format browsecomp-gptoss-clean-full.json — 372 examples, LLaMA-Factory sharegpt format… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-gptoss-clean-qwen35-sft.textn<1K0 likes170 downloads5mo agoHugging Face09bdanko /wixqa-gpt-oss-120b-all-MiniLM-L6-v2-pgvector-evalstext10K<n<100K0 likes163 downloads6mo agoHugging Face10baseten-admin /gpt-oss120b-generated-magpie-1m-v0.1text100K<n<1M2 likes139 downloads1y agoHugging Face11project-telos /gpt_oss_maze_acts_120_m11_v1tabularn<1K0 likes135 downloads1mo agoHugging Face12Jackrong /Natural-Reasoning-gpt-oss-120B-S1 Dataset Card: Natural-Reasoning-gpt-oss-120B-S1 📜 Dataset Overview This is a meticulously curated instruction fine-tuning dataset designed specifically for efficient knowledge distillation tasks. Built upon the first 100,000 questions from the large-scale reasoning corpus facebook/natural_reasoning (s1, I will process the remaining parts later), it aims to transfer the advanced, multi-step reasoning capabilities of the teacher model gpt-oss-120-high to a student model… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Natural-Reasoning-gpt-oss-120B-S1.texttext-generation10K<n<100K24 likes102 downloads1y agoHugging Face130xzanuee /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or heuristic filtering, Superior-Reasoning-SFT-gpt-oss-120b is constructed using a principled… See the full description on the dataset page: https://huggingface.co/datasets/0xzanuee/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M0 likes102 downloads8mo agoHugging Face14AlgoDriveAI /TinyMathStories_gpt-oss-20b TinyMathStories A TinyStories-style corpus extended with math and lightweight reasoning. This dataset keeps the child-level vocabulary and short narrative style of TinyStories (Microsoft Research, Eldan & Li, 2023) and mixes in basic numeracy (counting, addition/subtraction, simple equations, fractions, measurement) and short justifications—so tiny models can practice coherent English and early math/logic. Research, generation, and curation by AlgoDriveAI.Inspired by and… See the full description on the dataset page: https://huggingface.co/datasets/AlgoDriveAI/TinyMathStories_gpt-oss-20b.texttext-generation100K<n<1M0 likes101 downloads9mo agoHugging Face15Jackrong /gpt-oss-120B-distilled-reasoning GPT-oss-120B-Distilled-Reasoning-math Dataset Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines Fields: Generator, Category, Input, Output Core Statistics Generated complete reasoning processes and answers using gpt-oss-120b (MXFP4).The text length of the dataset reflects the depth and complexity of its content. I have statistically analyzed the lengths of the input (question), Reasoning, and Answer.To understand the data… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120B-distilled-reasoning.texttext-classification1K<n<10K20 likes87 downloads1y agoHugging Face16Chaiyphop /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or heuristic filtering, Superior-Reasoning-SFT-gpt-oss-120b is constructed using a principled… See the full description on the dataset page: https://huggingface.co/datasets/Chaiyphop/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M1 likes80 downloads8mo agoHugging Face17Jackrong /GPT-OSS-120B-Distilled-Reasoning-math GPT-oss-120B-Distilled-Reasoning-math Dataset Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines Fields: Generator, Category, Input, CoT_Native_Reasoning, Reasoning, Answer Core Statistics Generated complete reasoning processes and answers using gpt-oss-120b (MXFP4).The text length of the dataset reflects the depth and complexity of its content. I have statistically analyzed the lengths of the input (question), Reasoning, and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-120B-Distilled-Reasoning-math.textquestion-answering1K<n<10K9 likes78 downloads1y agoHugging Face18erenyeager-1 /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M0 likes76 downloads2mo agoHugging Face19NarsAI /Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b           📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or… See the full description on the dataset page: https://huggingface.co/datasets/NarsAI/Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M0 likes72 downloads8mo agoHugging Face20Jackrong /ShareGPT-gpt-oss-120B-reasoning ShareGPT-style multi-turn distillation with gpt-oss-120B (reasoning) teacher. There is a misannotation in the generation section: I used medium‑reasoning dialogues (which fit within the context window), whereas high‑reasoning multi‑turn conversations would exceed it. Originally, when annotating, I intended to use high‑reasoning ones. texttext-classification1K<n<10K6 likes68 downloads10mo agoHugging Face21Lilbullet /prompt-injection-artificial-GPTOSS120b Prompt Injection (Synthetic) — GPT-OSS-120b This dataset contains a small collection of synthetic user prompts and Noraml user prompts designed to finetune Large Language Models (LLMs) against malicious prompt-injection / jailbreak attempts, including cases that use obfuscation (e.g., Base64, leetspeak, typos, irregular spacing) to evade safety filters. Dataset Summary Source repository: Lilbullet/prompt-injection-artificial-GPTOSS120b Model used: GPT-OSS-120b Generation… See the full description on the dataset page: https://huggingface.co/datasets/Lilbullet/prompt-injection-artificial-GPTOSS120b.textn<1K1 likes68 downloads8mo agoHugging Face22Dogacel /nemotron-post-training-v2-gpt-oss-120b-regen Dataset Card for Nemotron Post Training v2 gpt-oss-120b Regen Regenerated responses from nvidia/Nemotron-Post-Training-Dataset-v2 dataset using gpt-oss-120b model. Parameter Value Max Tokens 8192 Temperature 1.0 Top-k 20 Top-p 0.95 Repetition Penalty 1.5 Dataset consists only the english samples from the Nemotron Post Training Dataset. 85% of the chat prompts have reasoning enabled, every other category has reasoning disabled. Category Value math… See the full description on the dataset page: https://huggingface.co/datasets/Dogacel/nemotron-post-training-v2-gpt-oss-120b-regen.texttext-generation100K<n<1M2 likes59 downloads5mo agoHugging Face23zyx1234 /MuSeR_GPT_OSS_120B_DistillationThis dataset contains ~100k synthetic medical queries and corresponding responses distilled from GPT-OSS-120B. The generation of synthetic medical queries follows an attribute-conditioned generation method proposed in paper Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning. We found that supervised fine-tuning on this dataset can substantially improve LLMs' medical conversational capabilities. See our paper and project page for more details. If… See the full description on the dataset page: https://huggingface.co/datasets/zyx1234/MuSeR_GPT_OSS_120B_Distillation.textquestion-answering10K<n<100K4 likes58 downloads9mo agoHugging Face24deburky /gpt-oss-claude-code gpt-oss-claude-code SFT dataset Fine-tuning dataset for deburky/gpt-oss-claude-code, a tool-use and agentic coding model based on openai/gpt-oss-20b. Overview 284 training / 71 validation examples Format: gpt-oss harmony (<|start|>, <|channel|>, <|end|> tokens) Mix of knowledge Q&A, coding tasks, and multi-step tool-use conversations Tool-use examples include Read — file reading with offset/limit Grep — pattern search across files Glob — file discovery Bash… See the full description on the dataset page: https://huggingface.co/datasets/deburky/gpt-oss-claude-code.texttext-generationn<1K0 likes57 downloads6mo agoHugging Face25Jackrong /GPT-OSS-20B-Distilled-Reasoning-Mini Dataset Card for Dataset Name GPT-OSS-20B Distilled Reasoning Dataset Mini (Multi-stage Evaluative Refinement Method for Reasoning Generation) Dataset Details and Description This is a high-quality instruction fine-tuning dataset constructed through knowledge distillation, featuring detailed Chain-of-Thought (CoT) reasoning processes. The dataset is designed to enhance the capabilities of smaller language models in complex reasoning, logical analysis, and instruction… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-20B-Distilled-Reasoning-Mini.tabulartext-classification1K<n<10K21 likes55 downloads1y agoHugging Face26kth8 /gpt-oss-20b-MedXpertQA-benchmarkBenchmark of openai/gpt-oss-20b against TsinghuaC3I/MedXpertQA dataset, "Text" subset, "test" split. Accuracy: 27.1%. Metric Value Correct 664 Incorrect 1785 Errors 1 Total samples 2450 Total completion tokens 3,163,003 Raw stats: { "accuracy": 0.271, "correct": 664, "incorrect": 1785, "error": 1, "total": 2450, "completion_tokens": 3163003 } tabular1K<n<10K0 likes52 downloads5mo agoHugging Face27anshy /Superior-Reasoning-SFT-gpt-oss-120b-random-shuffled Superior-Reasoning-SFT-gpt-oss-120b &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that… See the full description on the dataset page: https://huggingface.co/datasets/anshy/Superior-Reasoning-SFT-gpt-oss-120b-random-shuffled.texttext-generation100K<n<1M0 likes48 downloads4mo agoHugging Face28hardrave /bushcraft_survival_gpt_oss_data_distilled 🧟 ZombieLLM — Bushcraft Survival Distilled by GPT-OSS-20B A distilled instruction–response dataset built for the ZombieLLM project.We reanimated CoT_Reasoning_Bushcraft_Survival by keeping its survival questions and replacing the original chain-of-thought (CoT) answers with concise, high-quality final-only responses generated by GPT-OSS-20B using the Harmony chat template. This dataset was used in the domain tuning stage of ZombieLLM, injecting survival knowledge and grounding… See the full description on the dataset page: https://huggingface.co/datasets/hardrave/bushcraft_survival_gpt_oss_data_distilled.text1K<n<10K2 likes47 downloads1y agoHugging Face29rl-rag /gpt_oss_120b_sf_all_correcttabular10K<n<100K0 likes47 downloads6mo agoHugging Face30Kiria-Nozan /TRIM-gpt-oss-120b-separate-neighbors-only-para-random-feature-num TRIM Agent Reasoning Messages (HF Public Export) This directory is a Hugging Face-friendly public export of the TRIM agent reasoning SFT data. What Is Included Provider: vllm Model: gpt-oss-120b SFT mode: local_neighbor_only Splits present: train Records in this export manifest: 10056 Tasks in this split: AMES, BBB_Martins, Bioavailability_Ma, CYP2C9_Substrate_CarbonMangels, CYP2D6_Substrate_CarbonMangels, CYP3A4_Substrate_CarbonMangels, Carcinogens_Lagunin, ClinTox… See the full description on the dataset page: https://huggingface.co/datasets/Kiria-Nozan/TRIM-gpt-oss-120b-separate-neighbors-only-para-random-feature-num.tabular10K<n<100K0 likes47 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.