CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rl-rag /hle-gpt-oss-120b-no-python-260222 hle-gpt-oss-120b-no-python-260222 Deep research agent evaluation on rl-rag/hle_text_only (test split). Results Metric Value pass@4 47.9% avg@4 26.6% Trajectory accuracy 26.6% (2292/8632) Questions 2158 Trajectories 8632 (4 per question) Avg tool calls 14.5 Full conversations ❌ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains huggingface.co Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-no-python-260222.tabular1K<n<10K1 likes6.6k downloads7mo agoHugging Face02masterpieceexternal /gpt-oss-20b-moe-expert-power-traces-320k GPT-OSS-20B MoE Expert Power Traces (320k, ChipWhisperer) This dataset contains analog power traces captured with a ChipWhisperer Husky while running forced single-expert MoE computations derived from openai/gpt-oss-20b on an NVIDIA H100. What is recorded Each trace corresponds to one capture trial where: A fixed expert id is selected (expert_00 ... expert_31). A random hidden-state tensor is generated once per trial. The selected expert computation is executed… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-moe-expert-power-traces-320k.audio-classification100K<n<1M0 likes5.9k downloads4mo agoHugging Face03Nilaksh404 /gpt-oss-120bdocumentn<1K0 likes3.2k downloads10mo agoHugging Face04Alibaba-Apsara /Superior-Reasoning-SFT-gpt-oss-120b-Logprob Superior-Reasoning-SFT-gpt-oss-120b-Logprob           🚀 Overview This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset. 🔗 Relationship to Main Dataset This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120bdataset. Records are linked via a unique sample_uuid. Main Dataset: Contains the text (prompts… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.texttext-generation100K<n<1M63 likes2.2k downloads8mo agoHugging Face05rl-rag /browsecomp-gpt-oss-120b-260222 browsecomp-gpt-oss-120b-260222 Deep research agent evaluation on data/browsecomp.jsonl (normal split). Results Metric Value pass@4 46.8% avg@4 23.9% Trajectory accuracy 23.9% (1211/5064) Questions 1266 Trajectories 5064 (4 per question) Avg tool calls 26.1 Full conversations ✅ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains huggingface.co Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-gpt-oss-120b-260222.tabular1K<n<10K0 likes1.8k downloads7mo agoHugging Face06erenyeager-1 /Superior-Reasoning-SFT-gpt-oss-120b-Logprob Superior-Reasoning-SFT-gpt-oss-120b-Logprob &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; 🚀 Overview This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset. 🔗 Relationship to Main Dataset This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120b dataset. Records are linked via a unique sample_uuid. Main… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.texttext-generation100K<n<1M0 likes1.6k downloads2mo agoHugging Face07rl-rag /browsecomp-no-scroll-gpt-oss-120b browsecomp-no-scroll-gpt-oss-120b Deep research agent evaluation on data/browsecomp.jsonl (normal split). Results Metric Value pass@4 46.0% avg@4 22.9% Trajectory accuracy 22.9% (1160/5064) Questions 1266 Trajectories 5064 (4 per question) Avg tool calls 27.0 Full conversations ❌ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains huggingface.co Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-no-scroll-gpt-oss-120b.tabular1K<n<10K0 likes1.5k downloads7mo agoHugging Face08Alibaba-Apsara /Superior-Reasoning-SFT-gpt-oss-120b Superior-Reasoning-SFT-gpt-oss-120b           📣 News Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30. 🚀 Overview The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b.texttext-generation100K<n<1M352 likes1.4k downloads8mo agoHugging Face09rl-rag /browsecomp-high-effort-gpt-oss-120b browsecomp-high-effort-gpt-oss-120b Deep research agent evaluation on data/browsecomp.jsonl (normal split). Results Metric Value pass@4 44.1% avg@4 22.9% Trajectory accuracy 22.9% (1158/5064) Questions 1266 Trajectories 5064 (4 per question) Avg tool calls 55.4 Full conversations ❌ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 100 Temperature 0.7 Blocked domains huggingface.co Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-high-effort-gpt-oss-120b.tabular1K<n<10K0 likes1.2k downloads7mo agoHugging Face10jacobmorrison /gpt-oss-20b-combined-outputstext1M<n<10M0 likes816 downloads1y agoHugging Face11OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B Medical-Reasoning-SFT-GPT-OSS-120B A high-quality synthetic dataset of medical reasoning conversations generated using OpenAI's gpt-oss-120B model with reasoning effort set to high, designed for supervised fine-tuning of large language models in healthcare applications. I used Intelligent-Internet/II-Medical-Reasoning-SFT as a seed dataset, so I would like to thank the authors and Intelligent-Internet for their great work. Dataset Statistics Total Samples: 200,927… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B.texttext-generation100K<n<1M255 likes720 downloads10mo agoHugging Face12aimosprite /gpt-oss-120b-high-reasoning-firsthalf-g2-statsPrivate artifact repo for GPT-OSS-120B G^2 collection. Selection: dataset: aimosprite/high-reasoning-eval usage half: partitions 0,1,2 selected traces: 183 Contents: g2_stats/: per-layer sum_g2 safetensors for routed MoE weights selected_traces_manifest.json: trace selection manifest selected_traces.jsonl: exact selected traces used for stats collection Run family: shared base strategy: zero target rank for downstream materialization: 256 model: unsloth/gpt-oss-120b-BF16 0 likes713 downloads6mo agoHugging Face13andyrdt /gpt-oss-20b-rollouts GPT-OSS-20B Rollouts Generated rollouts from GPT-OSS-20B with parsed Harmony channels (assistant thinking/final). Schema: user_content, system_reasoning_effort, assistant_thinking, assistant_content. Loading example: load_dataset("andyrdt/gpt-oss-20b-rollouts", "HarmBench", split="standard_train"). Notes This repository uses manual configuration to expose both subset (config) and split dropdowns in the viewer. Safety and jailbreak HarmBench: Safety prompts… See the full description on the dataset page: https://huggingface.co/datasets/andyrdt/gpt-oss-20b-rollouts.text1M<n<10M7 likes705 downloads9mo agoHugging Face14rl-rag /browsecomp-gpt-oss-120b-v2 BrowseComp GPT-oss-120B Evaluation (v2) Deep research agent evaluation on BrowseComp using GPT-oss-120B with the elastic-serving OSS engine. Results Metric Value pass@1 37.9% Questions 1,266 Correct 480/1,266 Success rate 1132/1,266 Avg tool calls ~61 Judge GPT-4.1 Model & Setup Setting Value Model GPT-oss-120B Engine elastic-serving OSS (Harmony protocol) Reasoning effort high Max turns unlimited (until model… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-gpt-oss-120b-v2.0 likes556 downloads6mo agoHugging Face15rl-rag /hle-gpt-oss-120b-with-python-260222 hle-gpt-oss-120b-with-python-260222 Deep research agent evaluation on unknown. Results Metric Value pass@4 39.5% avg@4 17.5% Trajectory accuracy 17.4% (1860/10660) Questions 1350 Trajectories 10660 (4 per question) Avg tool calls 0.0 Full conversations ❌ Model & Setup Model unknown Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domainsNone Tool Usage Tool Calls %… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-with-python-260222.tabular10K<n<100K0 likes521 downloads7mo agoHugging Face16cminst /SimpleDeco-gptoss20b0 likes514 downloads6mo agoHugging Face17rl-rag /browsecomp-high-effort-full-gpt-oss-120b browsecomp-high-effort-full-gpt-oss-120b Deep research agent evaluation on data/browsecomp.jsonl (normal split). Results Metric Value pass@1 20.9% avg@1 20.9% Trajectory accuracy 20.9% (264/1266) Questions 1266 Trajectories 1266 (1 per question) Avg tool calls 52.9 Full conversations ✅ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 100 Temperature 0.7 Blocked domains huggingface.co… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-high-effort-full-gpt-oss-120b.tabular1K<n<10K0 likes506 downloads6mo agoHugging Face18project-telos /gpt_oss_20b_doorkey_boundary_activationstabularn<1K0 likes476 downloads3mo agoHugging Face19baseten-admin /gpt-oss120b-generated-perfectblendtext100K<n<1M1 likes467 downloads1y agoHugging Face20rl-rag /browsecomp-oss-env-high-effort-gpt-oss-120b browsecomp-oss-env-high-effort-gpt-oss-120b Deep research agent evaluation on data/browsecomp.jsonl (normal split). Results Metric Value pass@1 19.4% avg@1 19.4% Trajectory accuracy 19.4% (245/1266) Questions 1266 Trajectories 1266 (1 per question) Avg tool calls 52.5 Full conversations ✅ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 100 Temperature 0.7 Blocked domains huggingface.co… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-oss-env-high-effort-gpt-oss-120b.tabular1K<n<10K0 likes394 downloads6mo agoHugging Face21masterpieceexternal /gpt-oss-20b-moe-expert-power-traces-320k-ds16k GPT-OSS-20B MoE Expert Power Traces (Downsampled to 16k) Downsampled variant of the 320k expert-trace capture set. Source Raw source dataset (same captures): 32 experts (expert_00..expert_31) 10,000 traces per expert 320,000 total traces raw trace length ~195k samples per trace Downsampling method Each raw trace was resampled to exactly 16384 samples using linear interpolation (np.interp) matching the trainer resampling step. No baseline normalization and no… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-moe-expert-power-traces-320k-ds16k.audio-classification100K<n<1M0 likes362 downloads7mo agoHugging Face22zerostratos /gpt-oss-sampledtext1M<n<10M0 likes305 downloads1y agoHugging Face23twinkle-ai /gpt-oss-eval-logs-and-scores This repository contains the detailed evaluation results of gpt-oss models, tested using Twinkle Eval, a robust and efficient AI evaluation tool developed by Twinkle AI. Each entry includes per-question scores across multiple benchmark suites. text1K<n<10K1 likes299 downloads1y agoHugging Face24tytodd /gpt-oss-120b-10k gpt-oss-120b-10k Repo: tytodd/gpt-oss-120b-10k Config: /Users/tytodd/Desktop/Modaic/code/core/probe-lab/configs/datasets/10k/10k.yaml Model: openai/gpt-oss-120b Runtime: Modal local vLLM on localhost benchmark train val ood all customer_support_tickets_gorkem 1.70% 0.00% 1.55% mfrc 0.00% 0.00% 0.00% go_emotions 11.56% 6.90% 11.15% customer_support_tickets_en 30.27% 20.69% 29.41% aes2_essay_scoring 26.87% 20.69% 26.32% ultrafeedback 35.71% 48.28% 36.84%… See the full description on the dataset page: https://huggingface.co/datasets/tytodd/gpt-oss-120b-10k.text10K<n<100K0 likes278 downloads6mo agoHugging Face25OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B-Small Medical-Reasoning-SFT-GPT-OSS-120B-Small A filtered and processed version of OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B optimized for training efficiency. Dataset Description This dataset contains high-quality medical reasoning conversations with the following modifications: Length Filtering: Only includes samples where assistant responses are between 1000 and 10000 characters Reasoning Extraction: Reasoning content from <think> tags has been extracted into a separate… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-Small.texttext-generation100K<n<1M3 likes276 downloads9mo agoHugging Face26tomaarsen /zelo-scores-10kx100-gpt-oss-20btabular100K<n<1M0 likes259 downloads5mo agoHugging Face27PleIAs /gpt-oss20b-samples-dedupA simple deduplicated variant of https://huggingface.co/datasets/jxm/gpt-oss20b-samples Given the predictability of synthetic data we opted for a simple strategy: keeping the unique combinations of first and last ten words. Total count of unique occurrences is available in the column occurrence_count. text100K<n<1M5 likes257 downloads1y agoHugging Face28Jackrong /gpt-oss-120b-reasoning-STEM-5K GPT-OSS-120B-Distilled-Reasoning-STEM Dataset 1) Dataset Overview Data Source Model: gpt-oss-120b-high Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics) Data Format: `JSON Lines Fields: generator, category, input, CoT_Native——reasoning, answer (Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.) 2) Design Goals (Motivation) This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.textquestion-answering1K<n<10K12 likes249 downloads1y agoHugging Face29OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B-V2 Medical-Reasoning-SFT-GPT-OSS-120B-V2 A large-scale medical reasoning dataset generated using openai/gpt-oss-120b, containing over 506,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. GPT-OSS-120B is OpenAI's state-of-the-art open-weight model, achieving near-parity with closed models on reasoning benchmarks while being Apache 2.0 licensed. Dataset Overview Metric Value Model openai/gpt-oss-120b Total Samples 506… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-V2.texttext-generation100K<n<1M9 likes213 downloads8mo agoHugging Face30twinkle-ai /gpt-oss-120b-mandarin-thinking-eval-logs-and-scorestabular100K<n<1M0 likes181 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.