datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hle-gpt-oss-120b-no-python-260222
hle-gpt-oss-120b-no-python-260222
Deep research agent evaluation on rl-rag/hle_text_only (test split).
Results
Metric
Value
pass@4
47.9%
avg@4
26.6%
Trajectory accuracy
26.6% (2292/8632)
Questions
2158
Trajectories
8632 (4 per question)
Avg tool calls
14.5
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-no-python-260222.gpt-oss-120bSuperior-Reasoning-SFT-gpt-oss-120b-Logprob
Superior-Reasoning-SFT-gpt-oss-120b-Logprob
🚀 Overview
This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset.
🔗 Relationship to Main Dataset
This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120bdataset. Records are linked via a unique sample_uuid.
Main Dataset: Contains the text (prompts… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.browsecomp-gpt-oss-120b-260222
browsecomp-gpt-oss-120b-260222
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
46.8%
avg@4
23.9%
Trajectory accuracy
23.9% (1211/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
26.1
Full conversations
✅
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-gpt-oss-120b-260222.Superior-Reasoning-SFT-gpt-oss-120b-Logprob
Superior-Reasoning-SFT-gpt-oss-120b-Logprob
🚀 Overview
This dataset contains the token-level log-probabilities generated by the teacher model (gpt-oss-120b) for the reasoning samples in the main Superior-Reasoning-SFT-gpt-oss-120b Dataset.
🔗 Relationship to Main Dataset
This dataset is a companion to the main Superior-Reasoning-SFT-gpt-oss-120b dataset. Records are linked via a unique sample_uuid.
Main… See the full description on the dataset page: https://huggingface.co/datasets/erenyeager-1/Superior-Reasoning-SFT-gpt-oss-120b-Logprob.browsecomp-no-scroll-gpt-oss-120b
browsecomp-no-scroll-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
46.0%
avg@4
22.9%
Trajectory accuracy
22.9% (1160/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
27.0
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-no-scroll-gpt-oss-120b.Superior-Reasoning-SFT-gpt-oss-120b
Superior-Reasoning-SFT-gpt-oss-120b
📣 News
Our dataset ranked #1 on the Hugging Face Datasets Trending leaderboard from January 20 to January 30.
🚀 Overview
The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought (Long-CoT) models. Unlike standard distilled datasets that rely on random sampling or… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b.browsecomp-high-effort-gpt-oss-120b
browsecomp-high-effort-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
44.1%
avg@4
22.9%
Trajectory accuracy
22.9% (1158/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
55.4
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
100
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-high-effort-gpt-oss-120b.gpt-oss-20b-combined-outputsMedical-Reasoning-SFT-GPT-OSS-120B
Medical-Reasoning-SFT-GPT-OSS-120B
A high-quality synthetic dataset of medical reasoning conversations generated using OpenAI's gpt-oss-120B model with reasoning effort set to high, designed for supervised fine-tuning of large language models in healthcare applications. I used Intelligent-Internet/II-Medical-Reasoning-SFT as a seed dataset, so I would like to thank the authors and Intelligent-Internet for their great work.
Dataset Statistics
Total Samples: 200,927… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B.gpt-oss-20b-rollouts
GPT-OSS-20B Rollouts
Generated rollouts from GPT-OSS-20B with parsed Harmony channels (assistant thinking/final).
Schema: user_content, system_reasoning_effort, assistant_thinking, assistant_content.
Loading example: load_dataset("andyrdt/gpt-oss-20b-rollouts", "HarmBench", split="standard_train").
Notes
This repository uses manual configuration to expose both subset (config) and split dropdowns in the viewer.
Safety and jailbreak
HarmBench: Safety prompts… See the full description on the dataset page: https://huggingface.co/datasets/andyrdt/gpt-oss-20b-rollouts.hle-gpt-oss-120b-with-python-260222
hle-gpt-oss-120b-with-python-260222
Deep research agent evaluation on unknown.
Results
Metric
Value
pass@4
39.5%
avg@4
17.5%
Trajectory accuracy
17.4% (1860/10660)
Questions
1350
Trajectories
10660 (4 per question)
Avg tool calls
0.0
Full conversations
❌
Model & Setup
Model
unknown
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domainsNone
Tool Usage
Tool
Calls
%… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-with-python-260222.browsecomp-high-effort-full-gpt-oss-120b
browsecomp-high-effort-full-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@1
20.9%
avg@1
20.9%
Trajectory accuracy
20.9% (264/1266)
Questions
1266
Trajectories
1266 (1 per question)
Avg tool calls
52.9
Full conversations
✅
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
100
Temperature
0.7
Blocked domains
huggingface.co… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-high-effort-full-gpt-oss-120b.gpt_oss_20b_doorkey_boundary_activationsgpt-oss120b-generated-perfectblendbrowsecomp-oss-env-high-effort-gpt-oss-120b
browsecomp-oss-env-high-effort-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@1
19.4%
avg@1
19.4%
Trajectory accuracy
19.4% (245/1266)
Questions
1266
Trajectories
1266 (1 per question)
Avg tool calls
52.5
Full conversations
✅
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
100
Temperature
0.7
Blocked domains
huggingface.co… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-oss-env-high-effort-gpt-oss-120b.gpt-oss-sampledgpt-oss-eval-logs-and-scores
This repository contains the detailed evaluation results of gpt-oss models, tested using Twinkle Eval, a robust and efficient AI evaluation tool developed by Twinkle AI. Each entry includes per-question scores across multiple benchmark suites.
Medical-Reasoning-SFT-GPT-OSS-120B-Small
Medical-Reasoning-SFT-GPT-OSS-120B-Small
A filtered and processed version of OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B optimized for training efficiency.
Dataset Description
This dataset contains high-quality medical reasoning conversations with the following modifications:
Length Filtering: Only includes samples where assistant responses are between 1000 and 10000 characters
Reasoning Extraction: Reasoning content from <think> tags has been extracted into a separate… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-Small.gpt-oss-120b-10k
gpt-oss-120b-10k
Repo: tytodd/gpt-oss-120b-10k
Config: /Users/tytodd/Desktop/Modaic/code/core/probe-lab/configs/datasets/10k/10k.yaml
Model: openai/gpt-oss-120b
Runtime: Modal local vLLM on localhost
benchmark
train
val
ood
all
customer_support_tickets_gorkem
1.70%
0.00%
1.55%
mfrc
0.00%
0.00%
0.00%
go_emotions
11.56%
6.90%
11.15%
customer_support_tickets_en
30.27%
20.69%
29.41%
aes2_essay_scoring
26.87%
20.69%
26.32%
ultrafeedback
35.71%
48.28%
36.84%… See the full description on the dataset page: https://huggingface.co/datasets/tytodd/gpt-oss-120b-10k.zelo-scores-10kx100-gpt-oss-20bgpt-oss20b-samples-dedupA simple deduplicated variant of https://huggingface.co/datasets/jxm/gpt-oss20b-samples
Given the predictability of synthetic data we opted for a simple strategy: keeping the unique combinations of first and last ten words. Total count of unique occurrences is available in the column occurrence_count.
gpt-oss-120b-reasoning-STEM-5K
GPT-OSS-120B-Distilled-Reasoning-STEM Dataset
1) Dataset Overview
Data Source Model: gpt-oss-120b-high
Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics)
Data Format: `JSON Lines
Fields: generator, category, input, CoT_Native——reasoning, answer
(Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.)
2) Design Goals (Motivation)
This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.Medical-Reasoning-SFT-GPT-OSS-120B-V2
Medical-Reasoning-SFT-GPT-OSS-120B-V2
A large-scale medical reasoning dataset generated using openai/gpt-oss-120b, containing over 506,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
GPT-OSS-120B is OpenAI's state-of-the-art open-weight model, achieving near-parity with closed models on reasoning benchmarks while being Apache 2.0 licensed.
Dataset Overview
Metric
Value
Model
openai/gpt-oss-120b
Total Samples
506… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-V2.gpt-oss-120b-mandarin-thinking-eval-logs-and-scoresgpt-oss-20b-mandarin-thinking-eval-logs-and-scoresbrowsecomp-gptoss-clean-qwen35-sft
BrowseComp GPT-oss SFT Data (Qwen3.5 Format)
Multi-turn SFT training data for Qwen3.5 models, converted from GPT-oss-120B
BrowseComp trajectories. Available in two formats.
Files
OpenAI Messages Format (recommended for general use)
browsecomp-gptoss-clean-full-messages.json — 372 examples, standard messages format with tool_calls
LLaMA-Factory ShareGPT Format
browsecomp-gptoss-clean-full.json — 372 examples, LLaMA-Factory sharegpt format… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-gptoss-clean-qwen35-sft.gpt-oss-20bwixqa-gpt-oss-120b-all-MiniLM-L6-v2-pgvector-evalsgpt-oss120b-generated-magpie-1m-v0.1
