datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hle-gpt-oss-120b-no-python-260222
hle-gpt-oss-120b-no-python-260222
Deep research agent evaluation on rl-rag/hle_text_only (test split).
Results
Metric
Value
pass@4
47.9%
avg@4
26.6%
Trajectory accuracy
26.6% (2292/8632)
Questions
2158
Trajectories
8632 (4 per question)
Avg tool calls
14.5
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-no-python-260222.browsecomp-gpt-oss-120b-260222
browsecomp-gpt-oss-120b-260222
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
46.8%
avg@4
23.9%
Trajectory accuracy
23.9% (1211/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
26.1
Full conversations
✅
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-gpt-oss-120b-260222.browsecomp-no-scroll-gpt-oss-120b
browsecomp-no-scroll-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
46.0%
avg@4
22.9%
Trajectory accuracy
22.9% (1160/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
27.0
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-no-scroll-gpt-oss-120b.browsecomp-high-effort-gpt-oss-120b
browsecomp-high-effort-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
44.1%
avg@4
22.9%
Trajectory accuracy
22.9% (1158/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
55.4
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
100
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-high-effort-gpt-oss-120b.hle-gpt-oss-120b-with-python-260222
hle-gpt-oss-120b-with-python-260222
Deep research agent evaluation on unknown.
Results
Metric
Value
pass@4
39.5%
avg@4
17.5%
Trajectory accuracy
17.4% (1860/10660)
Questions
1350
Trajectories
10660 (4 per question)
Avg tool calls
0.0
Full conversations
❌
Model & Setup
Model
unknown
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domainsNone
Tool Usage
Tool
Calls
%… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-with-python-260222.browsecomp-high-effort-full-gpt-oss-120b
browsecomp-high-effort-full-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@1
20.9%
avg@1
20.9%
Trajectory accuracy
20.9% (264/1266)
Questions
1266
Trajectories
1266 (1 per question)
Avg tool calls
52.9
Full conversations
✅
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
100
Temperature
0.7
Blocked domains
huggingface.co… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-high-effort-full-gpt-oss-120b.gpt_oss_20b_doorkey_boundary_activationsbrowsecomp-oss-env-high-effort-gpt-oss-120b
browsecomp-oss-env-high-effort-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@1
19.4%
avg@1
19.4%
Trajectory accuracy
19.4% (245/1266)
Questions
1266
Trajectories
1266 (1 per question)
Avg tool calls
52.5
Full conversations
✅
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
100
Temperature
0.7
Blocked domains
huggingface.co… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-oss-env-high-effort-gpt-oss-120b.zelo-scores-10kx100-gpt-oss-20bgpt-oss-120b-mandarin-thinking-eval-logs-and-scoresgpt-oss-20b-mandarin-thinking-eval-logs-and-scoresgpt_oss_maze_acts_120_m11_v1health-qa-gpt-oss-120bv3-eval-judge-gpt-oss-20bGPT-OSS-20B-benchmark-rollouts-512-tokens
GPT-OSS-20B Benchmark Rollouts (512 tokens)
This dataset contains text generation outputs from OpenAI's GPT-OSS-20B model across multiple evaluation benchmarks, with generation limited to 512 tokens.
Dataset Description
The dataset captures GPT-OSS-20B's text generation behavior when responding to prompts from established AI evaluation benchmarks. Each example includes the original prompt, the model's generated response, and token statistics.
Benchmark Coverage… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/GPT-OSS-20B-benchmark-rollouts-512-tokens.nz_research_commons_gpt_oss_120b_openrouter_failed_rows_1200gpt-oss-20b-rolloutsnz_research_commons_gpt_oss_120b_openrouter_results_1200gpt-oss-20b-MedXpertQA-benchmarkBenchmark of openai/gpt-oss-20b against TsinghuaC3I/MedXpertQA dataset, "Text" subset, "test" split.
Accuracy: 27.1%.
Metric
Value
Correct
664
Incorrect
1785
Errors
1
Total samples
2450
Total completion tokens
3,163,003
Raw stats:
{
"accuracy": 0.271,
"correct": 664,
"incorrect": 1785,
"error": 1,
"total": 2450,
"completion_tokens": 3163003
}
GPT-OSS-20B-Distilled-Reasoning-Mini
Dataset Card for Dataset Name
GPT-OSS-20B Distilled Reasoning Dataset Mini
(Multi-stage Evaluative Refinement Method for Reasoning Generation)
Dataset Details and Description
This is a high-quality instruction fine-tuning dataset constructed through knowledge distillation, featuring detailed Chain-of-Thought (CoT) reasoning processes. The dataset is designed to enhance the capabilities of smaller language models in complex reasoning, logical analysis, and instruction… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-20B-Distilled-Reasoning-Mini.gpt_oss_120b_lcbv6_hardest_to_easiest_s_0_e_65_4kx64x5_t_1_gepasimpleqa_verified_gpt-oss_scored
simpleqa_verified_gpt-oss_scored Evaluation Results
Eval created with evaljobs.
This dataset contains evaluation results for the model(s) hf-inference-providers/openai/gpt-oss-20b:cheapest,hf-inference-providers/openai/gpt-oss-120b:cheapest using the eval script simpleqa_verified-integration-tests.
To browse the results interactively, visit this Space.
How to Run This Eval
pip install git+https://github.com/dvsrepo/evaljobs.git
export HF_TOKEN=your_token_here… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/simpleqa_verified_gpt-oss_scored.distill-cl-biography-independent-ab-v4-gpt-oss-120b-v37-3550p32v-a
Independent synthetic biographies, role A
This public dataset is the authenticated 3,550-person role-A prefix used
by the scratch GPT-2-medium A/B experiment. It contains 32 independently
generated biography views per person (113,600 rows) and a separate canonical
four-question QA bundle per person (14,200 rows).
The biographies were generated with the pinned openai/gpt-oss-120b v37
workflow. The biographies configuration exposes split train; the qa
configuration exposes split… See the full description on the dataset page: https://huggingface.co/datasets/jhdlee/distill-cl-biography-independent-ab-v4-gpt-oss-120b-v37-3550p32v-a.open-scholar-gpt-oss-120b
open-scholar-gpt-oss-120b
Deep research agent evaluation on data/drtulu_open_scholar.jsonl (normal split).
Results
Metric
Value
pass@1
0.0%
avg@1
0.0%
Trajectory accuracy
0.0% (0/11854)
Questions
11854
Trajectories
11854 (1 per question)
Avg tool calls
17.3
Full conversations
✅
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
None
Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/open-scholar-gpt-oss-120b.TRIM-gpt-oss-120b-separate-neighbors-only-para-random-feature-num
TRIM Agent Reasoning Messages (HF Public Export)
This directory is a Hugging Face-friendly public export of the TRIM agent reasoning SFT data.
What Is Included
Provider: vllm
Model: gpt-oss-120b
SFT mode: local_neighbor_only
Splits present: train
Records in this export manifest: 10056
Tasks in this split: AMES, BBB_Martins, Bioavailability_Ma, CYP2C9_Substrate_CarbonMangels, CYP2D6_Substrate_CarbonMangels, CYP3A4_Substrate_CarbonMangels, Carcinogens_Lagunin, ClinTox… See the full description on the dataset page: https://huggingface.co/datasets/Kiria-Nozan/TRIM-gpt-oss-120b-separate-neighbors-only-para-random-feature-num.distill-cl-biography-independent-ab-v4-gpt-oss-120b-v37-3550p32v-b
Independent synthetic biographies, role B
This public dataset is the authenticated 3,550-person role-B prefix used
by the scratch GPT-2-medium A/B experiment. It contains 32 independently
generated biography views per person (113,600 rows) and a separate canonical
four-question QA bundle per person (14,200 rows).
The biographies were generated with the pinned openai/gpt-oss-120b v37
workflow. The biographies configuration exposes split train; the qa
configuration exposes split… See the full description on the dataset page: https://huggingface.co/datasets/jhdlee/distill-cl-biography-independent-ab-v4-gpt-oss-120b-v37-3550p32v-b.gpt_oss_120b_sf_all_correctgpt_oss_doorkey_action_distributions
GPT-OSS-20B DoorKey action-distribution time series
This dataset contains sentence-prefix next-action readouts for 46 fixed DoorKey
environment states drawn from 31 trajectories in
project-telos/trajectories_key_door_100.
It also contains the corresponding offline BEAST change-point results.
The dataset has 7,084 positions. Position 0 is the readout before any reasoning
text is revealed. Each subsequent position reveals one additional reasoning
sentence from the same model… See the full description on the dataset page: https://huggingface.co/datasets/project-telos/gpt_oss_doorkey_action_distributions.jailbreak-gpt-oss-120b-highbfcl-gpt-oss-20b-test
bfcl-gpt-oss-20b-test Evaluation Results
Eval created with evaljobs.
This dataset contains evaluation results for the model(s) hf-inference-providers/openai/gpt-oss-20b:fastest using the eval inspect_evals/bfcl from Inspect Evals.
To browse the results interactively, visit this Space.
Command
This eval was run with:
evaljobs inspect_evals/bfcl \
--model hf-inference-providers/openai/gpt-oss-20b:fastest \
--name bfcl-gpt-oss-20b-test \
--limit 50… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/bfcl-gpt-oss-20b-test.
