CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rl-rag /hle-gpt-oss-120b-no-python-260222 hle-gpt-oss-120b-no-python-260222 Deep research agent evaluation on rl-rag/hle_text_only (test split). Results Metric Value pass@4 47.9% avg@4 26.6% Trajectory accuracy 26.6% (2292/8632) Questions 2158 Trajectories 8632 (4 per question) Avg tool calls 14.5 Full conversations ❌ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains huggingface.co Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-no-python-260222.tabular1K<n<10K1 likes6.6k downloads7mo agoHugging Face02rl-rag /browsecomp-gpt-oss-120b-260222 browsecomp-gpt-oss-120b-260222 Deep research agent evaluation on data/browsecomp.jsonl (normal split). Results Metric Value pass@4 46.8% avg@4 23.9% Trajectory accuracy 23.9% (1211/5064) Questions 1266 Trajectories 5064 (4 per question) Avg tool calls 26.1 Full conversations ✅ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains huggingface.co Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-gpt-oss-120b-260222.tabular1K<n<10K0 likes1.8k downloads7mo agoHugging Face03rl-rag /browsecomp-no-scroll-gpt-oss-120b browsecomp-no-scroll-gpt-oss-120b Deep research agent evaluation on data/browsecomp.jsonl (normal split). Results Metric Value pass@4 46.0% avg@4 22.9% Trajectory accuracy 22.9% (1160/5064) Questions 1266 Trajectories 5064 (4 per question) Avg tool calls 27.0 Full conversations ❌ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains huggingface.co Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-no-scroll-gpt-oss-120b.tabular1K<n<10K0 likes1.5k downloads7mo agoHugging Face04rl-rag /browsecomp-high-effort-gpt-oss-120b browsecomp-high-effort-gpt-oss-120b Deep research agent evaluation on data/browsecomp.jsonl (normal split). Results Metric Value pass@4 44.1% avg@4 22.9% Trajectory accuracy 22.9% (1158/5064) Questions 1266 Trajectories 5064 (4 per question) Avg tool calls 55.4 Full conversations ❌ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 100 Temperature 0.7 Blocked domains huggingface.co Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-high-effort-gpt-oss-120b.tabular1K<n<10K0 likes1.2k downloads7mo agoHugging Face05rl-rag /hle-gpt-oss-120b-with-python-260222 hle-gpt-oss-120b-with-python-260222 Deep research agent evaluation on unknown. Results Metric Value pass@4 39.5% avg@4 17.5% Trajectory accuracy 17.4% (1860/10660) Questions 1350 Trajectories 10660 (4 per question) Avg tool calls 0.0 Full conversations ❌ Model & Setup Model unknown Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domainsNone Tool Usage Tool Calls %… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-with-python-260222.tabular10K<n<100K0 likes521 downloads7mo agoHugging Face06rl-rag /browsecomp-high-effort-full-gpt-oss-120b browsecomp-high-effort-full-gpt-oss-120b Deep research agent evaluation on data/browsecomp.jsonl (normal split). Results Metric Value pass@1 20.9% avg@1 20.9% Trajectory accuracy 20.9% (264/1266) Questions 1266 Trajectories 1266 (1 per question) Avg tool calls 52.9 Full conversations ✅ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 100 Temperature 0.7 Blocked domains huggingface.co… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-high-effort-full-gpt-oss-120b.tabular1K<n<10K0 likes506 downloads6mo agoHugging Face07project-telos /gpt_oss_20b_doorkey_boundary_activationstabularn<1K0 likes476 downloads3mo agoHugging Face08rl-rag /browsecomp-oss-env-high-effort-gpt-oss-120b browsecomp-oss-env-high-effort-gpt-oss-120b Deep research agent evaluation on data/browsecomp.jsonl (normal split). Results Metric Value pass@1 19.4% avg@1 19.4% Trajectory accuracy 19.4% (245/1266) Questions 1266 Trajectories 1266 (1 per question) Avg tool calls 52.5 Full conversations ✅ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 100 Temperature 0.7 Blocked domains huggingface.co… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-oss-env-high-effort-gpt-oss-120b.tabular1K<n<10K0 likes394 downloads6mo agoHugging Face09tomaarsen /zelo-scores-10kx100-gpt-oss-20btabular100K<n<1M0 likes259 downloads5mo agoHugging Face10twinkle-ai /gpt-oss-120b-mandarin-thinking-eval-logs-and-scorestabular100K<n<1M0 likes181 downloads7mo agoHugging Face11twinkle-ai /gpt-oss-20b-mandarin-thinking-eval-logs-and-scorestabular100K<n<1M0 likes177 downloads7mo agoHugging Face12project-telos /gpt_oss_maze_acts_120_m11_v1tabularn<1K0 likes106 downloads1mo agoHugging Face13andersonbcdefg /health-qa-gpt-oss-120btabular100K<n<1M2 likes100 downloads1y agoHugging Face14ACSci /v3-eval-judge-gpt-oss-20btabular10K<n<100K0 likes93 downloads5mo agoHugging Face15AmanPriyanshu /GPT-OSS-20B-benchmark-rollouts-512-tokens GPT-OSS-20B Benchmark Rollouts (512 tokens) This dataset contains text generation outputs from OpenAI's GPT-OSS-20B model across multiple evaluation benchmarks, with generation limited to 512 tokens. Dataset Description The dataset captures GPT-OSS-20B's text generation behavior when responding to prompts from established AI evaluation benchmarks. Each example includes the original prompt, the model's generated response, and token statistics. Benchmark Coverage… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/GPT-OSS-20B-benchmark-rollouts-512-tokens.tabular10K<n<100K0 likes87 downloads1y agoHugging Face16dinushiTJ /nz_research_commons_gpt_oss_120b_openrouter_failed_rows_1200tabularn<1K0 likes79 downloads2mo agoHugging Face17mntss /gpt-oss-20b-rolloutstabular1M<n<10M0 likes78 downloads1y agoHugging Face18dinushiTJ /nz_research_commons_gpt_oss_120b_openrouter_results_1200tabular1K<n<10K0 likes74 downloads2mo agoHugging Face19kth8 /gpt-oss-20b-MedXpertQA-benchmarkBenchmark of openai/gpt-oss-20b against TsinghuaC3I/MedXpertQA dataset, "Text" subset, "test" split. Accuracy: 27.1%. Metric Value Correct 664 Incorrect 1785 Errors 1 Total samples 2450 Total completion tokens 3,163,003 Raw stats: { "accuracy": 0.271, "correct": 664, "incorrect": 1785, "error": 1, "total": 2450, "completion_tokens": 3163003 } tabular1K<n<10K0 likes57 downloads5mo agoHugging Face20Jackrong /GPT-OSS-20B-Distilled-Reasoning-Mini Dataset Card for Dataset Name GPT-OSS-20B Distilled Reasoning Dataset Mini (Multi-stage Evaluative Refinement Method for Reasoning Generation) Dataset Details and Description This is a high-quality instruction fine-tuning dataset constructed through knowledge distillation, featuring detailed Chain-of-Thought (CoT) reasoning processes. The dataset is designed to enhance the capabilities of smaller language models in complex reasoning, logical analysis, and instruction… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-20B-Distilled-Reasoning-Mini.tabulartext-classification1K<n<10K21 likes54 downloads1y agoHugging Face21anirudhb11 /gpt_oss_120b_lcbv6_hardest_to_easiest_s_0_e_65_4kx64x5_t_1_gepatabular1K<n<10K0 likes53 downloads7mo agoHugging Face22dvilasuero /simpleqa_verified_gpt-oss_scored simpleqa_verified_gpt-oss_scored Evaluation Results Eval created with evaljobs. This dataset contains evaluation results for the model(s) hf-inference-providers/openai/gpt-oss-20b:cheapest,hf-inference-providers/openai/gpt-oss-120b:cheapest using the eval script simpleqa_verified-integration-tests. To browse the results interactively, visit this Space. How to Run This Eval pip install git+https://github.com/dvsrepo/evaljobs.git export HF_TOKEN=your_token_here… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/simpleqa_verified_gpt-oss_scored.tabularn<1K0 likes51 downloads10mo agoHugging Face23jhdlee /distill-cl-biography-independent-ab-v4-gpt-oss-120b-v37-3550p32v-a Independent synthetic biographies, role A This public dataset is the authenticated 3,550-person role-A prefix used by the scratch GPT-2-medium A/B experiment. It contains 32 independently generated biography views per person (113,600 rows) and a separate canonical four-question QA bundle per person (14,200 rows). The biographies were generated with the pinned openai/gpt-oss-120b v37 workflow. The biographies configuration exposes split train; the qa configuration exposes split… See the full description on the dataset page: https://huggingface.co/datasets/jhdlee/distill-cl-biography-independent-ab-v4-gpt-oss-120b-v37-3550p32v-a.tabular100K<n<1M0 likes47 downloads28d agoHugging Face24rl-rag /open-scholar-gpt-oss-120b open-scholar-gpt-oss-120b Deep research agent evaluation on data/drtulu_open_scholar.jsonl (normal split). Results Metric Value pass@1 0.0% avg@1 0.0% Trajectory accuracy 0.0% (0/11854) Questions 11854 Trajectories 11854 (1 per question) Avg tool calls 17.3 Full conversations ✅ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains None Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/open-scholar-gpt-oss-120b.tabular10K<n<100K0 likes46 downloads7mo agoHugging Face25Kiria-Nozan /TRIM-gpt-oss-120b-separate-neighbors-only-para-random-feature-num TRIM Agent Reasoning Messages (HF Public Export) This directory is a Hugging Face-friendly public export of the TRIM agent reasoning SFT data. What Is Included Provider: vllm Model: gpt-oss-120b SFT mode: local_neighbor_only Splits present: train Records in this export manifest: 10056 Tasks in this split: AMES, BBB_Martins, Bioavailability_Ma, CYP2C9_Substrate_CarbonMangels, CYP2D6_Substrate_CarbonMangels, CYP3A4_Substrate_CarbonMangels, Carcinogens_Lagunin, ClinTox… See the full description on the dataset page: https://huggingface.co/datasets/Kiria-Nozan/TRIM-gpt-oss-120b-separate-neighbors-only-para-random-feature-num.tabular10K<n<100K0 likes46 downloads5mo agoHugging Face26jhdlee /distill-cl-biography-independent-ab-v4-gpt-oss-120b-v37-3550p32v-b Independent synthetic biographies, role B This public dataset is the authenticated 3,550-person role-B prefix used by the scratch GPT-2-medium A/B experiment. It contains 32 independently generated biography views per person (113,600 rows) and a separate canonical four-question QA bundle per person (14,200 rows). The biographies were generated with the pinned openai/gpt-oss-120b v37 workflow. The biographies configuration exposes split train; the qa configuration exposes split… See the full description on the dataset page: https://huggingface.co/datasets/jhdlee/distill-cl-biography-independent-ab-v4-gpt-oss-120b-v37-3550p32v-b.tabular100K<n<1M0 likes46 downloads28d agoHugging Face27rl-rag /gpt_oss_120b_sf_all_correcttabular10K<n<100K0 likes44 downloads6mo agoHugging Face28project-telos /gpt_oss_doorkey_action_distributions GPT-OSS-20B DoorKey action-distribution time series This dataset contains sentence-prefix next-action readouts for 46 fixed DoorKey environment states drawn from 31 trajectories in project-telos/trajectories_key_door_100. It also contains the corresponding offline BEAST change-point results. The dataset has 7,084 positions. Position 0 is the readout before any reasoning text is revealed. Each subsequent position reveals one additional reasoning sentence from the same model… See the full description on the dataset page: https://huggingface.co/datasets/project-telos/gpt_oss_doorkey_action_distributions.tabular10K<n<100K0 likes43 downloads1mo agoHugging Face29lvogel123 /jailbreak-gpt-oss-120b-hightabular1K<n<10K0 likes39 downloads11mo agoHugging Face30dvilasuero /bfcl-gpt-oss-20b-test bfcl-gpt-oss-20b-test Evaluation Results Eval created with evaljobs. This dataset contains evaluation results for the model(s) hf-inference-providers/openai/gpt-oss-20b:fastest using the eval inspect_evals/bfcl from Inspect Evals. To browse the results interactively, visit this Space. Command This eval was run with: evaljobs inspect_evals/bfcl \ --model hf-inference-providers/openai/gpt-oss-20b:fastest \ --name bfcl-gpt-oss-20b-test \ --limit 50… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/bfcl-gpt-oss-20b-test.tabularn<1K0 likes35 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.