CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01csoai /gspc-oss GSPC — openness bank (OSSBench) Council of AI measurement bank. Measurement, not certification. Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026. Live measurement. This bank stands behind the openness row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=openness (family, kind, status and n are on that row, never typed here;… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-oss.tabularquestion-answeringn<1K0 likes735 downloads4d agoHugging Face02Jackrong /gpt-oss-120b-reasoning-STEM-5K GPT-OSS-120B-Distilled-Reasoning-STEM Dataset 1) Dataset Overview Data Source Model: gpt-oss-120b-high Task Type: STEM Reasoning and Problem Solving (Science, Technology, Engineering & Mathematics) Data Format: `JSON Lines Fields: generator, category, input, CoT_Native——reasoning, answer (Consistent with the math dataset, splitting the original 'output' into 'reasoning' and 'answer' for COT/SFT scenarios.) 2) Design Goals (Motivation) This dataset targets… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120b-reasoning-STEM-5K.textquestion-answering1K<n<10K12 likes250 downloads1y agoHugging Face03OpenMed /Medical-Reasoning-SFT-GPT-OSS-120B-V2 Medical-Reasoning-SFT-GPT-OSS-120B-V2 A large-scale medical reasoning dataset generated using openai/gpt-oss-120b, containing over 506,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. GPT-OSS-120B is OpenAI's state-of-the-art open-weight model, achieving near-parity with closed models on reasoning benchmarks while being Apache 2.0 licensed. Dataset Overview Metric Value Model openai/gpt-oss-120b Total Samples 506… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-GPT-OSS-120B-V2.texttext-generation100K<n<1M9 likes222 downloads8mo agoHugging Face04Tonic /Health-Bench-Eval-OSS-2025-07 Dataset Card for HealthBench Dataset Summary HealthBench is a benchmark dataset developed by OpenAI in collaboration with 262 physicians from 60 countries to evaluate AI systems in health-related conversational scenarios. It contains 5,000 multi-turn health conversations in a JSONL file (2025-05-07-06-14-12_oss_eval.jsonl), simulating interactions between AI models and users (laypersons or clinicians). Each conversation includes a user prompt, a candidate model response… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/Health-Bench-Eval-OSS-2025-07.texttext-generation1K<n<10K4 likes151 downloads1y agoHugging Face05Hugodonotexit /Superior-Reasoning-SFT-gpt-oss-120b-split-en Superior-Reasoning SFT (stage1 + stage2) with <think> split and English filtering Summary This dataset is a processed derivative of Alibaba-Apsara/Superior-Reasoning-SFT-gpt-oss-120b (subsets stage1 and stage2, train split). It restructures each example into three fields: input: the original input reasoning: the content extracted from <think> ... </think> within the original output (inner text only) output: the remainder of the original output after removing all <think>… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/Superior-Reasoning-SFT-gpt-oss-120b-split-en.texttext-generation100K<n<1M2 likes113 downloads8mo agoHugging Face06Jackrong /gpt-oss-120B-distilled-reasoning GPT-oss-120B-Distilled-Reasoning-math Dataset Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines Fields: Generator, Category, Input, Output Core Statistics Generated complete reasoning processes and answers using gpt-oss-120b (MXFP4).The text length of the dataset reflects the depth and complexity of its content. I have statistically analyzed the lengths of the input (question), Reasoning, and Answer.To understand the data… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120B-distilled-reasoning.texttext-classification1K<n<10K20 likes87 downloads1y agoHugging Face07Jackrong /GPT-OSS-120B-Distilled-Reasoning-math GPT-oss-120B-Distilled-Reasoning-math Dataset Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines Fields: Generator, Category, Input, CoT_Native_Reasoning, Reasoning, Answer Core Statistics Generated complete reasoning processes and answers using gpt-oss-120b (MXFP4).The text length of the dataset reflects the depth and complexity of its content. I have statistically analyzed the lengths of the input (question), Reasoning, and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-120B-Distilled-Reasoning-math.textquestion-answering1K<n<10K9 likes78 downloads1y agoHugging Face08zyx1234 /MuSeR_GPT_OSS_120B_DistillationThis dataset contains ~100k synthetic medical queries and corresponding responses distilled from GPT-OSS-120B. The generation of synthetic medical queries follows an attribute-conditioned generation method proposed in paper Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning. We found that supervised fine-tuning on this dataset can substantially improve LLMs' medical conversational capabilities. See our paper and project page for more details. If… See the full description on the dataset page: https://huggingface.co/datasets/zyx1234/MuSeR_GPT_OSS_120B_Distillation.textquestion-answering10K<n<100K4 likes58 downloads9mo agoHugging Face09Jackrong /GPT-OSS-20B-Distilled-Reasoning-Mini Dataset Card for Dataset Name GPT-OSS-20B Distilled Reasoning Dataset Mini (Multi-stage Evaluative Refinement Method for Reasoning Generation) Dataset Details and Description This is a high-quality instruction fine-tuning dataset constructed through knowledge distillation, featuring detailed Chain-of-Thought (CoT) reasoning processes. The dataset is designed to enhance the capabilities of smaller language models in complex reasoning, logical analysis, and instruction… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/GPT-OSS-20B-Distilled-Reasoning-Mini.tabulartext-classification1K<n<10K21 likes55 downloads1y agoHugging Face10emgena /omnimcp_eu_vat_oss_calculator_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_eu_vat_oss_calculator_teaser.texttext-generationn<1K0 likes52 downloads9d agoHugging Face11Jackrong /gpt-oss-120B-distilled-math-OpenAI-Harmony 📚 Dataset Overview Data Source Model: gpt-oss-120bTask Type: Mathematical Problem SolvingData Format: JSON Lines (.jsonl)Fields: Generator, Category, Input, Output Note: If you are using this template for training, please make sure the format is correct before starting.Since this template is still under continuous improvement and learning, it may not be fully complete yet. I appreciate your understanding. 📈 Core Statistics Generated complete reasoning processes… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/gpt-oss-120B-distilled-math-OpenAI-Harmony.texttext-classification1K<n<10K6 likes46 downloads1y agoHugging Face12MapleBi /MetaRAG_Cross-Issue_OSSQA MetaRAG Cross-Issue OSSQA Dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA MetaRAG Cross-Issue OSSQA is an English open-source software issue question-answering and retrieval benchmark. Each example asks a question grounded in one GitHub issue and requires evidence from a related issue. The data contains explicit cross-issue references and a three-document silver evidence path. Dataset configurations Configuration Splits Rows… See the full description on the dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA.tabularquestion-answering10K<n<100K0 likes44 downloads1mo agoHugging Face13wanglab /eurorad-gpt-oss-training-data Benchmarking and Adapting On-Device Large Language Models for Clinical Decision Support Authors Alif Munim* 1, Jun Ma* 1,2, Omar Ibrahim* 1, Alhusain Abdalla* 1, Shuolin Yin3, Leo Chen4, Bo Wang† 1,5,6,7,8 * Equal contribution     † Corresponding author 1AI Collaborative Centre, University Health Network, Toronto, Canada 2Princess Margaret Cancer Centre, University Health Network, Toronto, Canada 3Department of… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/eurorad-gpt-oss-training-data.texttext-generation1K<n<10K2 likes43 downloads7mo agoHugging Face14LeBrony /buddha_oss_dataset 장아함경 Buddha QA Dataset (Complete) / Agama Sutra Buddha QA Dataset 한국어 설명 | English Description 한국어 🙏 개요 이 데이터셋은 장아함경(長阿含經) 제1-3권 전체를 기반으로 생성된 한국어 불교 질문-답변 데이터셋입니다. GPT-4.1-2025-04-14 모델과 고급 추론 시스템(fill_thinking.py)을 사용하여 생성되었으며, 현대인이 이해하기 쉽도록 해석된 붓다의 가르침을 포함합니다. 📊 데이터셋 통계 총 QA 쌍: 335개 훈련 데이터: 268개 (80%) 검증 데이터: 67개 (20%) 출처: 장아함경 제1-3권 (총 70개 경전) 생성 모델: GPT-4.1-2025-04-14 추론 시스템: fill_thinking.py 기반 🔧 데이터 형식 각 레코드는 다음과 같은 3-message… See the full description on the dataset page: https://huggingface.co/datasets/LeBrony/buddha_oss_dataset.textquestion-answeringn<1K0 likes34 downloads1y agoHugging Face15omareng /eurorad-gpt-oss-training-data Eurorad Medical Radiology Training Dataset with GPT-OSS 120B Reasoning Training dataset used for fine-tuning GPT-OSS 20B for medical radiology diagnosis tasks. Dataset Description This dataset contains 1,894 medical radiology cases from Eurorad, each enhanced with detailed diagnostic reasoning generated by GPT-OSS 120B. The dataset was used to train the model available at omareng/on-device-LLM-gpt-oss-20b. Dataset Structure Each row contains: case_id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/omareng/eurorad-gpt-oss-training-data.texttext-generation1K<n<10K0 likes33 downloads10mo agoHugging Face16Ericwang /gpt-oss-distilled-redteam2k GPT-OSS Distilled RedTeam-2K Dataset This is a preliminary experimental subset of a larger dataset. For the full dataset and additional information, see: Nemotron Nano 2 Safety Distill — GPT-OSS . ⚠️ Content Warning: This dataset contains potentially harmful or policy-violating prompts (e.g., animal abuse, violence, privacy violations). The content includes sensitive safety-related queries and should be used responsibly for research purposes only. Overview This… See the full description on the dataset page: https://huggingface.co/datasets/Ericwang/gpt-oss-distilled-redteam2k.texttext-generation1K<n<10K1 likes32 downloads11mo agoHugging Face17IIGroup /gpt-oss-benchmark-responses gpt-oss-20b Benchmark Responses Dataset Overview This dataset contains responses generated by the gpt-oss-20b model on multiple benchmark tests, showcasing its performance in mathematical reasoning, language understanding, and cross-domain knowledge tasks. All responses are generated with a maximum length of 16K tokens. The included benchmarks are: (TODO) HLE (Humanity's Last Exam): A multimodal benchmark with 2,500 multiple-choice and short-answer questions spanning… See the full description on the dataset page: https://huggingface.co/datasets/IIGroup/gpt-oss-benchmark-responses.textquestion-answeringn<1K8 likes30 downloads1y agoHugging Face18IIGroup /s1K-1.1-gpt-oss-20b s1K-1.1-gpt-oss-20b Dataset Summary The s1K-1.1-gpt-oss-20b dataset extends the simplescaling/s1K-1.1 dataset by incorporating reasoning trajectories generated by the openai/gpt-oss-20b model. This dataset contains questions primarily from mathematical problem-solving domains, along with responses and reasoning trajectories generated by the gpt-oss-20b model. The dataset is designed to facilitate research into model reasoning capabilities, test-time scaling, and… See the full description on the dataset page: https://huggingface.co/datasets/IIGroup/s1K-1.1-gpt-oss-20b.textquestion-answering1K<n<10K3 likes30 downloads1y agoHugging Face19cybertruck32489 /gpt-oss-reasoning-ru-mediumНовая версия датасета на 25к строк textquestion-answering10K<n<100K1 likes21 downloads1y agoHugging Face20OctoMed /GPT-OSS-120B-Reasoning OctoMed/GPT-OSS-120B-Reasoning Single-turn instruction-following examples with explicit chain-of-thought reasoning, converted to OctoMed format for SFT training. Source Derived from Jackrong/gpt-oss-120b-Reasoning-Instruction by Jackrong. All credit for the original data collection and model distillation goes to the original authors. Format Each example contains: question: the instruction / question text answer: the final answer extracted after reasoning… See the full description on the dataset page: https://huggingface.co/datasets/OctoMed/GPT-OSS-120B-Reasoning.textquestion-answering10K<n<100K0 likes20 downloads5mo agoHugging Face21introvoyz041 /Medical-Reasoning-SFT-GPT-OSS-120B-V2 Medical-Reasoning-SFT-GPT-OSS-120B-V2 A large-scale medical reasoning dataset generated using openai/gpt-oss-120b, containing over 506,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. GPT-OSS-120B is OpenAI's state-of-the-art open-weight model, achieving near-parity with closed models on reasoning benchmarks while being Apache 2.0 licensed. Dataset Overview Metric Value Model openai/gpt-oss-120b Total Samples 506… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/Medical-Reasoning-SFT-GPT-OSS-120B-V2.texttext-generation100K<n<1M1 likes14 downloads5mo agoHugging Face22cybertruck32489 /gpt-oss-reasoning-ru-nanoВторая версия датасета на 1000 примеров. textquestion-answering1K<n<10K0 likes13 downloads1y agoHugging Face23Hugodonotexit /filtered-awesome-chatgpt-propmts-oss-120b Filtered Awesome ChatGPT Prompts – Model Outputs Dataset Overview This dataset contains model-generated responses to prompts from the fka/awesome-chatgpt-prompts Hugging Face dataset. Each prompt was sent to the openai/gpt-oss-120b model via the OpenRouter API. The resulting dataset was then filtered to remove: Non English outputs with high language-detection confidence (fastText score < 0.7) Very short outputs (≤ 10 words) The goal of this dataset is to provide a… See the full description on the dataset page: https://huggingface.co/datasets/Hugodonotexit/filtered-awesome-chatgpt-propmts-oss-120b.texttext-generationn<1K1 likes9 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.