datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-o1-verifiable-problem
Introduction
This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes.
For details, see our paper and GitHub repository.
Citation
If you find our data useful, please consider citing our work!
@misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.polaris_filtered_nemotron_medium_math_verifiable
Polaris Filtered Nemotron Medium Sympy Verifiable (v2)
This dataset is a curated subset of reasoning data from nvidia/Nemotron-Math-v2, specifically filtered for mathematical verifiability (verified using math verify-based equivalence), not having tool-reliance (TIR), and decontamination against the POLAIRS (POLARIS-Project/Polaris-Dataset-53K) dataset.
Dataset Summary
Total Original Samples: 2,424,392
Final Kept Samples: 357,790 (14.8%)
Target Reasoning Length: 4k-8k… See the full description on the dataset page: https://huggingface.co/datasets/devvrit/polaris_filtered_nemotron_medium_math_verifiable.polaris_filtered_nemotron_medium_sympy_verifiable
Polaris Filtered Nemotron Medium Sympy Verifiable
This dataset is a curated subset of reasoning data from nvidia/Nemotron-Math-v2, specifically filtered for mathematical verifiability (verified using sympy-based equivalence), not having tool-reliance (TIR), and decontamination against the POLAIRS (POLARIS-Project/Polaris-Dataset-53K) dataset.
Dataset Summary
Total Original Samples: 2,500,820
Final Kept Samples: 263,123 (10.5%)
Target Reasoning Length: Optimized for… See the full description on the dataset page: https://huggingface.co/datasets/devvrit/polaris_filtered_nemotron_medium_sympy_verifiable.Medical-o1-verifiable-problem-Thai
Introduction
This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes.
For details, see our paper and GitHub repository.
Citation
If you find our data useful, please consider citing our work!
@misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Medical-o1-verifiable-problem-Thai.verifiable-ai-provenance-bench
Verifiable AI Provenance Bench (TTTPS)
25 real timestamp-provenance receipts generated on 2026-08-04 by calling the
live KPP (Kenosian Protocol Platform) provenance API
(POST /v1/anchor, POST /v1/verify), which implements the TTTPS (Time-Token
Tamper-evident Provenance Seal) scheme. Each row is one real API round trip:
a content_hash was submitted to /v1/anchor, the returned receipt_id was then
submitted to /v1/verify, and both raw responses are recorded.
This dataset was built… See the full description on the dataset page: https://huggingface.co/datasets/Pittro/verifiable-ai-provenance-bench.verifiable-by-construction
Synthetic clinical question set for verbatim-citation evaluation
222 synthetic clinical questions, each written from one section of a published
cardiometabolic practice guideline.
Questions carry the identifier of their source section, not its text. The
guidelines are copyrighted and cannot be redistributed, so you need your own copy of
them. The harness rebuilds the section tree from them, and the labels resolve against
that tree.
Paper: arXiv:2609.15964
Code and harness:… See the full description on the dataset page: https://huggingface.co/datasets/oberst-lab/verifiable-by-construction.medical-o1-verifiable-problem-mk
Dataset Card for Dataset Name
This is a preview of a Macedonian translation of the medical-o1-verifiable-problem dataset by Freedom Intelligence.
Note that this preview currently contains 1068 rows.
Dataset Details
Dataset Structure
Each example consists of a question and a verifiable answer.
Dataset Creation
For methodological details regarding the creation of the original dataset, please refer to the original paper.
Machine translation was… See the full description on the dataset page: https://huggingface.co/datasets/ilijalichkovski/medical-o1-verifiable-problem-mk.cleaned_NuminaMath-RL-Verifiable_with_proofnumina-cot-verifiable-10kcleaned_NuminaMath-RL-Verifiableverifiable-coding-problems-python-prefopenr1-verifiable_8ktrain_verifiable
