datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.OpenO1-SFTThis repository contains the dataset used for fine-tuning a language model using SFT for Chain-of-Thought Activation from the paper Generative AI Act II: Test Time Scaling Drives Cognition Engineering.
Code: https://github.com/GAIR-NLP/cognition-engineering
🎉🎉🎉This repository contains the dataset used for fine-tuning a language model using SFT for Chain-of-Thought Activation.
🌈🌈🌈The dataset is designed to enhance the model's ability to generate coherent and logical reasoning sequences.… See the full description on the dataset page: https://huggingface.co/datasets/O1-OPEN/OpenO1-SFT.medical-o1-verifiable-problem
Introduction
This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes.
For details, see our paper and GitHub repository.
Citation
If you find our data useful, please consider citing our work!
@misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.Medical-o1-Reasoning-SFT-Thai
Medical-GPT-Reasoning-Thai
Dataset Summary
This dataset contains medical Q&A data in JSON format, designed for fine-tuning AI models in medical reasoning and response generation.representing a medical question, complex chain-of-thought reasoning, and a concise response. All content is in Thai language.
The dataset is derived from a larger medical Q&A collection and has been processed to ensure JSON validity, with multi-line objects combined into single valid entries.… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Medical-o1-Reasoning-SFT-Thai.Vietnamese-o1-journeyOriginal dataset: https://huggingface.co/datasets/GAIR/o1-journey
This dataset is a Vietnamese translated version of GAIR/o1-journey. Please cite the original dataset if you find it useful in your work.
Translated to Vietnamese with context-aware using gemini-flash-2.0-exp via this repo:
https://github.com/vTuanpham/Large_dataset_translator.
Roughly 2 hours for 649 examples.
Open_o1_sft_Pro_translated_jp
概要
このデータセットはOpen_o1_sft_ProデータセットをQwen社のQwen2.5-14B-Instructを用いて日本語に翻訳したものになります。
テンプレート
テンプレートは以下です。
{"conversations": [{"role": "user", "content": "入力"}, {"role": "assistant", "thought": "思考",
"content": "出力"}, ...],
"id": id(整数),
"dataset": "元データセットの名前"}
ライセンス
ライセンスは元データセットに準じます。
謝辞
データセットの製作者様,Qwenの開発者様,計算資源を貸してくださったVolt mindの皆様に感謝します。
medical-o1-reasoning-SFT-TH
Translated medical-o1-reasoning-SFT (TH)
Overview
This repository contains a translated version of the medical-o1-reasoning-SFT.
Changes in This Version
Full translation of all Question, Complex_CoT, and Response fields from English to Thai.
Preserved original dataset structure.
Using Gemini 2.0 Flash for translation.
Some data has been excluded due to translation limitations.
Original Dataset Source
The original dataset was released under an… See the full description on the dataset page: https://huggingface.co/datasets/RJTPP/medical-o1-reasoning-SFT-TH.Diabetica-o1-SFT
Diabetica-o1-SFT
Diabetica: Adapting Large Language Model to Enhance Multiple Medical Tasks in Diabetes Care and Management
Code |Paper
Introduction
Specifically, we use Deepseek-R1-Distilled-Qwen-32B as our teacher model. Our data augmentation strategy follows a two-step approach: (1) We prompt Qwen2.5-72B-Instruct to generate diverse synthetic questions based on existing datasets. (2) We then use Deepseek-R1-Distilled-Qwen-32B to generate responses for… See the full description on the dataset page: https://huggingface.co/datasets/WaltonFuture/Diabetica-o1-SFT.medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/Tapashpaul029/medical-o1-reasoning-SFT.Medical-o1-verifiable-problem-Thai
Introduction
This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes.
For details, see our paper and GitHub repository.
Citation
If you find our data useful, please consider citing our work!
@misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/Medical-o1-verifiable-problem-Thai.medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an LLM… See the full description on the dataset page: https://huggingface.co/datasets/dolbi14/medical-o1-reasoning-SFT.medical-o1-reasoning-SFT
简介
该数据集用于微调 HuatuoGPT-o1,这是一个专为高级医疗推理设计的大语言模型。该数据集是使用 GPT-4o 构建的,它搜索可验证的医疗问题的解决方案,并通过医疗验证器进行验证。
更多详细信息,请参阅我们的论文和GitHub 仓库。
引用
如果您觉得我们的数据对您有帮助,请考虑引用我们的工作!
@misc{chen2024huatuogpto1medicalcomplexreasoning,
title={HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs},
author={Junying Chen and Zhenyang Cai and Ke Ji and Xidong Wang and Wanlong Liu and Rongsheng Wang and Jianye Hou and Benyou Wang},
year={2024},
eprint={2412.18925}… See the full description on the dataset page: https://huggingface.co/datasets/toimc/medical-o1-reasoning-SFT.medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an LLM… See the full description on the dataset page: https://huggingface.co/datasets/Hellrabbit/medical-o1-reasoning-SFT.medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an LLM… See the full description on the dataset page: https://huggingface.co/datasets/liupc/medical-o1-reasoning-SFT.medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an LLM… See the full description on the dataset page: https://huggingface.co/datasets/BlueGod01/medical-o1-reasoning-SFT.medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/thongfamilynguyen1126/medical-o1-reasoning-SFT.medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an LLM… See the full description on the dataset page: https://huggingface.co/datasets/TianchenTang/medical-o1-reasoning-SFT.medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an LLM… See the full description on the dataset page: https://huggingface.co/datasets/vBasura/medical-o1-reasoning-SFT.medical-o1-verifiable-problem-mk
Dataset Card for Dataset Name
This is a preview of a Macedonian translation of the medical-o1-verifiable-problem dataset by Freedom Intelligence.
Note that this preview currently contains 1068 rows.
Dataset Details
Dataset Structure
Each example consists of a question and a verifiable answer.
Dataset Creation
For methodological details regarding the creation of the original dataset, please refer to the original paper.
Machine translation was… See the full description on the dataset page: https://huggingface.co/datasets/ilijalichkovski/medical-o1-verifiable-problem-mk.medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an LLM… See the full description on the dataset page: https://huggingface.co/datasets/Fadil369/medical-o1-reasoning-SFT.medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an LLM… See the full description on the dataset page: https://huggingface.co/datasets/CrossNow/medical-o1-reasoning-SFT.Pashto-Medical-o1-Reasoning-SFT-Dataset
Pashto Medical o1 Reasoning SFT Dataset
This dataset provides medical instruction-tuning data featuring chain-of-thought (CoT) reasoning steps in Pashto, structured for Supervised Fine-Tuning (SFT) of large language models.
Dataset Structure
The dataset contains conversational message formats with step-by-step reasoning encapsulated via <think> blocks, followed by the final expert medical response.
Data Fields
Question: The medical question or… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Medical-o1-Reasoning-SFT-Dataset.
