datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.CMB
CMB: A Comprehensive Medical Benchmark in Chinese
🌐 Github • 🌐 Website • 🤗 HuggingFace
🌈 Update
[2024.02.21] The answers to the CMB-Exam test has been updated and some errors caused by omissions in version management have been fixed.
[2024.01.08] In order to facilitate testing, we disclose the answers to the CMB-Exam test
[2023.09.22] CMB is included in OpenCompass.
[2023.08.21] Paper released.
[2023.08.01] 🎉🎉🎉 CMB is published!🎉🎉🎉
🌐… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/CMB.MileBench
MileBench
Introduction
We introduce MileBench, a pioneering benchmark designed to test the MultImodal Long-contExt capabilities of MLLMs.
This benchmark comprises not only multimodal long contexts, but also multiple tasks requiring both comprehension and generation.
We establish two distinct evaluation sets, diagnostic and realistic, to systematically assess MLLMs’ long-context adaptation capacity and their ability to completetasks in long-context scenarios
To… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/MileBench.freebase
Freebase
Dataset Description
Large-scale knowledge base (archived by Google)
Original Source: http://commondatastorage.googleapis.com/freebase-public/rdf/freebase-rdf-latest.gz
Dataset Summary
This dataset contains RDF triples from Freebase converted to HuggingFace
dataset format for easy use in machine learning pipelines.
Format: Originally ntriples, converted to HuggingFace Dataset
Size: 300.0 GB (extracted)
Entities: ~50M
Triples: ~3B
Original License:
CC… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/freebase.PubMedVision
News
[2025/02/18]: We add the original captions of PubMedVision in PubMedVision_Original_Caption.json, as well as the Chinese version of PubMedVision in PubMedVision_Chinese.json.
[2024/07/01]: We add annotations for 'body_part' and 'modality' of images, utilizing the HuatuoGPT-Vision-7B model.
PubMedVision
PubMedVision is a large-scale medical VQA dataset. We extracted high-quality image-text pairs from PubMed and used GPT-4V to reformat them to enhance their quality.… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/PubMedVision.Medical-R1-Distill-Data
Introduction
This dataset is an SFT dataset distilled from Deepseek-R1 (Full Power Version), based on medical verifiable problems from HuatuoGPT-o1.
The Chinese version of the dataset is available at FreedomIntelligence/Medical-R1-Distill-Data-Chinese.
The distillation originates from the native Deepseek-R1 API requests. We hope this distilled dataset can help initialize your models with the reasoning chain from R1. You can also use our previously built medical verified long… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data.huatuo_encyclopedia_qa
Dataset Card for Huatuo_encyclopedia_qa
Dataset Summary
This dataset has a total of 364,420 pieces of medical QA data, some of which have multiple questions in different ways. We extract medical QA pairs from plain texts (e.g., medical encyclopedias and medical articles). We collected 8,699 encyclopedia entries for diseases and 2,736 encyclopedia entries for medicines on Chinese Wikipedia. Moreover, we crawled 226,432 high-quality medical articles from the Qianwen Health… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_encyclopedia_qa.ALLaVA-4V
📚 ALLaVA-4V Data
Generation Pipeline
LAION
We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here.
Vison-FLAN
We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here.
Wizard
We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo.
Dataset Cards
All datasets can be found here.
The structure of naming is shown below:
ALLaVA-4V
├──… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V.medical-o1-verifiable-problem
Introduction
This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes.
For details, see our paper and GitHub repository.
Citation
If you find our data useful, please consider citing our work!
@misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.Freebase
Freebase
Dataset Description
Large-scale knowledge base (archived by Google)
Original Source: http://commondatastorage.googleapis.com/freebase-public/rdf/freebase-rdf-latest.gz
Dataset Summary
This dataset contains RDF triples from Freebase converted to HuggingFace
dataset format for easy use in machine learning pipelines.
Format: Originally ntriples, converted to HuggingFace Dataset
Size: 300.0 GB (extracted)
Entities: ~50M
Triples: ~3B
Original License:
CC… See the full description on the dataset page: https://huggingface.co/datasets/Yariz/Freebase.TCM-Pretrain-Data-ShizhenGPT
📚 Introduction
This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Pretrain-Data-ShizhenGPT.Huatuo26M-Lite
Huatuo26M-Lite 📚
Table of Contents 🗂
Dataset Description 📝
Dataset Information ℹ️
Data Distribution 📊
Usage 🔧
Citation 📖
Dataset Description 📝
Huatuo26M-Lite is a refined and optimized dataset based on the Huatuo26M dataset, which has undergone multiple purification processes and rewrites. It has more data dimensions and higher data quality. We welcome you to try using it.
Dataset Information ℹ️
Dataset Name: Huatuo26M-Lite
Version:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Huatuo26M-Lite.swe-zero-free
SWE-ZERO-Free, 12M agentic coding trajectories made without an LLM
SWE-ZERO-12M showed you can generate agentic SWE data without Docker, just by sticking to shell commands that need no setup. We took that idea and asked whether you even need the model. You don't. So this is SWE-ZERO, but free.
This dataset has 12,238,610 mini-swe-agent trajectories across 119,084 real GitHub PRs. Same source and same format as SWE-ZERO. The only difference is how they get made. SWE-ZERO samples… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/swe-zero-free.huatuo_consultation_qa
Dataset Card for huatuo_consultation_qa
Dataset Summary
We collected data from a website for medical consultation , consisting of many online consultation records by medical experts. Each record is a QA pair: a patient raises a question and a medical doctor answers the question. The basic information of doctors (including name, hospital organization, and department) was recorded.
We directly crawl patient’s questions and doctor’s answers as QA pairs, getting 32,708,346… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_consultation_qa.huatuo_knowledge_graph_qa
Dataset Card for Huatuo_knowledge_graph_qa
Dataset Summary
We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map.
Dataset Creation
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.TCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.HuatuoGPT2-SFT-GPT4-140K
HuatuoGPT2-SFT-GPT4-140K
140K Chinese medical instructions generated by GPT-4, based on questions from HuatuoGPT Dataset.
This dataset contains supervised fine-tuning instructions for HuatuoGPT2, designed to enhance the model's ability to follow instructions in real medical scenarios. We have made all the data (142,248 entries) in this dataset publicly available.
Repository
Github: https://github.com/FreedomIntelligence/HuatuoGPT-II
Citation… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-SFT-GPT4-140K.huatuo26M-testdatasets
Dataset Card for huatuo26M-testdatasets
Dataset Summary
We are pleased to announce the release of our evaluation dataset, a subset of the Huatuo-26M. This dataset contains 6,000 entries that we used for Natural Language Generation (NLG) experimentation in our associated research paper.
We encourage researchers and developers to use this evaluation dataset to gauge the performance of their own models. This is not only a chance to assess the accuracy and relevancy of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo26M-testdatasets.Caselaw_Access_Project
The Caselaw Access Project
In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/
Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here:… See the full description on the dataset page: https://huggingface.co/datasets/free-law/Caselaw_Access_Project.Evol-Instruct-Chinese-GPT4The dataset is created by (1) translating English questions of Evol-instruct-70k into Chinese and (2) requesting GPT4 to generate Chinese responses.
For more details, please refer to:
Repository:
https://github.com/FreedomIntelligence/AceGPT
https://github.com/FreedomIntelligence/LLMZoo
Paper:
AceGPT, Localizing Large Language Models in Arabic
Phoenix: Democratizing ChatGPT across Languages
BibTeX entry and citation info
@article{huang2023acegpt,
title={AceGPT, Localizing… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Evol-Instruct-Chinese-GPT4.Caselaw_Access_Project_embeddings
The Caselaw Access Project
In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/
Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here:… See the full description on the dataset page: https://huggingface.co/datasets/free-law/Caselaw_Access_Project_embeddings.Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1
Dataset Description:
Teaches the model to follow arbitrary text formatting instructions (bullet styles, numbering, delimiters, heading formats, inline emphasis, web-answer structure, etc.) for targeted chat behaviors. Uses explicit Regex and string matching for the reward signal.
This dataset is ready for commercial or non-commercial uses.
Dataset Owner(s):
NVIDIA Corporation
Dataset Creation Date:
Created on: April 10, 2026
Last Modified on: April… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1.Pashto-Free-Hand-Reasoning-Dataset
Pashto Free-Hand Reasoning SFT Dataset 🧠♻️
This dataset contains high-quality, long-form SFT (Supervised Fine-Tuning) conversational data in Pashto, featuring unconstrained, natural model reasoning (<think> blocks) paired with standardized chat responses.
🔄 The 3R Approach (Recycle, Reuse, Reason)
Instead of discarding legacy QA pairs, this dataset follows a 3R data philosophy:
Recycle: Taking older, simple, or raw legacy Pashto questions.
Reuse: Re-processing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Free-Hand-Reasoning-Dataset.RAG-Instruct
Introduction
RAG-Instruct is a RAG dataset designed to comprehensively enhance LLM RAG capabilities, synthesized using GPT-4o. This dataset is based on the Wikipedia corpus and This dataset is based on the Wikipedia corpus and offers the advantages of query-document scenario diversity and task diversity.
The RAG-Instruct dataset can significantly enhance the RAG ability of LLMs and make remarkable improvements in RAG performance across various tasks.
Model
WQA (acc)
PQA (acc)… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/RAG-Instruct.HuatuoGPT2-Pretraining-Instruction
HuatuoGPT2-Pretraining-Instruction-5200K
Here are the pre-training instructions for HuatuoGPT-II, developed with 5.2 million medical corpus using ChatGPT.
This dataset is used to incorporate extensive medical knowledge and enable a one-stage medical adaptation. All our data have been made publicly accessible.
Data Volume
The following table details the volume and distribution of pre-training data for HuatuoGPT2:
Data Source
Data Volume
Medical_Web_Corpus_cn… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-Pretraining-Instruction.swe-zero-free-v2
SWE-ZERO-Free, 12M agentic coding trajectories made without an LLM
SWE-ZERO-12M showed you can generate agentic SWE data without Docker, just by sticking to shell commands that need no setup. We took that idea and asked whether you even need the model. You don't. So this is SWE-ZERO, but free.
This dataset has 12,177,154 mini-swe-agent trajectories across ~121,800 real GitHub PRs. Same source and same format as SWE-ZERO. The only difference is how they get made. SWE-ZERO samples… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/swe-zero-free-v2.Medical-R1-Distill-Data-Chinese
Introduction
This dataset is an SFT dataset distilled from Deepseek-R1 (Full Power Version), based on Chinese medical verifiable problems from HuatuoGPT-o1.
The distillation originates from the native Deepseek-R1 API requests. We hope this distilled dataset can help initialize your models with the reasoning chain from R1. You can also use our previously built medical verified long reasoning chains based on GPT-4o on medical-o1-reasoning-SFT.
For details, see our paper and GitHub… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data-Chinese.ScalpelBench
ScalpelBench
ScalpelBench is a compact instruction-tuning corpus developed for controlled
studies of model compression, with a particular focus on layer pruning,
post-pruning recovery, and capability retention. The released corpus contains
approximately 0.1B tokens of instruction-response data spanning general
English, Chinese, mathematical reasoning, and code generation.
Mixture Design
The mixture proportions follow high-level capability-balancing principles… See the full description on the dataset page: https://huggingface.co/datasets/freeai-org/ScalpelBench.monitorability-as-a-free-gift-data-training-data
Monitorability as a free gift training data reordered, or normalized during packaging.
Configurations
Config
Rows
Purpose
Original file
all
18,591
Main all-domain experiment
combined_dataset.parquet
no_if
13,591
All-domain experiment without instruction following
combined_dataset_noif.parquet
instruction_following
5,000
Instruction-following experiments
instruction_following_ai2_5000.parquet
math
5,000
Main math experiments
skywork_math.parquet… See the full description on the dataset page: https://huggingface.co/datasets/polaris-73/monitorability-as-a-free-gift-data-training-data.OnePO-Medical-20K
OnePO-Medical-20K
📄 Paper |
💻 GitHub
⚡ Introduction
OnePO-Medical-20K is the medical RL dataset released with OnePO, containing 20,338 medical tasks across multiple languages.
One stage, no preceding SFT. OnePO adapts pretrained models to medicine through a single reinforcement-learning stage.
Two complementary task types. Multiple-choice questions provide verifiable answers. Open-ended conversations provide scoring rubrics.
Teacher guidance included. Each task includes a… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/OnePO-Medical-20K.
