datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.CMB
CMB: A Comprehensive Medical Benchmark in Chinese
🌐 Github • 🌐 Website • 🤗 HuggingFace
🌈 Update
[2024.02.21] The answers to the CMB-Exam test has been updated and some errors caused by omissions in version management have been fixed.
[2024.01.08] In order to facilitate testing, we disclose the answers to the CMB-Exam test
[2023.09.22] CMB is included in OpenCompass.
[2023.08.21] Paper released.
[2023.08.01] 🎉🎉🎉 CMB is published!🎉🎉🎉
🌐… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/CMB.MileBench
MileBench
Introduction
We introduce MileBench, a pioneering benchmark designed to test the MultImodal Long-contExt capabilities of MLLMs.
This benchmark comprises not only multimodal long contexts, but also multiple tasks requiring both comprehension and generation.
We establish two distinct evaluation sets, diagnostic and realistic, to systematically assess MLLMs’ long-context adaptation capacity and their ability to completetasks in long-context scenarios
To… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/MileBench.PubMedVision
News
[2025/02/18]: We add the original captions of PubMedVision in PubMedVision_Original_Caption.json, as well as the Chinese version of PubMedVision in PubMedVision_Chinese.json.
[2024/07/01]: We add annotations for 'body_part' and 'modality' of images, utilizing the HuatuoGPT-Vision-7B model.
PubMedVision
PubMedVision is a large-scale medical VQA dataset. We extracted high-quality image-text pairs from PubMed and used GPT-4V to reformat them to enhance their quality.… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/PubMedVision.Medical-R1-Distill-Data
Introduction
This dataset is an SFT dataset distilled from Deepseek-R1 (Full Power Version), based on medical verifiable problems from HuatuoGPT-o1.
The Chinese version of the dataset is available at FreedomIntelligence/Medical-R1-Distill-Data-Chinese.
The distillation originates from the native Deepseek-R1 API requests. We hope this distilled dataset can help initialize your models with the reasoning chain from R1. You can also use our previously built medical verified long… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data.ALLaVA-4V
📚 ALLaVA-4V Data
Generation Pipeline
LAION
We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here.
Vison-FLAN
We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here.
Wizard
We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo.
Dataset Cards
All datasets can be found here.
The structure of naming is shown below:
ALLaVA-4V
├──… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V.medical-o1-verifiable-problem
Introduction
This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes.
For details, see our paper and GitHub repository.
Citation
If you find our data useful, please consider citing our work!
@misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.Huatuo26M-Lite
Huatuo26M-Lite 📚
Table of Contents 🗂
Dataset Description 📝
Dataset Information ℹ️
Data Distribution 📊
Usage 🔧
Citation 📖
Dataset Description 📝
Huatuo26M-Lite is a refined and optimized dataset based on the Huatuo26M dataset, which has undergone multiple purification processes and rewrites. It has more data dimensions and higher data quality. We welcome you to try using it.
Dataset Information ℹ️
Dataset Name: Huatuo26M-Lite
Version:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Huatuo26M-Lite.freebase_qaFreebaseQA is for open-domain factoid question answering (QA) tasks over structured knowledge bases, like Freebase The data set is generated by matching trivia-type question-answer pairs with subject-predicateobject triples in Freebase.TCM-Instruction-Tuning-ShizhenGPT
📚 Introduction
This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM.
For details, see our paper and GitHub repository.
📊 Dataset Overview
The open-sourced fine-tuning dataset consists of three parts:
Modality
Data Quantity
TCM Text Instructions
📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.HuatuoGPT2-SFT-GPT4-140K
HuatuoGPT2-SFT-GPT4-140K
140K Chinese medical instructions generated by GPT-4, based on questions from HuatuoGPT Dataset.
This dataset contains supervised fine-tuning instructions for HuatuoGPT2, designed to enhance the model's ability to follow instructions in real medical scenarios. We have made all the data (142,248 entries) in this dataset publicly available.
Repository
Github: https://github.com/FreedomIntelligence/HuatuoGPT-II
Citation… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-SFT-GPT4-140K.ApolloMoEDataset
Democratizing Medical LLMs For Much More Languages
Covering 12 Major Languages including English, Chinese, French, Hindi, Spanish, Arabic, Russian, Japanese, Korean, German, Italian, Portuguese and 38 Minor Languages So far.
📃 Paper • 🌐 Demo • 🤗 ApolloMoEDataset • 🤗 ApolloMoEBench • 🤗 Models •🌐 Apollo • 🌐 ApolloMoE
🌈 Update
[2024.10.15] ApolloMoE repo is published!🎉
Languages Coverage
12 Major Languages and 38 Minor Languages
Click to… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ApolloMoEDataset.Medical_Multimodal_Evaluation_Data
Evaluation Guide
This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks.
To get started:
Download the dataset and extract the images.zip file.
Find evaluation code on our GitHub: HuatuoGPT-Vision.
This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.EchoX-Dialougues
EchoX-Dialogues: Training Data for EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs
🐈⬛ Github | 📃 Paper | 🚀 Space
🧠 EchoX-8B | 🧠 EchoX-3B | 📦 EchoX-Dialogues-Plus
EchoX-Dialogues provides the primary speech dialogue data used to train EchoX, restricted to S2T (speech → text) in this repository.
All input speech is synthetic; text is derived from public sources with multi-stage cleaning and rewriting. Most turns include asr /… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/EchoX-Dialougues.GPQA-diamond-freeRAG-Instruct
Introduction
RAG-Instruct is a RAG dataset designed to comprehensively enhance LLM RAG capabilities, synthesized using GPT-4o. This dataset is based on the Wikipedia corpus and This dataset is based on the Wikipedia corpus and offers the advantages of query-document scenario diversity and task diversity.
The RAG-Instruct dataset can significantly enhance the RAG ability of LLMs and make remarkable improvements in RAG performance across various tasks.
Model
WQA (acc)
PQA (acc)… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/RAG-Instruct.HuatuoGPT2-Pretraining-Instruction
HuatuoGPT2-Pretraining-Instruction-5200K
Here are the pre-training instructions for HuatuoGPT-II, developed with 5.2 million medical corpus using ChatGPT.
This dataset is used to incorporate extensive medical knowledge and enable a one-stage medical adaptation. All our data have been made publicly accessible.
Data Volume
The following table details the volume and distribution of pre-training data for HuatuoGPT2:
Data Source
Data Volume
Medical_Web_Corpus_cn… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-Pretraining-Instruction.ApolloMoEBench
Democratizing Medical LLMs For Much More Languages
Covering 12 Major Languages including English, Chinese, French, Hindi, Spanish, Arabic, Russian, Japanese, Korean, German, Italian, Portuguese and 38 Minor Languages So far.
📃 Paper • 🌐 Demo • 🤗 ApolloMoEDataset • 🤗 ApolloMoEBench • 🤗 Models •🌐 Apollo • 🌐 ApolloMoE
🌈 Update
[2024.10.15] ApolloMoE repo is published!🎉
Languages Coverage
12 Major Languages and 38 Minor Languages
Click to… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ApolloMoEBench.HiMed
HiMed
HiMed is a Hindi medical dataset and benchmark suite covering both Western medicine and Indian systems of medicine.It consists of two parts:
HiMed-Trad: traditional Indian medicine
HiMed-West: Western medicine under Hindi prompts
Repository Layout
All released files are under data/:
data/
├── HiMed-Trad_Bench.json
├── HiMed-Trad_Corpus.json
├── HiMed-West_Bench.json
├── HiMed-West_Corpus.json
└── HiMed-West_Exam.json
We define multiple Hugging Face dataset… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HiMed.Medical-R1-Distill-Data-Chinese
Introduction
This dataset is an SFT dataset distilled from Deepseek-R1 (Full Power Version), based on Chinese medical verifiable problems from HuatuoGPT-o1.
The distillation originates from the native Deepseek-R1 API requests. We hope this distilled dataset can help initialize your models with the reasoning chain from R1. You can also use our previously built medical verified long reasoning chains based on GPT-4o on medical-o1-reasoning-SFT.
For details, see our paper and GitHub… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data-Chinese.ScalpelBench
ScalpelBench
ScalpelBench is a compact instruction-tuning corpus developed for controlled
studies of model compression, with a particular focus on layer pruning,
post-pruning recovery, and capability retention. The released corpus contains
approximately 0.1B tokens of instruction-response data spanning general
English, Chinese, mathematical reasoning, and code generation.
Mixture Design
The mixture proportions follow high-level capability-balancing principles… See the full description on the dataset page: https://huggingface.co/datasets/freeai-org/ScalpelBench.OnePO-Medical-20K
OnePO-Medical-20K
📄 Paper |
💻 GitHub
⚡ Introduction
OnePO-Medical-20K is the medical RL dataset released with OnePO, containing 20,338 medical tasks across multiple languages.
One stage, no preceding SFT. OnePO adapts pretrained models to medicine through a single reinforcement-learning stage.
Two complementary task types. Multiple-choice questions provide verifiable answers. Open-ended conversations provide scoring rubrics.
Teacher guidance included. Each task includes a… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/OnePO-Medical-20K.freebase-neo4j-graph
Freebase → Neo4j Graph
A property-graph conversion of the final Freebase RDF dump (English-filtered),
including proper resolution of Freebase's Compound Value Type (CVT) nodes,
ready for import into Neo4j or use as a general-purpose large knowledge graph.
Freebase was a large collaborative knowledge base, discontinued by Google in
2016. This dataset is derived from the last publicly available RDF dump
(freebase-rdf-latest.gz, 1.9B raw triples), filtered to English-language… See the full description on the dataset page: https://huggingface.co/datasets/ksk-1729/freebase-neo4j-graph.ALLaVA-4V-Chinese
ALLaVA-4V for Chinese
This is the Chinese version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Chinese through ChatGPT and instructed ChatGPT not to translate content related to OCR.
The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V.
Citation
If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Chinese.Regulations-of-the-Free-State-Militia
Dataset Card: Regulations of the Free State Militia – Binding Constitutional Law
⚖️ STATUS: BINDING CONSTITUTIONAL LAW ⚖️
Links
Google Docs (Original): The Regulations of the Free State Militia
Hugging Face Dataset: https://huggingface.co/datasets/RaddicalSilly/Regulations-of-the-Free-State-Militia
Status Declaration
These Regulations are binding constitutional law.
The Regulations of the Free State Militia fulfill the Second Amendment's… See the full description on the dataset page: https://huggingface.co/datasets/RaddicalSilly/Regulations-of-the-Free-State-Militia.mascarade-freecad-dataset
Mascarade — FreeCAD / OpenSCAD / CAD parametric Q&A
Description
Q&A bilingue (FR/EN) sur la CAO paramétrique : scripting FreeCAD Python (Part, PartDesign, Sketcher, Draft), code OpenSCAD, CadQuery, modélisation 3D pour impression, design-for-manufacturing.
Ce dataset fait partie de la famille Mascarade, un corpus thématique destiné au fine-tuning LoRA de modèles compacts (cibles : Qwen 2.5-32B, Gemma 3n-E4B, Qwen3 4B) pour des assistants spécialisés en électronique… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/mascarade-freecad-dataset.research
Freederia Research Archive Dataset Card
Freederia is a large-scale research-data archive for AI agents, RAG builders, search systems, and technical-intelligence workflows.
The archive contains synthetic exploratory research records, problem-anchored technical records, technical-intelligence reports, public HTML articles, metadata indexes, ontology graphs, quality reports, manifests, ledger events, source or basis records, and machine-readable package files.
This Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/freederia/research.ALLaVA-4V-Arabic
ALLaVA-4V for Arabic
This is the Arabic version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Arabic through ChatGPT and instructed ChatGPT not to translate content related to OCR.
The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V.
Citation
If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of Hong… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Arabic.wangchanx-seed-free-synthetic-instruct-thai-120k
Dataset Card for WangchanX Seed-Free Synthetic Instruct Thai 120k
Dataset Summary
This dataset contains about 120k synthetic instruction-following samples in Thai, generated using a novel seed-free approach. It covers a wide range of domains derived from Wikipedia, including both general knowledge and Thai-specific cultural topics. The dataset is designed for instruction-tuning Thai language models to improve their ability to understand and generate Thai text in various… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/wangchanx-seed-free-synthetic-instruct-thai-120k.MatCha
Dataset Description
Materials characterization plays a key role in understanding the processing–microstructure–property relationships that guide material design and optimization. While multimodal large language models (MLLMs) have shown promise in generative and predictive tasks, their ability to interpret real-world characterization imaging data remains underexplored.
MatCha is the first benchmark designed specifically for materials characterization image understanding. It… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/MatCha.
