CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FreedomIntelligence /medical-o1-reasoning-SFT News [2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data. [2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1. [2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.textquestion-answering10K<n<100K1.2k likes21k downloads1y agoHugging Face02FreedomIntelligence /ShareGPT-4o-Image 📚 ShareGPT-4o-Image ShareGPT-4o-Image is a large-scale and high-quality image generation dataset, where all images are produced by GPT-4o’s image generation capabilities. This dataset is designed to align open multimodal models with GPT-4o’s strengths in visual content creation. It includes 45K text-to-image and 46K text-and-image-to-image samples, making it a useful resource for enhancing multimodal models in both image generation and editing tasks. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ShareGPT-4o-Image.texttext-to-image10K<n<100K102 likes5.3k downloads1y agoHugging Face03FreedomIntelligence /CMB CMB: A Comprehensive Medical Benchmark in Chinese 🌐 Github • 🌐 Website • 🤗 HuggingFace 🌈 Update [2024.02.21] The answers to the CMB-Exam test has been updated and some errors caused by omissions in version management have been fixed. [2024.01.08] In order to facilitate testing, we disclose the answers to the CMB-Exam test [2023.09.22] CMB is included in OpenCompass. [2023.08.21] Paper released. [2023.08.01] 🎉🎉🎉 CMB is published!🎉🎉🎉 🌐… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/CMB.textquestion-answeringn<1K38 likes3.1k downloads2y agoHugging Face04FreedomIntelligence /TalkVid TalkVid Dataset This repository hosts the TalkVid dataset. Paper: TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis Arxiv paper: https://arxiv.org/abs/2508.13618 Project Page: https://freedomintelligence.github.io/talk-vid GitHub: https://github.com/FreedomIntelligence/TalkVid Abstract Audio-driven talking head synthesis has achieved remarkable photorealism, yet state-of-the-art (SOTA) models exhibit a critical failure: they lack… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TalkVid.audioimage-to-videon<1K22 likes3k downloads1y agoHugging Face05FreedomIntelligence /alpaca-gpt4-chineseThe dataset is used in the research related to MultilingualSIFT. text10K<n<100K16 likes2.1k downloads3y agoHugging Face06FreedomIntelligence /PubMedVision News [2025/02/18]: We add the original captions of PubMedVision in PubMedVision_Original_Caption.json, as well as the Chinese version of PubMedVision in PubMedVision_Chinese.json. [2024/07/01]: We add annotations for 'body_part' and 'modality' of images, utilizing the HuatuoGPT-Vision-7B model. PubMedVision PubMedVision is a large-scale medical VQA dataset. We extracted high-quality image-text pairs from PubMed and used GPT-4V to reformat them to enhance their quality.… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/PubMedVision.imagequestion-answering1M<n<10M107 likes1.7k downloads2y agoHugging Face07FreedomIntelligence /ApolloCorpus Multilingual Medicine: Model, Dataset, Benchmark, Code Covering English, Chinese, French, Hindi, Spanish, Hindi, Arabic So far 👨🏻‍💻Github •📃 Paper • 🌐 Demo • 🤗 ApolloCorpus • 🤗 XMedBench 中文 | English 🌈 Update [2024.03.07] Paper released. [2024.02.12] ApolloCorpus and XMedBench is published!🎉 [2024.01.23] Apollo repo is published!🎉 Results Apollo-0.5B • 🤗 Apollo-1.8B • 🤗 Apollo-2B • 🤗 Apollo-6B • 🤗 Apollo-7B… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ApolloCorpus.text1M<n<10M41 likes1.1k downloads2y agoHugging Face08FreedomIntelligence /Medical-R1-Distill-Data Introduction This dataset is an SFT dataset distilled from Deepseek-R1 (Full Power Version), based on medical verifiable problems from HuatuoGPT-o1. The Chinese version of the dataset is available at FreedomIntelligence/Medical-R1-Distill-Data-Chinese. The distillation originates from the native Deepseek-R1 API requests. We hope this distilled dataset can help initialize your models with the reasoning chain from R1. You can also use our previously built medical verified long… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data.textquestion-answering10K<n<100K77 likes945 downloads2y agoHugging Face09FreedomIntelligence /huatuo_encyclopedia_qa Dataset Card for Huatuo_encyclopedia_qa Dataset Summary This dataset has a total of 364,420 pieces of medical QA data, some of which have multiple questions in different ways. We extract medical QA pairs from plain texts (e.g., medical encyclopedias and medical articles). We collected 8,699 encyclopedia entries for diseases and 2,736 encyclopedia entries for medicines on Chinese Wikipedia. Moreover, we crawled 226,432 high-quality medical articles from the Qianwen Health… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_encyclopedia_qa.texttext-generation100K<n<1M91 likes815 downloads3y agoHugging Face10FreedomIntelligence /medical-o1-verifiable-problem Introduction This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes. For details, see our paper and GitHub repository. Citation If you find our data useful, please consider citing our work! @misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.textquestion-answering10K<n<100K124 likes784 downloads2y agoHugging Face11FreedomIntelligence /ALLaVA-4V 📚 ALLaVA-4V Data Generation Pipeline LAION We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here. Vison-FLAN We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here. Wizard We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo. Dataset Cards All datasets can be found here. The structure of naming is shown below: ALLaVA-4V ├──… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V.imagequestion-answering100K<n<1M97 likes738 downloads1y agoHugging Face12FreedomIntelligence /Huatuo26M-Lite Huatuo26M-Lite 📚 Table of Contents 🗂 Dataset Description 📝 Dataset Information ℹ️ Data Distribution 📊 Usage 🔧 Citation 📖 Dataset Description 📝 Huatuo26M-Lite is a refined and optimized dataset based on the Huatuo26M dataset, which has undergone multiple purification processes and rewrites. It has more data dimensions and higher data quality. We welcome you to try using it. Dataset Information ℹ️ Dataset Name: Huatuo26M-Lite Version:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Huatuo26M-Lite.tabulartext-classification100K<n<1M69 likes669 downloads3y agoHugging Face13FreedomIntelligence /TCM-Pretrain-Data-ShizhenGPT 📚 Introduction This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Pretrain-Data-ShizhenGPT.texttext-generation1M<n<10M10 likes663 downloads1y agoHugging Face14FreedomIntelligence /HuatuoGPT-sft-data-v1text100K<n<1M84 likes610 downloads3y agoHugging Face15gt-free-ocr-metrics /omnidocbench-render-compare OmniDocBench Render-and-Compare This dataset contains the rendered HTML reconstructions and comparison images produced by a render-and-compare pipeline — a reference-free visual similarity evaluation framework for OCR systems. Overview The pipeline processes each page of OmniDocBench through a Qwen3.5-122B-A10B OCR model, renders the structured output back to a PNG via HTML (reconstructed.png), and compares it against the original page scan (masked_original.png) using… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare.imageother10K<n<100K0 likes511 downloads5mo agoHugging Face16FreedomIntelligence /sharegpt-chineseChinese ShareGPT data translated by gpt-3.5-turbo.The dataset is used in the research related to MultilingualSIFT. text10K<n<100K16 likes432 downloads3y agoHugging Face17FreedomIntelligence /TCM-Instruction-Tuning-ShizhenGPT 📚 Introduction This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced fine-tuning dataset consists of three parts: Modality Data Quantity TCM Text Instructions 📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.textquestion-answering100K<n<1M13 likes420 downloads1y agoHugging Face18FreedomIntelligence /huatuo_consultation_qa Dataset Card for huatuo_consultation_qa Dataset Summary We collected data from a website for medical consultation , consisting of many online consultation records by medical experts. Each record is a QA pair: a patient raises a question and a medical doctor answers the question. The basic information of doctors (including name, hospital organization, and department) was recorded. We directly crawl patient’s questions and doctor’s answers as QA pairs, getting 32,708,346… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_consultation_qa.texttext-generation10M<n<100M16 likes417 downloads3y agoHugging Face19FreedomIntelligence /huatuo_knowledge_graph_qa Dataset Card for Huatuo_knowledge_graph_qa Dataset Summary We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map. Dataset Creation Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.texttext-generation100K<n<1M52 likes414 downloads3y agoHugging Face20FreedomIntelligence /Medical_Multimodal_Evaluation_Data Evaluation Guide This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks. To get started: Download the dataset and extract the images.zip file. Find evaluation code on our GitHub: HuatuoGPT-Vision. This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.imageimage-to-text10K<n<100K29 likes352 downloads2y agoHugging Face21TableSenseAI /FreeformTableQAtext1K<n<10K0 likes337 downloads1y agoHugging Face22FreedomIntelligence /CoD-PatientSymDisease Citation @misc{chen2024codinterpretablemedicalagent, title={CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis}, author={Junying Chen and Chi Gui and Anningzhe Gao and Ke Ji and Xidong Wang and Xiang Wan and Benyou Wang}, year={2024}, eprint={2407.13301}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2407.13301}, } texttoken-classification10K<n<100K14 likes336 downloads2y agoHugging Face23FreedomIntelligence /alpaca-gpt4-indonesianThe dataset is used in the research related to MultilingualSIFT. text10K<n<100K16 likes317 downloads3y agoHugging Face24FreedomIntelligence /huatuo26M-testdatasets Dataset Card for huatuo26M-testdatasets Dataset Summary We are pleased to announce the release of our evaluation dataset, a subset of the Huatuo-26M. This dataset contains 6,000 entries that we used for Natural Language Generation (NLG) experimentation in our associated research paper. We encourage researchers and developers to use this evaluation dataset to gauge the performance of their own models. This is not only a chance to assess the accuracy and relevancy of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo26M-testdatasets.texttext-generation1K<n<10K22 likes283 downloads3y agoHugging Face25FreedomIntelligence /ApolloMoEDataset Democratizing Medical LLMs For Much More Languages Covering 12 Major Languages including English, Chinese, French, Hindi, Spanish, Arabic, Russian, Japanese, Korean, German, Italian, Portuguese and 38 Minor Languages So far. 📃 Paper • 🌐 Demo • 🤗 ApolloMoEDataset • 🤗 ApolloMoEBench • 🤗 Models •🌐 Apollo • 🌐 ApolloMoE 🌈 Update [2024.10.15] ApolloMoE repo is published!🎉 Languages Coverage 12 Major Languages and 38 Minor Languages Click to… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ApolloMoEDataset.textquestion-answering100K<n<1M6 likes278 downloads2y agoHugging Face26FreedomIntelligence /evol-instruct-chineseThe dataset is used in the research related to MultilingualSIFT. text10K<n<100K10 likes265 downloads3y agoHugging Face27FreedomIntelligence /DxBench Citation @misc{chen2024codinterpretablemedicalagent, title={CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis}, author={Junying Chen and Chi Gui and Anningzhe Gao and Ke Ji and Xidong Wang and Xiang Wan and Benyou Wang}, year={2024}, eprint={2407.13301}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2407.13301}, } texttoken-classification1K<n<10K9 likes262 downloads2y agoHugging Face28FreedomIntelligence /XMedbench Multilingual Medicine: Model, Dataset, Benchmark, Code Covering English, Chinese, French, Hindi, Spanish, Hindi, Arabic So far 👨🏻‍💻Github •📃 Paper • 🤗 ApolloCorpus • 🤗 XMedBench 中文 | English 🌈 Update [2024.03.07] Paper released. [2024.02.12] ApolloCorpus and XMedBench is published!🎉 [2024.01.23] Apollo repo is published!🎉 Results Usage Zip File Data category Data: EN: MedQA-USMLE MedMCQA PubMedQA:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/XMedbench.text10K<n<100K14 likes245 downloads2y agoHugging Face29FreedomIntelligence /alpaca-gpt4-arabicThe dataset is used in the research related to MultilingualSIFT. text10K<n<100K12 likes236 downloads3y agoHugging Face30FreedomIntelligence /BlendNet 📚 BlendNet The dataset contains $12k$ samples. To balance cost savings with data quality and scale, we manually annotated $2k$ samples and used GPT-4o to annotate the remaining $10k$ samples. For more details, please visit our GitHub repository or refer to our arXiv paper. 📖 Citation @misc{du2024blenderllmtraininglargelanguage, title={BlenderLLM: Training Large Language Models for Computer-Aided Design with Self-improvement}, author={Yuhao Du and… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/BlendNet.tabular10K<n<100K12 likes224 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.