CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FreedomIntelligence /medical-o1-reasoning-SFT News [2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data. [2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1. [2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.textquestion-answering10K<n<100K1.2k likes21k downloads1y agoHugging Face02FreedomIntelligence /CMB CMB: A Comprehensive Medical Benchmark in Chinese 🌐 Github • 🌐 Website • 🤗 HuggingFace 🌈 Update [2024.02.21] The answers to the CMB-Exam test has been updated and some errors caused by omissions in version management have been fixed. [2024.01.08] In order to facilitate testing, we disclose the answers to the CMB-Exam test [2023.09.22] CMB is included in OpenCompass. [2023.08.21] Paper released. [2023.08.01] 🎉🎉🎉 CMB is published!🎉🎉🎉 🌐… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/CMB.textquestion-answeringn<1K38 likes3.2k downloads2y agoHugging Face03FreedomIntelligence /MileBench MileBench Introduction We introduce MileBench, a pioneering benchmark designed to test the MultImodal Long-contExt capabilities of MLLMs. This benchmark comprises not only multimodal long contexts, but also multiple tasks requiring both comprehension and generation. We establish two distinct evaluation sets, diagnostic and realistic, to systematically assess MLLMs’ long-context adaptation capacity and their ability to completetasks in long-context scenarios To… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/MileBench.visual-question-answering1K<n<10K9 likes2k downloads2y agoHugging Face04CleverThis /freebase Freebase Dataset Description Large-scale knowledge base (archived by Google) Original Source: http://commondatastorage.googleapis.com/freebase-public/rdf/freebase-rdf-latest.gz Dataset Summary This dataset contains RDF triples from Freebase converted to HuggingFace dataset format for easy use in machine learning pipelines. Format: Originally ntriples, converted to HuggingFace Dataset Size: 300.0 GB (extracted) Entities: ~50M Triples: ~3B Original License: CC… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/freebase.texttext-generation1B<n<10B1 likes1.7k downloads10mo agoHugging Face05FreedomIntelligence /PubMedVision News [2025/02/18]: We add the original captions of PubMedVision in PubMedVision_Original_Caption.json, as well as the Chinese version of PubMedVision in PubMedVision_Chinese.json. [2024/07/01]: We add annotations for 'body_part' and 'modality' of images, utilizing the HuatuoGPT-Vision-7B model. PubMedVision PubMedVision is a large-scale medical VQA dataset. We extracted high-quality image-text pairs from PubMed and used GPT-4V to reformat them to enhance their quality.… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/PubMedVision.imagequestion-answering1M<n<10M107 likes1.7k downloads2y agoHugging Face06FreedomIntelligence /Medical-R1-Distill-Data Introduction This dataset is an SFT dataset distilled from Deepseek-R1 (Full Power Version), based on medical verifiable problems from HuatuoGPT-o1. The Chinese version of the dataset is available at FreedomIntelligence/Medical-R1-Distill-Data-Chinese. The distillation originates from the native Deepseek-R1 API requests. We hope this distilled dataset can help initialize your models with the reasoning chain from R1. You can also use our previously built medical verified long… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data.textquestion-answering10K<n<100K77 likes969 downloads2y agoHugging Face07FreedomIntelligence /huatuo_encyclopedia_qa Dataset Card for Huatuo_encyclopedia_qa Dataset Summary This dataset has a total of 364,420 pieces of medical QA data, some of which have multiple questions in different ways. We extract medical QA pairs from plain texts (e.g., medical encyclopedias and medical articles). We collected 8,699 encyclopedia entries for diseases and 2,736 encyclopedia entries for medicines on Chinese Wikipedia. Moreover, we crawled 226,432 high-quality medical articles from the Qianwen Health… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_encyclopedia_qa.texttext-generation100K<n<1M91 likes877 downloads3y agoHugging Face08FreedomIntelligence /ALLaVA-4V 📚 ALLaVA-4V Data Generation Pipeline LAION We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here. Vison-FLAN We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here. Wizard We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo. Dataset Cards All datasets can be found here. The structure of naming is shown below: ALLaVA-4V ├──… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V.imagequestion-answering100K<n<1M98 likes829 downloads1y agoHugging Face09FreedomIntelligence /medical-o1-verifiable-problem Introduction This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes. For details, see our paper and GitHub repository. Citation If you find our data useful, please consider citing our work! @misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.textquestion-answering10K<n<100K124 likes778 downloads2y agoHugging Face10Yariz /Freebase Freebase Dataset Description Large-scale knowledge base (archived by Google) Original Source: http://commondatastorage.googleapis.com/freebase-public/rdf/freebase-rdf-latest.gz Dataset Summary This dataset contains RDF triples from Freebase converted to HuggingFace dataset format for easy use in machine learning pipelines. Format: Originally ntriples, converted to HuggingFace Dataset Size: 300.0 GB (extracted) Entities: ~50M Triples: ~3B Original License: CC… See the full description on the dataset page: https://huggingface.co/datasets/Yariz/Freebase.texttext-generation1B<n<10B0 likes768 downloads8mo agoHugging Face11FreedomIntelligence /TCM-Pretrain-Data-ShizhenGPT 📚 Introduction This dataset is the pre-training dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source the largest existing TCM corpus dataset (over 5B tokens) from TCM-related websites and books. Additionally, we also open-source the largest scale TCM image-text pretraining dataset. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced pre-training dataset consists of five parts:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Pretrain-Data-ShizhenGPT.texttext-generation1M<n<10M10 likes673 downloads1y agoHugging Face12FreedomIntelligence /Huatuo26M-Lite Huatuo26M-Lite 📚 Table of Contents 🗂 Dataset Description 📝 Dataset Information ℹ️ Data Distribution 📊 Usage 🔧 Citation 📖 Dataset Description 📝 Huatuo26M-Lite is a refined and optimized dataset based on the Huatuo26M dataset, which has undergone multiple purification processes and rewrites. It has more data dimensions and higher data quality. We welcome you to try using it. Dataset Information ℹ️ Dataset Name: Huatuo26M-Lite Version:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Huatuo26M-Lite.tabulartext-classification100K<n<1M69 likes628 downloads3y agoHugging Face13arpandeepk /swe-zero-free SWE-ZERO-Free, 12M agentic coding trajectories made without an LLM SWE-ZERO-12M showed you can generate agentic SWE data without Docker, just by sticking to shell commands that need no setup. We took that idea and asked whether you even need the model. You don't. So this is SWE-ZERO, but free. This dataset has 12,238,610 mini-swe-agent trajectories across 119,084 real GitHub PRs. Same source and same format as SWE-ZERO. The only difference is how they get made. SWE-ZERO samples… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/swe-zero-free.texttext-generation10M<n<100M0 likes569 downloads4mo agoHugging Face14FreedomIntelligence /huatuo_consultation_qa Dataset Card for huatuo_consultation_qa Dataset Summary We collected data from a website for medical consultation , consisting of many online consultation records by medical experts. Each record is a QA pair: a patient raises a question and a medical doctor answers the question. The basic information of doctors (including name, hospital organization, and department) was recorded. We directly crawl patient’s questions and doctor’s answers as QA pairs, getting 32,708,346… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_consultation_qa.texttext-generation10M<n<100M16 likes449 downloads3y agoHugging Face15FreedomIntelligence /huatuo_knowledge_graph_qa Dataset Card for Huatuo_knowledge_graph_qa Dataset Summary We built this QA dataset based on the medical knowledge map, with a total of 798,444 pieces of data, in which the questions are constructed by means of templates, and the answers are the contents of the entries in the knowledge map. Dataset Creation Source Data… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo_knowledge_graph_qa.texttext-generation100K<n<1M52 likes424 downloads3y agoHugging Face16FreedomIntelligence /TCM-Instruction-Tuning-ShizhenGPT 📚 Introduction This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced fine-tuning dataset consists of three parts: Modality Data Quantity TCM Text Instructions 📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.textquestion-answering100K<n<1M13 likes420 downloads1y agoHugging Face17FreedomIntelligence /HuatuoGPT2-SFT-GPT4-140K HuatuoGPT2-SFT-GPT4-140K 140K Chinese medical instructions generated by GPT-4, based on questions from HuatuoGPT Dataset. This dataset contains supervised fine-tuning instructions for HuatuoGPT2, designed to enhance the model's ability to follow instructions in real medical scenarios. We have made all the data (142,248 entries) in this dataset publicly available. Repository Github: https://github.com/FreedomIntelligence/HuatuoGPT-II Citation… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-SFT-GPT4-140K.question-answering14 likes415 downloads2y agoHugging Face18FreedomIntelligence /huatuo26M-testdatasets Dataset Card for huatuo26M-testdatasets Dataset Summary We are pleased to announce the release of our evaluation dataset, a subset of the Huatuo-26M. This dataset contains 6,000 entries that we used for Natural Language Generation (NLG) experimentation in our associated research paper. We encourage researchers and developers to use this evaluation dataset to gauge the performance of their own models. This is not only a chance to assess the accuracy and relevancy of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/huatuo26M-testdatasets.texttext-generation1K<n<10K22 likes298 downloads3y agoHugging Face19free-law /Caselaw_Access_Projectgated The Caselaw Access Project In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/ Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here:… See the full description on the dataset page: https://huggingface.co/datasets/free-law/Caselaw_Access_Project.texttext-generation1M<n<10M102 likes232 downloads3y agoHugging Face20FreedomIntelligence /Evol-Instruct-Chinese-GPT4The dataset is created by (1) translating English questions of Evol-instruct-70k into Chinese and (2) requesting GPT4 to generate Chinese responses. For more details, please refer to: Repository: https://github.com/FreedomIntelligence/AceGPT https://github.com/FreedomIntelligence/LLMZoo Paper: AceGPT, Localizing Large Language Models in Arabic Phoenix: Democratizing ChatGPT across Languages BibTeX entry and citation info @article{huang2023acegpt, title={AceGPT, Localizing… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Evol-Instruct-Chinese-GPT4.texttext-generation10K<n<100K47 likes223 downloads3y agoHugging Face21free-law /Caselaw_Access_Project_embeddings The Caselaw Access Project In collaboration with Ravel Law, Harvard Law Library digitized over 40 million U.S. court decisions consisting of 6.7 million cases from the last 360 years into a dataset that is widely accessible to use. Access a bulk download of the data through the Caselaw Access Project API (CAPAPI): https://case.law/caselaw/ Find more information about accessing state and federal written court decisions of common law through the bulk data service documentation here:… See the full description on the dataset page: https://huggingface.co/datasets/free-law/Caselaw_Access_Project_embeddings.texttext-generation1M<n<10M9 likes222 downloads3y agoHugging Face22nvidia /Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1 Dataset Description: Teaches the model to follow arbitrary text formatting instructions (bullet styles, numbering, delimiters, heading formats, inline emphasis, web-answer structure, etc.) for targeted chat behaviors. Uses explicit Regex and string matching for the reward signal. This dataset is ready for commercial or non-commercial uses. Dataset Owner(s): NVIDIA Corporation Dataset Creation Date: Created on: April 10, 2026 Last Modified on: April… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Free-Form-Formatting-v1.texttext-generation1K<n<10K2 likes218 downloads4mo agoHugging Face23nassimjp /Pashto-Free-Hand-Reasoning-Dataset Pashto Free-Hand Reasoning SFT Dataset 🧠♻️ This dataset contains high-quality, long-form SFT (Supervised Fine-Tuning) conversational data in Pashto, featuring unconstrained, natural model reasoning (<think> blocks) paired with standardized chat responses. 🔄 The 3R Approach (Recycle, Reuse, Reason) Instead of discarding legacy QA pairs, this dataset follows a 3R data philosophy: Recycle: Taking older, simple, or raw legacy Pashto questions. Reuse: Re-processing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Free-Hand-Reasoning-Dataset.texttext-generation1K<n<10K0 likes215 downloads8d agoHugging Face24FreedomIntelligence /RAG-Instruct Introduction RAG-Instruct is a RAG dataset designed to comprehensively enhance LLM RAG capabilities, synthesized using GPT-4o. This dataset is based on the Wikipedia corpus and This dataset is based on the Wikipedia corpus and offers the advantages of query-document scenario diversity and task diversity. The RAG-Instruct dataset can significantly enhance the RAG ability of LLMs and make remarkable improvements in RAG performance across various tasks. Model WQA (acc) PQA (acc)… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/RAG-Instruct.textquestion-answering10K<n<100K43 likes189 downloads2y agoHugging Face25FreedomIntelligence /HuatuoGPT2-Pretraining-Instruction HuatuoGPT2-Pretraining-Instruction-5200K Here are the pre-training instructions for HuatuoGPT-II, developed with 5.2 million medical corpus using ChatGPT. This dataset is used to incorporate extensive medical knowledge and enable a one-stage medical adaptation. All our data have been made publicly accessible. Data Volume The following table details the volume and distribution of pre-training data for HuatuoGPT2: Data Source Data Volume Medical_Web_Corpus_cn… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-Pretraining-Instruction.textquestion-answering1M<n<10M13 likes183 downloads2y agoHugging Face26arpandeepk /swe-zero-free-v2 SWE-ZERO-Free, 12M agentic coding trajectories made without an LLM SWE-ZERO-12M showed you can generate agentic SWE data without Docker, just by sticking to shell commands that need no setup. We took that idea and asked whether you even need the model. You don't. So this is SWE-ZERO, but free. This dataset has 12,177,154 mini-swe-agent trajectories across ~121,800 real GitHub PRs. Same source and same format as SWE-ZERO. The only difference is how they get made. SWE-ZERO samples… See the full description on the dataset page: https://huggingface.co/datasets/arpandeepk/swe-zero-free-v2.texttext-generation10M<n<100M0 likes165 downloads4mo agoHugging Face27FreedomIntelligence /Medical-R1-Distill-Data-Chinese Introduction This dataset is an SFT dataset distilled from Deepseek-R1 (Full Power Version), based on Chinese medical verifiable problems from HuatuoGPT-o1. The distillation originates from the native Deepseek-R1 API requests. We hope this distilled dataset can help initialize your models with the reasoning chain from R1. You can also use our previously built medical verified long reasoning chains based on GPT-4o on medical-o1-reasoning-SFT. For details, see our paper and GitHub… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data-Chinese.textquestion-answering10K<n<100K48 likes163 downloads2y agoHugging Face28freeai-org /ScalpelBench ScalpelBench ScalpelBench is a compact instruction-tuning corpus developed for controlled studies of model compression, with a particular focus on layer pruning, post-pruning recovery, and capability retention. The released corpus contains approximately 0.1B tokens of instruction-response data spanning general English, Chinese, mathematical reasoning, and code generation. Mixture Design The mixture proportions follow high-level capability-balancing principles… See the full description on the dataset page: https://huggingface.co/datasets/freeai-org/ScalpelBench.texttext-generation100K<n<1M1 likes139 downloads24d agoHugging Face29polaris-73 /monitorability-as-a-free-gift-data-training-data Monitorability as a free gift training data reordered, or normalized during packaging. Configurations Config Rows Purpose Original file all 18,591 Main all-domain experiment combined_dataset.parquet no_if 13,591 All-domain experiment without instruction following combined_dataset_noif.parquet instruction_following 5,000 Instruction-following experiments instruction_following_ai2_5000.parquet math 5,000 Main math experiments skywork_math.parquet… See the full description on the dataset page: https://huggingface.co/datasets/polaris-73/monitorability-as-a-free-gift-data-training-data.texttext-generation10K<n<100K0 likes136 downloads2mo agoHugging Face30FreedomIntelligence /OnePO-Medical-20K OnePO-Medical-20K 📄 Paper | 💻 GitHub ⚡ Introduction OnePO-Medical-20K is the medical RL dataset released with OnePO, containing 20,338 medical tasks across multiple languages. One stage, no preceding SFT. OnePO adapts pretrained models to medicine through a single reinforcement-learning stage. Two complementary task types. Multiple-choice questions provide verifiable answers. Open-ended conversations provide scoring rubrics. Teacher guidance included. Each task includes a… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/OnePO-Medical-20K.textquestion-answering10K<n<100K2 likes102 downloads2h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.