CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01FreedomIntelligence /medical-o1-reasoning-SFT News [2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data. [2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1. [2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.textquestion-answering10K<n<100K1.2k likes21k downloads1y agoHugging Face02FreedomIntelligence /CMB CMB: A Comprehensive Medical Benchmark in Chinese 🌐 Github • 🌐 Website • 🤗 HuggingFace 🌈 Update [2024.02.21] The answers to the CMB-Exam test has been updated and some errors caused by omissions in version management have been fixed. [2024.01.08] In order to facilitate testing, we disclose the answers to the CMB-Exam test [2023.09.22] CMB is included in OpenCompass. [2023.08.21] Paper released. [2023.08.01] 🎉🎉🎉 CMB is published!🎉🎉🎉 🌐… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/CMB.textquestion-answeringn<1K38 likes3.2k downloads2y agoHugging Face03FreedomIntelligence /MileBench MileBench Introduction We introduce MileBench, a pioneering benchmark designed to test the MultImodal Long-contExt capabilities of MLLMs. This benchmark comprises not only multimodal long contexts, but also multiple tasks requiring both comprehension and generation. We establish two distinct evaluation sets, diagnostic and realistic, to systematically assess MLLMs’ long-context adaptation capacity and their ability to completetasks in long-context scenarios To… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/MileBench.visual-question-answering1K<n<10K9 likes2k downloads2y agoHugging Face04FreedomIntelligence /PubMedVision News [2025/02/18]: We add the original captions of PubMedVision in PubMedVision_Original_Caption.json, as well as the Chinese version of PubMedVision in PubMedVision_Chinese.json. [2024/07/01]: We add annotations for 'body_part' and 'modality' of images, utilizing the HuatuoGPT-Vision-7B model. PubMedVision PubMedVision is a large-scale medical VQA dataset. We extracted high-quality image-text pairs from PubMed and used GPT-4V to reformat them to enhance their quality.… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/PubMedVision.imagequestion-answering1M<n<10M107 likes1.7k downloads2y agoHugging Face05FreedomIntelligence /Medical-R1-Distill-Data Introduction This dataset is an SFT dataset distilled from Deepseek-R1 (Full Power Version), based on medical verifiable problems from HuatuoGPT-o1. The Chinese version of the dataset is available at FreedomIntelligence/Medical-R1-Distill-Data-Chinese. The distillation originates from the native Deepseek-R1 API requests. We hope this distilled dataset can help initialize your models with the reasoning chain from R1. You can also use our previously built medical verified long… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data.textquestion-answering10K<n<100K77 likes969 downloads2y agoHugging Face06FreedomIntelligence /ALLaVA-4V 📚 ALLaVA-4V Data Generation Pipeline LAION We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here. Vison-FLAN We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here. Wizard We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo. Dataset Cards All datasets can be found here. The structure of naming is shown below: ALLaVA-4V ├──… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V.imagequestion-answering100K<n<1M98 likes829 downloads1y agoHugging Face07FreedomIntelligence /medical-o1-verifiable-problem Introduction This dataset features open-ended medical problems designed to improve LLMs' medical reasoning. Each entry includes a open-ended question and a ground-truth answer based on challenging medical exams. The verifiable answers enable checking LLM outputs, refining their reasoning processes. For details, see our paper and GitHub repository. Citation If you find our data useful, please consider citing our work! @misc{chen2024huatuogpto1medicalcomplexreasoning… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-verifiable-problem.textquestion-answering10K<n<100K124 likes778 downloads2y agoHugging Face08FreedomIntelligence /Huatuo26M-Lite Huatuo26M-Lite 📚 Table of Contents 🗂 Dataset Description 📝 Dataset Information ℹ️ Data Distribution 📊 Usage 🔧 Citation 📖 Dataset Description 📝 Huatuo26M-Lite is a refined and optimized dataset based on the Huatuo26M dataset, which has undergone multiple purification processes and rewrites. It has more data dimensions and higher data quality. We welcome you to try using it. Dataset Information ℹ️ Dataset Name: Huatuo26M-Lite Version:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Huatuo26M-Lite.tabulartext-classification100K<n<1M69 likes628 downloads3y agoHugging Face09KelvinJiang /freebase_qaFreebaseQA is for open-domain factoid question answering (QA) tasks over structured knowledge bases, like Freebase The data set is generated by matching trivia-type question-answer pairs with subject-predicateobject triples in Freebase.question-answering10K<n<100K8 likes544 downloads3y agoHugging Face10FreedomIntelligence /TCM-Instruction-Tuning-ShizhenGPT 📚 Introduction This dataset is a fine-tuning dataset for ShizhenGPT, a multimodal LLM for Traditional Chinese Medicine (TCM). We open-source 245K multimodal Chinese medicine instruction data, including text instructions, visual instructions, and signal instructions for TCM. For details, see our paper and GitHub repository. 📊 Dataset Overview The open-sourced fine-tuning dataset consists of three parts: Modality Data Quantity TCM Text Instructions 📝 Text… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TCM-Instruction-Tuning-ShizhenGPT.textquestion-answering100K<n<1M13 likes420 downloads1y agoHugging Face11FreedomIntelligence /HuatuoGPT2-SFT-GPT4-140K HuatuoGPT2-SFT-GPT4-140K 140K Chinese medical instructions generated by GPT-4, based on questions from HuatuoGPT Dataset. This dataset contains supervised fine-tuning instructions for HuatuoGPT2, designed to enhance the model's ability to follow instructions in real medical scenarios. We have made all the data (142,248 entries) in this dataset publicly available. Repository Github: https://github.com/FreedomIntelligence/HuatuoGPT-II Citation… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-SFT-GPT4-140K.question-answering14 likes415 downloads2y agoHugging Face12FreedomIntelligence /ApolloMoEDataset Democratizing Medical LLMs For Much More Languages Covering 12 Major Languages including English, Chinese, French, Hindi, Spanish, Arabic, Russian, Japanese, Korean, German, Italian, Portuguese and 38 Minor Languages So far. 📃 Paper • 🌐 Demo • 🤗 ApolloMoEDataset • 🤗 ApolloMoEBench • 🤗 Models •🌐 Apollo • 🌐 ApolloMoE 🌈 Update [2024.10.15] ApolloMoE repo is published!🎉 Languages Coverage 12 Major Languages and 38 Minor Languages Click to… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ApolloMoEDataset.textquestion-answering100K<n<1M6 likes328 downloads2y agoHugging Face13FreedomIntelligence /Medical_Multimodal_Evaluation_Data Evaluation Guide This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks. To get started: Download the dataset and extract the images.zip file. Find evaluation code on our GitHub: HuatuoGPT-Vision. This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.imageimage-to-text10K<n<100K29 likes314 downloads2y agoHugging Face14FreedomIntelligence /EchoX-Dialougues EchoX-Dialogues: Training Data for EchoX: Towards Mitigating Acoustic-Semantic Gap via Echo Training for Speech-to-Speech LLMs 🐈‍⬛ Github | 📃 Paper | 🚀 Space  🧠 EchoX-8B | 🧠 EchoX-3B | 📦 EchoX-Dialogues-Plus  EchoX-Dialogues provides the primary speech dialogue data used to train EchoX, restricted to S2T (speech → text) in this repository. All input speech is synthetic; text is derived from public sources with multi-stage cleaning and rewriting. Most turns include asr /… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/EchoX-Dialougues.automatic-speech-recognition4 likes244 downloads1y agoHugging Face15nikhilchandak /GPQA-diamond-freetextquestion-answeringn<1K0 likes191 downloads1y agoHugging Face16FreedomIntelligence /RAG-Instruct Introduction RAG-Instruct is a RAG dataset designed to comprehensively enhance LLM RAG capabilities, synthesized using GPT-4o. This dataset is based on the Wikipedia corpus and This dataset is based on the Wikipedia corpus and offers the advantages of query-document scenario diversity and task diversity. The RAG-Instruct dataset can significantly enhance the RAG ability of LLMs and make remarkable improvements in RAG performance across various tasks. Model WQA (acc) PQA (acc)… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/RAG-Instruct.textquestion-answering10K<n<100K43 likes189 downloads2y agoHugging Face17FreedomIntelligence /HuatuoGPT2-Pretraining-Instruction HuatuoGPT2-Pretraining-Instruction-5200K Here are the pre-training instructions for HuatuoGPT-II, developed with 5.2 million medical corpus using ChatGPT. This dataset is used to incorporate extensive medical knowledge and enable a one-stage medical adaptation. All our data have been made publicly accessible. Data Volume The following table details the volume and distribution of pre-training data for HuatuoGPT2: Data Source Data Volume Medical_Web_Corpus_cn… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-Pretraining-Instruction.textquestion-answering1M<n<10M13 likes183 downloads2y agoHugging Face18FreedomIntelligence /ApolloMoEBench Democratizing Medical LLMs For Much More Languages Covering 12 Major Languages including English, Chinese, French, Hindi, Spanish, Arabic, Russian, Japanese, Korean, German, Italian, Portuguese and 38 Minor Languages So far. 📃 Paper • 🌐 Demo • 🤗 ApolloMoEDataset • 🤗 ApolloMoEBench • 🤗 Models •🌐 Apollo • 🌐 ApolloMoE 🌈 Update [2024.10.15] ApolloMoE repo is published!🎉 Languages Coverage 12 Major Languages and 38 Minor Languages Click to… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ApolloMoEBench.textquestion-answering10K<n<100K0 likes173 downloads2y agoHugging Face19FreedomIntelligence /HiMed HiMed HiMed is a Hindi medical dataset and benchmark suite covering both Western medicine and Indian systems of medicine.It consists of two parts: HiMed-Trad: traditional Indian medicine HiMed-West: Western medicine under Hindi prompts Repository Layout All released files are under data/: data/ ├── HiMed-Trad_Bench.json ├── HiMed-Trad_Corpus.json ├── HiMed-West_Bench.json ├── HiMed-West_Corpus.json └── HiMed-West_Exam.json We define multiple Hugging Face dataset… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HiMed.textquestion-answering100K<n<1M2 likes170 downloads9mo agoHugging Face20FreedomIntelligence /Medical-R1-Distill-Data-Chinese Introduction This dataset is an SFT dataset distilled from Deepseek-R1 (Full Power Version), based on Chinese medical verifiable problems from HuatuoGPT-o1. The distillation originates from the native Deepseek-R1 API requests. We hope this distilled dataset can help initialize your models with the reasoning chain from R1. You can also use our previously built medical verified long reasoning chains based on GPT-4o on medical-o1-reasoning-SFT. For details, see our paper and GitHub… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical-R1-Distill-Data-Chinese.textquestion-answering10K<n<100K48 likes163 downloads2y agoHugging Face21freeai-org /ScalpelBench ScalpelBench ScalpelBench is a compact instruction-tuning corpus developed for controlled studies of model compression, with a particular focus on layer pruning, post-pruning recovery, and capability retention. The released corpus contains approximately 0.1B tokens of instruction-response data spanning general English, Chinese, mathematical reasoning, and code generation. Mixture Design The mixture proportions follow high-level capability-balancing principles… See the full description on the dataset page: https://huggingface.co/datasets/freeai-org/ScalpelBench.texttext-generation100K<n<1M1 likes139 downloads25d agoHugging Face22FreedomIntelligence /OnePO-Medical-20K OnePO-Medical-20K 📄 Paper | 💻 GitHub ⚡ Introduction OnePO-Medical-20K is the medical RL dataset released with OnePO, containing 20,338 medical tasks across multiple languages. One stage, no preceding SFT. OnePO adapts pretrained models to medicine through a single reinforcement-learning stage. Two complementary task types. Multiple-choice questions provide verifiable answers. Open-ended conversations provide scoring rubrics. Teacher guidance included. Each task includes a… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/OnePO-Medical-20K.textquestion-answering10K<n<100K5 likes102 downloads15h agoHugging Face23ksk-1729 /freebase-neo4j-graph Freebase → Neo4j Graph A property-graph conversion of the final Freebase RDF dump (English-filtered), including proper resolution of Freebase's Compound Value Type (CVT) nodes, ready for import into Neo4j or use as a general-purpose large knowledge graph. Freebase was a large collaborative knowledge base, discontinued by Google in 2016. This dataset is derived from the last publicly available RDF dump (freebase-rdf-latest.gz, 1.9B raw triples), filtered to English-language… See the full description on the dataset page: https://huggingface.co/datasets/ksk-1729/freebase-neo4j-graph.textgraph-ml100M<n<1B0 likes90 downloads1mo agoHugging Face24FreedomIntelligence /ALLaVA-4V-Chinese ALLaVA-4V for Chinese This is the Chinese version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Chinese through ChatGPT and instructed ChatGPT not to translate content related to OCR. The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V. Citation If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Chinese.imagequestion-answering100K<n<1M16 likes88 downloads2y agoHugging Face25RaddicalSilly /Regulations-of-the-Free-State-Militia Dataset Card: Regulations of the Free State Militia – Binding Constitutional Law ⚖️ STATUS: BINDING CONSTITUTIONAL LAW ⚖️ Links Google Docs (Original): The Regulations of the Free State Militia Hugging Face Dataset: https://huggingface.co/datasets/RaddicalSilly/Regulations-of-the-Free-State-Militia Status Declaration These Regulations are binding constitutional law. The Regulations of the Free State Militia fulfill the Second Amendment's… See the full description on the dataset page: https://huggingface.co/datasets/RaddicalSilly/Regulations-of-the-Free-State-Militia.documenttext-generation1K<n<10K1 likes72 downloads4mo agoHugging Face26Ailiance-fr /mascarade-freecad-dataset Mascarade — FreeCAD / OpenSCAD / CAD parametric Q&A Description Q&A bilingue (FR/EN) sur la CAO paramétrique : scripting FreeCAD Python (Part, PartDesign, Sketcher, Draft), code OpenSCAD, CadQuery, modélisation 3D pour impression, design-for-manufacturing. Ce dataset fait partie de la famille Mascarade, un corpus thématique destiné au fine-tuning LoRA de modèles compacts (cibles : Qwen 2.5-32B, Gemma 3n-E4B, Qwen3 4B) pour des assistants spécialisés en électronique… See the full description on the dataset page: https://huggingface.co/datasets/Ailiance-fr/mascarade-freecad-dataset.texttext-generation1K<n<10K0 likes58 downloads5mo agoHugging Face27freederia /research Freederia Research Archive Dataset Card Freederia is a large-scale research-data archive for AI agents, RAG builders, search systems, and technical-intelligence workflows. The archive contains synthetic exploratory research records, problem-anchored technical records, technical-intelligence reports, public HTML articles, metadata indexes, ontology graphs, quality reports, manifests, ledger events, source or basis records, and machine-readable package files. This Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/freederia/research.text-generation1M<n<10M0 likes57 downloads3mo agoHugging Face28FreedomIntelligence /ALLaVA-4V-Arabic ALLaVA-4V for Arabic This is the Arabic version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Arabic through ChatGPT and instructed ChatGPT not to translate content related to OCR. The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V. Citation If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of Hong… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Arabic.imagequestion-answering100K<n<1M4 likes48 downloads2y agoHugging Face29airesearch /wangchanx-seed-free-synthetic-instruct-thai-120k Dataset Card for WangchanX Seed-Free Synthetic Instruct Thai 120k Dataset Summary This dataset contains about 120k synthetic instruction-following samples in Thai, generated using a novel seed-free approach. It covers a wide range of domains derived from Wikipedia, including both general knowledge and Thai-specific cultural topics. The dataset is designed for instruction-tuning Thai language models to improve their ability to understand and generate Thai text in various… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/wangchanx-seed-free-synthetic-instruct-thai-120k.tabulartext-generation100K<n<1M3 likes46 downloads2y agoHugging Face30FreedomIntelligence /MatCha Dataset Description Materials characterization plays a key role in understanding the processing–microstructure–property relationships that guide material design and optimization. While multimodal large language models (MLLMs) have shown promise in generative and predictive tasks, their ability to interpret real-world characterization imaging data remains underexplored. MatCha is the first benchmark designed specifically for materials characterization image understanding. It… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/MatCha.question-answering5 likes40 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.