CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01medalpaca /medical_meadow_health_advice Health Advice Dataset Summary This is the dataset use in the paper: Detecting Causal Language Use in Science Findings. It was cleaned and formated to fit into the alpaca template. Citation Information @inproceedings{yu-etal-2019-detecting, title = "Detecting Causal Language Use in Science Findings", author = "Yu, Bei and Li, Yingya and Wang, Jun", booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_health_advice.textquestion-answering1K<n<10K11 likes2.9k downloads3y agoHugging Face02Amod /mental_health_counseling_conversationsgated Amod/mental_health_counseling_conversations This dataset is a compilation of high-quality, real one-on-one mental health counseling conversations between individuals and licensed professionals. Each exchange is structured as a clear question–answer pair, making it directly suitable for fine-tuning or instruction-tuning language models that need to handle sensitive, empathetic, and contextually aware dialogue. Since its public release in 2023, it has been downloaded over 100,000… See the full description on the dataset page: https://huggingface.co/datasets/Amod/mental_health_counseling_conversations.texttext-generation1K<n<10K502 likes1.5k downloads10mo agoHugging Face03lintw /VL-Health VL-Health Dataset Overview The VL-Health dataset is designed for multi-stage training of unified LVLMs in the medical domain. It consists of two key phases: Alignment – Focused on training image captioning capabilities and learning representations of input visual information. Instruct Fine-Tuning – Designed for enhancing the model's ability to handle various vision-language tasks, including both visual comprehension and visual generation tasks. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/lintw/VL-Health.textquestion-answering100K<n<1M18 likes1.2k downloads2y agoHugging Face04Larxel /healthqa-br HealthQA-BR Resumo O HealthQA-BR é o primeiro benchmark de larga escala e abrangência para todo o Sistema Único de Saúde (SUS), projetado para medir o conhecimento clínico de Grandes Modelos de Linguagem (LLMs) frente aos desafios da saúde pública brasileira. Composto por 5.632 questões de múltipla escolha, o conjunto de dados é derivado de provas e concursos de licenciamento profissional e residência de abrangência nacional e de alto impacto no Brasil. Diferentemente de… See the full description on the dataset page: https://huggingface.co/datasets/Larxel/healthqa-br.textquestion-answering1K<n<10K4 likes828 downloads1y agoHugging Face05katielink /healthsearchqa HealthSearchQA Dataset of consumer health questions released by Google for the Med-PaLM paper (arXiv preprint). From the paper: We curated our own additional dataset consisting of 3,173 commonly searched consumer questions, referred to as HealthSearchQA. The dataset was curated using seed medical conditions and their associated symptoms. We used the seed data to retrieve publicly-available commonly searched questions generated by a search engine, which were displayed to all users… See the full description on the dataset page: https://huggingface.co/datasets/katielink/healthsearchqa.textquestion-answering1K<n<10K22 likes738 downloads3y agoHugging Face06UVSKKR /Ethical-Reasoning-in-Mental-Health-v1gatedThis repository contains the dataset for the paper EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI. Overview Ethical-Reasoning-in-Mental-Health-v1 (EthicsMH) is a carefully curated dataset focused on ethical decision-making scenarios in mental health contexts.This dataset captures the complexity of real-world dilemmas faced by therapists, psychiatrists, and AI systems when navigating critical issues such as confidentiality, autonomy, and bias. Each sample… See the full description on the dataset page: https://huggingface.co/datasets/UVSKKR/Ethical-Reasoning-in-Mental-Health-v1.textquestion-answeringn<1K4 likes555 downloads1y agoHugging Face07yesilhealth /Health_Benchmarks LLM Health Benchmarks Dataset by Yesil Science The LLM Health Benchmarks Dataset is a specialized resource for evaluating large language models (LLMs) in different medical specialties. It provides structured question-answer pairs designed to test the performance of AI models in understanding and generating domain-specific knowledge. Primary Purpose This dataset is built to: Benchmark LLMs in medical specialties and subfields. Assess the accuracy and contextual… See the full description on the dataset page: https://huggingface.co/datasets/yesilhealth/Health_Benchmarks.textquestion-answering1K<n<10K10 likes445 downloads1y agoHugging Face08yahskapar /HealthChat-11K HealthChat-11K This repository contains HealthChat-11K, a curated dataset of approximately 11,000 real-world conversations, composed of 25,000 user messages, where users seek healthcare information from Large Language Models (LLMs). The goal of this work is to provide a high-quality resource for systematically studying and improving health conversations involving humans and AI (e.g., LLMs). The dataset was presented in the paper: "What's Up, Doc?": Analyzing How Users Seek Health… See the full description on the dataset page: https://huggingface.co/datasets/yahskapar/HealthChat-11K.tabulartext-classification10K<n<100K12 likes375 downloads5mo agoHugging Face09NepaliAI /Nepali-HealthChattextquestion-answering10K<n<100K3 likes353 downloads3y agoHugging Face10forever-healthy /evipedia-reviews Evipedia Evidence Reviews The full public catalogue of evipedia.ai — a continuously-updated encyclopedia of evidence reviews on health & longevity interventions — as one record per review. Each record carries the review's metadata plus its complete Markdown body. Homepage / source: https://evipedia.ai Live file: https://evipedia.ai/evipedia-corpus.jsonl (this dataset mirrors it) Publisher: Forever Healthy License: CC BY 4.0 What's inside One JSON object per… See the full description on the dataset page: https://huggingface.co/datasets/forever-healthy/evipedia-reviews.texttext-retrievaln<1K2 likes320 downloads5h agoHugging Face11li-lab /HealMed HealMed (Human-verified Evaluation Across Languages for Medical AI) is a multilingual medical dataset featuring expert-verified translations for benchmarking multilingual medical AI systems. The dataset comprises translations from two complementary sources. A portion is based on the multilingual translations released by the GlobMed project (arXiv: 2601.02186), while the remainder was generated by our team using zero-shot machine translation to expand language coverage. Each translated… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/HealMed.textquestion-answering10K<n<100K2 likes300 downloads11d agoHugging Face12BAAI /IndustryInstruction_Health-Medicine IndustryInstruction: Health & Medicine This repository contains the IndustryInstruction: Health & Medicine domain subset of BAAI/IndustryInstruction. Refer to the parent dataset card for data construction, intended use, limitations, and licensing details. Citation If you use this dataset in your work, please cite IndustryInstruction: @misc{shi2024industryinstruction, title = {IndustryInstruction}, author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Health-Medicine.textquestion-answering100K<n<1M9 likes222 downloads1mo agoHugging Face13healthlifestyle /OpenMathReasoning OpenMathReasoning OpenMathReasoning is a large-scale math reasoning dataset for training large language models (LLMs). This dataset contains 306K unique mathematical problems sourced from AoPS forums with: 3.2M long chain-of-thought (CoT) solutions 1.7M long tool-integrated reasoning (TIR) solutions 566K samples that select the most promising solution out of many candidates (GenSelect) Additional 193K problems sourced from AoPS forums (problems only, no solutions) We used… See the full description on the dataset page: https://huggingface.co/datasets/healthlifestyle/OpenMathReasoning.textquestion-answering1M<n<10M0 likes220 downloads3mo agoHugging Face14MedSwin /HealthCareMagic-DistilledThis dataset undergone: Comprehensive data augmentation pipeline, Soft/hard label distillation from MedGemma-27B-Text-IT textquestion-answering10K<n<100K0 likes151 downloads8mo agoHugging Face15Ahmed-Selem /Shifaa_Arabic_Mental_Health_Consultations 🏥 Shifaa Arabic Mental Health Consultations 🧠 📌 Overview Shifaa Arabic Mental Health Consultations is a high-quality dataset designed to advance Arabic medical language models.This dataset provides 35,648 real-world medical consultations, covering a wide range of mental health concerns. 📊 Dataset Summary Size: 35,648 consultations Main Specializations: 7 Specific Diagnoses: 123 Languages: Arabic (العربية) Why This Dataset? 🔹 Lack of… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Mental_Health_Consultations.textquestion-answering10K<n<100K14 likes146 downloads2y agoHugging Face16Tonic /Health-Bench-Eval-OSS-2025-07 Dataset Card for HealthBench Dataset Summary HealthBench is a benchmark dataset developed by OpenAI in collaboration with 262 physicians from 60 countries to evaluate AI systems in health-related conversational scenarios. It contains 5,000 multi-turn health conversations in a JSONL file (2025-05-07-06-14-12_oss_eval.jsonl), simulating interactions between AI models and users (laypersons or clinicians). Each conversation includes a user prompt, a candidate model response… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/Health-Bench-Eval-OSS-2025-07.texttext-generation1K<n<10K4 likes140 downloads1y agoHugging Face17ShivomH /Mental-Health-Conversations Dataset Card This dataset consists of around 99k rows of mental health conversations. It is a cleaned version of "jerryjalapeno/nart-100k-synthetic". Source jerryjalapeno/nart-100k-synthetic texttext-generation10K<n<100K4 likes106 downloads1y agoHugging Face18NoirHuy /chatdoctor-healthcaremagic-112k-vi 🩺 ChatDoctor HealthCareMagic 112k (Vietnamese Translated) Tập dữ liệu hỏi đáp y khoa ChatDoctor HealthCareMagic 112k được dịch sang tiếng Việt chất lượng cao, phục vụ fine-tune các mô hình ngôn ngữ lớn (LLM) trong lĩnh vực y tế, chăm sóc sức khỏe và tư vấn y khoa tổng quát. 📌 Tổng quan dữ liệu Quy mô: 112,165 cặp hỏi - đáp y tế thực tế giữa bệnh nhân và bác sĩ. Phân chia: train: 106,556 mẫu (95%) validation: 5,609 mẫu (5%) Định dạng: Chuẩn Alpaca /… See the full description on the dataset page: https://huggingface.co/datasets/NoirHuy/chatdoctor-healthcaremagic-112k-vi.textquestion-answering100K<n<1M1 likes103 downloads7d agoHugging Face19emgena /omnimcp_healthtech_medops_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_healthtech_medops_teaser.texttext-generationn<1K1 likes102 downloads10d agoHugging Face20ai-enthusiasm-community /vietnamese_health_dataset Team and Homepage Official Website: https://aienthusiasm.vn Hugging Face Organization: https://huggingface.co/ai-enthusiasm-community Contact If you encounter any issues with the dataset or have any inquiries, please feel free to reach out to us via email at: aienthusiasm.team@gmail.com Dataset Structure The dataset is provided in a flattened tabular format, optimized for the Hugging Face Dataset Viewer and high-speed Parquet processing.… See the full description on the dataset page: https://huggingface.co/datasets/ai-enthusiasm-community/vietnamese_health_dataset.texttranslation100K<n<1M0 likes95 downloads4mo agoHugging Face21hllzmz /synthetic-mental-health-convos Synthetic Mental Health SFT Dataset Dataset Summary This dataset contains high-fidelity, synthetic patient-therapist dialogues designed for Supervised Fine-Tuning (SFT) of Large Language Models (LLMs) in the domain of mental health. The primary goal of this dataset is to train AI assistants to transition from "general knowledge" models to empathetic, supportive, and safety-conscious mental health companions. The dialogues cover a wide spectrum of mental health conditions… See the full description on the dataset page: https://huggingface.co/datasets/hllzmz/synthetic-mental-health-convos.texttext-generation1K<n<10K1 likes82 downloads10mo agoHugging Face22MedSwin /HealthCareMagicThis dataset undergone a comprehensive data augmentation pipeline. textquestion-answering100K<n<1M1 likes78 downloads8mo agoHugging Face23emgena /omnimcp_healthtech_hipaa_redactor_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_healthtech_hipaa_redactor_teaser.texttext-generationn<1K0 likes71 downloads10d agoHugging Face24emgena /omnimcp_healthtech_hl7_parser_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_healthtech_hl7_parser_teaser.texttext-generationn<1K0 likes67 downloads10d agoHugging Face25gvic-unb /public-health-news-DF-qa Public Health Brasília QA Corpus Dataset Summary Public Health Brasília QA Corpus is a dataset for evaluating Retrieval-Augmented Generation (RAG) systems over public health news articles from the Secretaria de Saúde do Distrito Federal (SES-DF), Brazil. It consists of two components: a QA evaluation corpus with location-aware question-answer pairs, and a knowledge base corpus of 1,688 public health news articles that serves as the retrieval source for the RAG… See the full description on the dataset page: https://huggingface.co/datasets/gvic-unb/public-health-news-DF-qa.textquestion-answering1K<n<10K0 likes65 downloads26d agoHugging Face26emgena /omnimcp_healthtech_fhir_validator_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_healthtech_fhir_validator_teaser.texttext-generationn<1K0 likes64 downloads10d agoHugging Face27emgena /omnimcp_healthtech_cohort_query_teaser 🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE: Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20! 📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_healthtech_cohort_query_teaser.texttext-generationn<1K0 likes62 downloads10d agoHugging Face28FunDialogues /healthcare-minor-consultation This Dialogue Comprised of fictitious examples of dialogues between a doctor and a patient during a minor medical consultation.. Check out the example below: "id": 1, "description": "Discussion about a common cold", "dialogue": "Patient: Doctor, I've been feeling congested and have a runny nose. What can I do to relieve these symptoms?\n\nDoctor: It sounds like you have a common cold. You can try over-the-counter decongestants to relieve congestion and saline nasal sprays to help… See the full description on the dataset page: https://huggingface.co/datasets/FunDialogues/healthcare-minor-consultation.tabularquestion-answeringn<1K5 likes61 downloads3y agoHugging Face29MaggiePai /mental_health_counseling_conversations Amod/mental_health_counseling_conversations This data is cloned from https://huggingface.co/datasets/Amod/mental_health_counseling_conversations Dataset Summary This dataset is a collection of questions and answers sourced from two online counseling and therapy platforms. The questions cover a wide range of mental health topics, and the answers are provided by qualified psychologists. The dataset is intended to be used for fine-tuning language models to improve their… See the full description on the dataset page: https://huggingface.co/datasets/MaggiePai/mental_health_counseling_conversations.texttext-generation1K<n<10K3 likes61 downloads2y agoHugging Face30hetanshwaghela /autoscientist-healthcare-reasoning 🩺 Adapted Healthcare Clinical-Reasoning (AutoScientist) Built with Adaptive Data by Adaption. A grounded, safety-blueprinted clinical-reasoning dataset — and a rigorous, fully-reproducible study of when data adaptation helps a small model, and when it doesn't. 📈 Adaptive Data quality Before → After Overall quality score 7.0 → 9.1 (+30%) Quality grade B → A Completion quality +37.9% Message quality +17.6% Percentile vs. reference corpus 15.3 → 33.0… See the full description on the dataset page: https://huggingface.co/datasets/hetanshwaghela/autoscientist-healthcare-reasoning.texttext-generation10K<n<100K0 likes60 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.