CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zabir1996 /mimic-medical-imaging-qa MIMIC Medical Imaging QA Dataset 5,207 Bloom's-taxonomy-stratified question--answer pairs derived from 23 medical imaging lectures (RPI BMED 2300). The dataset supports the paper "MIMIC: A Course-Derivation Pipeline and Benchmark for Slide-Anchored Tutoring with a Domain-Adapted Large Language Model" and was used to fine-tune MIMIC-LM, a domain-adapted Llama-3.1-8B-Instruct model for grounded medical imaging instruction. License The benchmark annotations, dataset… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/mimic-medical-imaging-qa.imagequestion-answering1K<n<10K3 likes800 downloads5mo agoHugging Face02microsoft /mediflow MediFlow A large-scale synthetic instruction dataset of 2.5M rows (~700k unique instructions) for clinical natural language processing covering 14 task types and 98 fine-grained input clinical documents. t-SNE 2D Plot of MediFlow Embeddings by Task Types Dataset Splits mediflow: 2.5M instruction data for SFT alignment. mediflow_dpo: ~135k top-quality instructions with GPT-4o generated rejected_output for DPO alignment. Main Columns instruction:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/mediflow.tabulartext-generation1M<n<10M53 likes290 downloads9mo agoHugging Face03joecwales /whiteglove-medical-medlineplus-2025 WhiteGlove Medical Knowledge Corpus MedlinePlus 2025 — Spectral Curation Pipeline Pipeline: WhiteGlove Spectral Curation | Domain: Medical | License: Public Domain (US Government) Dataset Summary A clean, deduplicated, semantically chunked medical knowledge corpus derived from the NIH MedlinePlus January 2025 ZIM archive. Produced by the WhiteGlove Spectral Curation Pipeline — an air-gapped, attribution-clean dataset factory built on SimHash-128 deduplication… See the full description on the dataset page: https://huggingface.co/datasets/joecwales/whiteglove-medical-medlineplus-2025.tabulartext-generation1K<n<10K0 likes279 downloads4mo agoHugging Face04Alaamer /medium-articles-posts-with-content Medium Articles Dataset Generator This project combines multiple datasets from Kaggle and Hugging Face to create a comprehensive collection of Medium articles. The combined dataset is available on Hugging Face Hub. Dataset Description This dataset is a unique compilation that not only combines multiple sources but also ensures data quality through normalization and deduplication. A key feature is that all entries in the text column are unique - there are no duplicate… See the full description on the dataset page: https://huggingface.co/datasets/Alaamer/medium-articles-posts-with-content.tabulartext-classification100K<n<1M3 likes217 downloads2y agoHugging Face05drkolesnikov /russian-nmo-medical-mcq Russian NMO Medical MCQ Choose language / Выберите язык: Русский | English Русский Это датасет русскоязычных медицинских тестовых вопросов НМО с вариантами ответа. В нем есть вопросы с одним правильным вариантом и вопросы с несколькими правильными вариантами. Датасет подготовлен так, чтобы его можно было сразу использовать для тонкой настройки LLM, проверки качества ответов и экспериментов с медицинским QA. Главная идея простая: дать модели вопрос, тему и варианты ответа… See the full description on the dataset page: https://huggingface.co/datasets/drkolesnikov/russian-nmo-medical-mcq.tabularquestion-answering1M<n<10M2 likes171 downloads4mo agoHugging Face06BrightData /IMDb-Media Dataset Card for "BrightData/IMDb-Media" Dataset Summary Explore feature films, TV series, episodes, mini-series, documentaries, and more with this IMDb dataset, comprising over 249K structured records and 32 data fields updated and refreshed regularly. Each entry includes all major data points such as timestamp, title, URLs, release date, IMDb rating, reviews, awards, origin, category/genre, budget, cast, director, images, videos and more. For a complete list of data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/IMDb-Media.tabulartext-classification100K<n<1M9 likes154 downloads2y agoHugging Face07TimotheeB /triage-medical-dataset Dataset release Version: 2026-03-19-v1 Published at: 2026-03-19T15:42:16+00:00 Repo: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset Dataset Card - POC Triage Medical Fiche unifiee: inventaire des sources, strategie de selection, schema, gouvernance. 1) Description Dataset bilingue FR/EN pour triage medical initial. Le pipeline produit deux artefacts principaux: SFT: paires instruction/reponse pour le fine-tuning supervise. DPO: paires… See the full description on the dataset page: https://huggingface.co/datasets/TimotheeB/triage-medical-dataset.tabulartext-generation10K<n<100K0 likes111 downloads6mo agoHugging Face08shaikat005 /medium-web-pentesting Medium Web Pentesting Articles Dataset Description A curated collection of 357 Medium articles focused on web penetration testing, scraped from Medium's search results for the query web pentesting. Each record includes article metadata and the opening snippet of the article body. This dataset is useful for NLP tasks such as topic modeling, text classification, content recommendation, and summarization within the cybersecurity domain. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/shaikat005/medium-web-pentesting.tabulartext-classificationn<1K0 likes109 downloads5mo agoHugging Face09Suhyunlee /cc-mediation CC-Mediation A cross-cultural conflict-mediation benchmark grounded in the Developmental Model of Intercultural Sensitivity (DMIS). Each scenario is a culturally grounded conflict dialogue with a mediation intervention and its post-intervention trajectory, organised as a preference pair (positive vs negative continuation) so downstream effects are measurable. 📄 Paper: CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation (Lee, Zhang, Yow, Deng;… See the full description on the dataset page: https://huggingface.co/datasets/Suhyunlee/cc-mediation.tabulartext-generation1K<n<10K0 likes97 downloads18d agoHugging Face10Metaskepsis /Olympiads_medium Numina-Olympiads Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers. Dataset Information Split: train Original size: 13284 Filtered size: 13240 Source: olympiads All examples contain valid boxed answers Dataset Description This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes: A mathematical word problem A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Olympiads_medium.tabulartext-generation10K<n<100K1 likes93 downloads2y agoHugging Face11DimitarV /bulgarian-medical-cpt-100m Bulgarian text for MOSS continued pretraining Exactly 100 million training tokens: 10M medical and 90M general Bulgarian. An additional 100,000 tokens are provided for validation (50k per source). No audio, instruction-response pairs, or generated answers. No model training has been performed as part of this dataset build. Medical data Exactly 10,000,000 training tokens and 50,000 additional validation tokens, including one <|im_end|> EOS per record. Extracted… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-medical-cpt-100m.tabulartext-generation100K<n<1M0 likes88 downloads14d agoHugging Face12Metaskepsis /Numina_medium Numina-Olympiads Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers. Dataset Information Split: train Original size: 37133 Filtered size: 37133 Source: olympiads All examples contain valid boxed answers Dataset Description This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes: A mathematical word problem A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Numina_medium.tabulartext-generation10K<n<100K0 likes83 downloads2y agoHugging Face13Saminx22 /medical_data_for_slm 🏥 Medical SLM Pretraining Dataset Card This dataset is a high-quality, cleaned collection of medical text designed for pretraining small language models (SLMs). It aggregates data from three primary authoritative sources, focusing on general medicine and clinical guidelines. 📊 Dataset Summary Total Documents: ~44,400 Estimated Tokens: ~44.7 Million Primary Language: English Configurations: documents: Raw cleaned text records. chunks: Tokenized and packed 1024-token… See the full description on the dataset page: https://huggingface.co/datasets/Saminx22/medical_data_for_slm.tabulartext-generation10K<n<100K1 likes66 downloads6mo agoHugging Face14DimitarV /bulgarian-medical-cpt-10m Bulgarian text for MOSS continued pretraining Exactly 10 million training tokens: 3M medical and 7M general Bulgarian. An additional 100,000 tokens are provided for validation (50k per source). No audio, instruction-response pairs, or generated answers. No model training has been performed as part of this dataset build. Medical data Exactly 3,000,000 training tokens and 50,000 additional validation tokens, including one <|im_end|> EOS per record. Extracted… See the full description on the dataset page: https://huggingface.co/datasets/DimitarV/bulgarian-medical-cpt-10m.tabulartext-generation10K<n<100K0 likes64 downloads14d agoHugging Face15JWei05 /Nemotron-Math-v2-Medium-10k Nemotron-Math-v2-Medium-10k A lightweight 10,500-problem subset of nvidia/Nemotron-Math-v2 for long-horizon Python-TIR reinforcement learning. It contains 1,500 problems from each metadata.reason_high_with_tool.pass bucket 1 through 7. A deterministic seed-42 shuffle assigns 500 examples to validation and 10,000 to train. This Hugging Face release intentionally contains no teacher traces. The full messages/tools aggregation is retained as a separate local artifact.… See the full description on the dataset page: https://huggingface.co/datasets/JWei05/Nemotron-Math-v2-Medium-10k.tabulartext-generation10K<n<100K0 likes60 downloads1mo agoHugging Face16room-b007 /test-medicina Medschool-Test, or "Test di Medicina" Is your LLM able to pass a National Entrance Exam for the Italian Medical School? This the GitHub repo for our Hugging Face dataset designed for evaluating Large Language Models (LLMs) on a broad range of questions from the national entrance exams for the Italian medical school (ORIGINAL WEBSITE). The dataset includes multiple-choice questions from various subjects such as biology, chemistry, physics, mathematics, world… See the full description on the dataset page: https://huggingface.co/datasets/room-b007/test-medicina.tabulartext-generation1K<n<10K4 likes54 downloads2y agoHugging Face17crawlfeeds /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes38 downloads6mo agoHugging Face18khopilot /khmer-medical-qa Khmer Medical Q&A Dataset Update Notice ✅ Dataset Updated (Aug 15, 2024): We identified and fixed an alignment issue. Everything is working properly now! Dataset Description This dataset contains 18,756 high-quality medical question-answer pairs translated from English to Khmer, designed for training medical AI assistants in the Khmer language. Features index: Sequential row index (0-18755) question_en: Medical question in English response_en:… See the full description on the dataset page: https://huggingface.co/datasets/khopilot/khmer-medical-qa.tabularquestion-answering10K<n<100K1 likes31 downloads1y agoHugging Face19crawlfeeds /Medium-Articles-Corpus Medium Articles Corpus (10K Sample) The Medium Articles Corpus is a massive, clean dataset of articles scraped from Medium.com. This sample version contains 10,000 articles + and is designed to showcase the quality and structure of the full corpus for researchers and developers. This is the subset from the large dataset https://crawlfeeds.com/websites/medium/text_data/medium_articles Dataset Features This dataset includes the following key features, provided in a… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medium-Articles-Corpus.imagetext-classification10K<n<100K2 likes30 downloads1y agoHugging Face20IAMRonHIT /MediFlowThinks MediFlow A large-scale synthetic instruction dataset of 2.5M rows (~700k unique instructions) for clinical natural language processing covering 14 task types and 98 fine-grained input clinical documents. t-SNE 2D Plot of MediFlow Embeddings by Task Types Dataset Splits mediflow: 2.5M instruction data for SFT alignment. mediflow_dpo: ~135k top-quality instructions with GPT-4o generated rejected_output for DPO alignment. Main Columns instruction:… See the full description on the dataset page: https://huggingface.co/datasets/IAMRonHIT/MediFlowThinks.tabulartext-generation1M<n<10M1 likes29 downloads8mo agoHugging Face21dhruveshpatel /star-mediumSynthetic data for the paper [2505.05755] Insertion Language Models: Sequence Generation with Arbitrary-Position Insertions. Project page: https://dhruveshp.com/projects/ilm tabulartext-generation10K<n<100K0 likes25 downloads1y agoHugging Face22RyanHalliwell /IMDb-Media Dataset Card for "BrightData/IMDb-Media" Dataset Summary Explore feature films, TV series, episodes, mini-series, documentaries, and more with this IMDb dataset, comprising over 249K structured records and 32 data fields updated and refreshed regularly. Each entry includes all major data points such as timestamp, title, URLs, release date, IMDb rating, reviews, awards, origin, category/genre, budget, cast, director, images, videos and more. For a complete list of… See the full description on the dataset page: https://huggingface.co/datasets/RyanHalliwell/IMDb-Media.tabulartext-classification100K<n<1M0 likes23 downloads2mo agoHugging Face23recogna-nlp /drbodebench_medicamentos Medication-Focused Clinical Benchmark from DrBodeBench Dataset Details To evaluate retrieval capabilities in higher-level reasoning scenarios, we created a second benchmark derived from the Portuguese medical benchmark DrBodeBench. This benchmark aggregates questions from Brazilian medical examinations, including the Revalida and the FUVEST direct-access residency exam. From DrBodeBench, we curated a specific subset of questions that exclusively pertains to… See the full description on the dataset page: https://huggingface.co/datasets/recogna-nlp/drbodebench_medicamentos.tabularquestion-answeringn<1K0 likes17 downloads6mo agoHugging Face24JingweiNi /ocr2_cf1900_k2_gpt55_medium_qwen35_error_steps_seed20260513 GPT-5.5 Medium Reannotation of Qwen3.5-Positive OCR2 Coding Steps This dataset follows the same 500-row parquet layout as JingweiNi/ocr2_cf1900_k2_qwen35_fp8_10k_seed20260513 and contains GPT-5.5 medium-reasoning reannotations for the 1,536 Qwen3.5-positive error steps. Summary Source dataset: JingweiNi/ocr2_cf1900_k2_qwen35_fp8_10k_seed20260513 Source rows: 500 K2-Think Codeforces traces Source manifest-selected Qwen3.5 labels: 10,000 steps GPT-5.5 reannotated… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ocr2_cf1900_k2_gpt55_medium_qwen35_error_steps_seed20260513.tabulartext-generationn<1K0 likes17 downloads4mo agoHugging Face25david-sprague /Medical-Health-QA-Articles-Dataset Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development. Dataset Overview Field Details Sources iCliniq, HealthTap, WebMD Total Records 1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.imagetext-classification1K<n<10K0 likes17 downloads4mo agoHugging Face26Vrda /real-world-medical-mistakes-dataset Real-World Medical Mistakes Dataset A curated dataset of 100 de-identified clinical reports from Internal Medicine and Emergency Departments, each containing a physician-inserted realistic medical error. Designed for training and evaluating AI systems that detect critical patient safety errors in clinical documentation. Dataset Description Overview This dataset was created as part of the Clinipal project — an AI-powered clinical error detection system. Three… See the full description on the dataset page: https://huggingface.co/datasets/Vrda/real-world-medical-mistakes-dataset.tabulartext-classificationn<1K1 likes14 downloads7mo agoHugging Face27SPAISS6F1 /spai-ss6-corpus-medical-o1-verifiable SPAI SS6 Medical O1 Verifiable Thai Index Index repo for the imported Thai medical verifiable-problem dataset config. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: medical_o1_verifiable_problem_thai Rows in canonical config: 40,906 Parquet size in canonical config: 0.00 GB… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-medical-o1-verifiable.tabulartext-generationn<1K0 likes14 downloads4mo agoHugging Face28SPAISS6F1 /medicine medicine Thai public medical and health web corpus collected for research and LLM dataset experimentation. Dataset Contents Split: train Records: 3035 deduplicated articles Format: Parquet Latest collection profile: free_1000 Latest generated at: 2026-06-06T16:25:20.801225+00:00 Source And Method URLs are collected from public sitemap XML files on configured Thai public sources, then crawled with robots.txt checks, rate limiting, Thai text… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/medicine.tabulartext-generation1K<n<10K0 likes12 downloads4mo agoHugging Face29SPAISS6F1 /spai-ss6-corpus-medical-health-web SPAI SS6 Thai Medical Health Web Corpus Thai public medical and health web articles collected by the local scraping pipeline. This is a lightweight index dataset repo. It does not duplicate the full corpus. The full Parquet data lives in the canonical repository config below. Canonical Data Canonical repo: SPAISS6F1/spai-ss6-llm-1b-thai-corpus Canonical config: default Rows in canonical config: 3,660 Parquet size in canonical config: 0.01 GB Source license: other… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/spai-ss6-corpus-medical-health-web.tabulartext-generationn<1K0 likes11 downloads4mo agoHugging Face30c00cjz00 /Medical-R1-Distill-Data-m1k Medical-R1-Distill-Data tabulartext-generationn<1K0 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.