datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MedQuad-MedicalQnADataset
Reference:
"A Question-Entailment Approach to Question Answering". Asma Ben Abacha and Dina Demner-Fushman. BMC Bioinformatics, 2019.
pending-medicare-provider-enrollment-data
Pending Medicare Provider Enrollment Data
This is a dated, source-receipted sample of behavioral-health NPIs newly present in CMS's pending first-time Medicare enrollment files on 2026-07-13, compared with the immediately prior 2026-07-09 publication.
Pending does not mean approved. A row indicates that a first-time Medicare enrollment application appeared in a CMS pending file. It does not prove enrollment, credentialing, licensure, a new practice, service availability… See the full description on the dataset page: https://huggingface.co/datasets/unitedideas/pending-medicare-provider-enrollment-data.Traditional-Chinese-Medicine-Multiple_choice_question
Discription
This dataset is sourced from the website of the Ministry of Examination, R.O.C (Taiwan) and contains past exam questions from the national Traditional Chinese Medicine examinations in Taiwan. The exam comprises six subjects. This dataset specifically includes questions from two subjects, including the History of Traditional Chinese Medicine, Basic Theories of Traditional Chinese Medicine, Neijing, Nanjing, Traditional Chinese Medicine Prescription Studies, and… See the full description on the dataset page: https://huggingface.co/datasets/Liavan/Traditional-Chinese-Medicine-Multiple_choice_question.LDS-retrain-bank-adamw-N16k-bs256-gpt2-mediumShifaa_Arabic_Medical_Consultations
Shifaa Arabic Medical Consultations 🏥📊
Overview 🌍
Shifaa is revolutionizing Arabic medical AI by addressing the critical gap in Arabic medical datasets. Our first contribution is the Shifaa Arabic Medical Consultations dataset, a comprehensive collection of 84,422 real-world medical consultations covering 16 Main Specializations and 585 Hierarchical Diagnoses.
🔍 Why is this dataset important?
First large-scale Arabic medical dataset for AI applications.… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Medical_Consultations.news_media_bias_and_factuality
News Media Factual Reporting and Political Bias
Dataset introduced in the paper "Mapping the Media Landscape: Predicting Factual Reporting and Political Bias Through Web Interactions" published in the CLEF 2024 main conference.
Similar to the news media reliability dataset, this dataset consists of a collections of 4K new media domains names with political bias and factual reporting labels.
Columns of the dataset:
source: domain name
bias: the political bias label. Values: "left"… See the full description on the dataset page: https://huggingface.co/datasets/sergioburdisso/news_media_bias_and_factuality.korean-medical-dialogue-summary-datasetTraditional-Chinese-Medicine-Exam
Coming Soon...
medicines_from_zakupki_gov_ruДанные для исследования существования focal points (https://www.jstor.org/stable/3132148) в гос. закупках лекарств в России.
mbib-base
Dataset Card for Media-Bias-Identification-Benchmark
Baseline
TaskModelMicro F1Macro F1
cognitive-bias ConvBERT/ConvBERT 0.7126 0.7664
fake-news Bart/RoBERTa-T 0.6811 0.7533
gender-bias RoBERTa-T/ELECTRA 0.8334 0.8211
hate-speech RoBERTA-T/Bart 0.8897 0.7310
linguistic-bias ConvBERT/Bart 0.7044 0.4995
political-bias ConvBERT/ConvBERT 0.7041 0.7110
racial-bias ConvBERT/ELECTRA 0.8772 0.6170
text-leve-bias ConvBERT/ConvBERT 0.7697… See the full description on the dataset page: https://huggingface.co/datasets/mediabiasgroup/mbib-base.IMDb-Media
Dataset Card for "BrightData/IMDb-Media"
Dataset Summary
Explore feature films, TV series, episodes, mini-series, documentaries, and more with this IMDb dataset, comprising over 249K structured records and 32 data fields updated and refreshed regularly.
Each entry includes all major data points such as timestamp, title, URLs, release date, IMDb rating, reviews, awards, origin, category/genre, budget, cast, director, images, videos and more.
For a complete list of data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/IMDb-Media.us-media-newspaper-publishing-telecom-layoffs-warn-act-notices-daily
US media, newspaper, publishing and telecom layoffs — the actual WARN Act filings, rebuilt every day
Last rebuilt: 2026-09-24. 846 layoff and closure notices filed by
newspapers and newspaper chains, broadcasters and TV station groups, film and game studios, magazine and book publishers, commercial printers, advertising and marketing agencies, and wireless, cable and telephone carriers and their call-centre contractors with US state labor departments — 101,859 workers,
265… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-media-newspaper-publishing-telecom-layoffs-warn-act-notices-daily.medical_textaliiihussain_social-media-viral-content-and-engagement-metrics
Social Media Viral Content & Engagement Metrics
What Makes Content Go Viral? Engagement, Sentiment, and Social Trends Dataset
Dataset Info
Source: Kaggle
Original Size: 0.07 MB
Kaggle Downloads: 1,836
Files: 1
Files
social_media_viral_content_dataset.csv
Mirrored from Kaggle
cits4012_A1_2026_medical_abstractsmedquad-medical-qa
MedQuAD Medical QA Dataset
This dataset is derived from the original MedQuAD medical question–answer corpus.
It is provided in a raw / lightly structured format, intended for:
Retrieval-Augmented Generation (RAG)
LoRA fine-tuning experiments
Medical question-answering systems
Agentic document intelligence pipelines
Data Notes
No aggressive cleaning has been applied
Original question–answer semantics are preserved
Users should apply task-specific preprocessing as needed… See the full description on the dataset page: https://huggingface.co/datasets/prithvi1029/medquad-medical-qa.medicalquestions
🤗 Dataset Card: fhirfly/medicalquestions
Dataset Overview
Dataset name: fhirfly/medicalquestions
Dataset size: 25,102 questions
Labels: 1 (medical), 0 (non-medical)
Distribution: Evenly distributed between medical and non-medical questions
Dataset Description
The fhirfly/medicalquestions dataset is a collection of 25,102 questions labeled as either medical or non-medical. The dataset aims to provide a diverse range of questions covering various medical and… See the full description on the dataset page: https://huggingface.co/datasets/fhirfly/medicalquestions.combined_medical_corpus
Combined Medical QA Corpus (FUNPANG + PubMedQA + MASHQA + MedQuAD)
Dataset Owner: Mukul B
This dataset contains a unified medical QA corpus built by combining four datasets:
Dataset Name
HF Source
clustered_FUNPANG_dataset_with_groups
mukulb/clustered_FUNPANG_dataset_with_groups
pubmedQA_dataset
mukulb/pubmedQA_dataset
clustered_MASHQA_with_groups
mukulb/clustered_MASHQA_with_groups
clustered_MEDQUAD_dataset_with_groups… See the full description on the dataset page: https://huggingface.co/datasets/mukulb/combined_medical_corpus.medical-mri
ACCESS REQUIREMENT - FOLLOW TO DOWNLOAD
This dataset requires following the author to access.
How to Access
Follow @shangshang on HuggingFace: https://huggingface.co/shangshang
Request access by commenting on the dataset page
Once approved, you will receive download permissions
Usage Agreement
For research and educational purposes only
Do not redistribute without permission
Cite the dataset in your work:
@misc{shangshang_dataset_2026… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/medical-mri.MedicalTranscriptions
Medical Transcriptions
Medical transcription data scraped from mtsamples.com
Content
This dataset contains sample medical transcriptions for various medical specialties.
More information can be found here
Due to data availability only transcripts for the following medical specialties were selected for the model training:
Surgery
Cardiovascular / Pulmonary
Orthopedic
Radiology
General Medicine
Gastroenterology
Neurology
Obstetrics / Gynecology
Urology
task_categories:… See the full description on the dataset page: https://huggingface.co/datasets/tchebonenko/MedicalTranscriptions.medical-diabetes
ACCESS REQUIREMENT - FOLLOW TO DOWNLOAD
This dataset requires following the author to access.
How to Access
Follow @shangshang on HuggingFace: https://huggingface.co/shangshang
Request access by commenting on the dataset page
Once approved, you will receive download permissions
Usage Agreement
For research and educational purposes only
Do not redistribute without permission
Cite the dataset in your work:
@misc{shangshang_dataset_2026… See the full description on the dataset page: https://huggingface.co/datasets/shangshang/medical-diabetes.vn-provinces-doh-medical-workforce
Vietnam DOH medical workforce by qualification
Vietnam DOH medical workforce by qualification. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (1008 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (96 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-doh-medical-workforce.multilingual-islr-mediapipe
Multilingual ISLR MediaPipe Landmarks
Dataset Description
This dataset combines frame-level MediaPipe Holistic landmarks derived from four isolated sign language recognition (ISLR) resources: INCLUDE-50, KSL, MINDS-Libras, and LIBRAS-UFOP. It provides a common tabular schema for research on landmark selection, temporal modeling, signer-independent evaluation, and multilingual transfer learning.
The release contains landmarks rather than source RGB videos. Every… See the full description on the dataset page: https://huggingface.co/datasets/danielelvs/multilingual-islr-mediapipe.DataComp_medium_pool_BLIP2_captions
Dataset Card for DataComp_medium_pool_BLIP2_captions
Dataset Summary
Supported Tasks and Leaderboards
We have used this dataset for pre-training CLIP models and found that it rivals or outperforms models trained on raw web captions on average across the 38 evaluation tasks proposed by DataComp.
Refer to the DataComp leaderboard (https://www.datacomp.ai/leaderboard.html) for the top baselines uncovered in our work.
Languages
Primarily English.… See the full description on the dataset page: https://huggingface.co/datasets/thaottn/DataComp_medium_pool_BLIP2_captions.Multilingual_medical_symptom_triage
tags:
- medical
- healthcare
- classification
- outbreak-detection
- triage
- multilingual
- adaption
- india
Multilingual Medical Symptom Triage Dataset
Dataset Description
A Mutlilingual medical triage dataset containing 9,064 patient
symptom descriptions in Hindi, English, and Hinglish (code-mixed
Hindi-English), paired with triage recommendations and rich
clinical metadata. Designed for training multilingual triage
classification models and… See the full description on the dataset page: https://huggingface.co/datasets/Tulsiandhare/Multilingual_medical_symptom_triage.medical-imaging-it-glossary
Medical Imaging IT Glossary — PACS, RIS, DICOM, HL7 (ES/EN)
A structured, bilingual (Spanish/English) reference glossary of standards,
systems, protocols and operational concepts used in medical imaging IT and
teleradiology: 53 terms across 23 categories, covering DICOM services and
data hierarchy, HL7 v2/FHIR message types, PACS/RIS/VNA systems, imaging
modalities (CT, MRI, US, PET, mammography, etc.), interoperability profiles
(IHE), and relevant Mexican/international… See the full description on the dataset page: https://huggingface.co/datasets/NODARISHUB/medical-imaging-it-glossary.vn-provinces-social-media-users-share
Vietnam social media users share
Vietnam social media users share. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (126 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (12 rows)
data/regions.csv
data/regions.dta… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-social-media-users-share.media_campaign_costmedical-cases-classification-tutorial
About
This is a pre-filtered and pre-split dataset for the HPE Generative AI "Medical Transcript Classification" tutorials.
No-Code Version (UI Only)
Notebooks Version
Bitext-media-llm-chatbot-training-dataset
Bitext - Media Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [media] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-media-llm-chatbot-training-dataset.
