datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vidore_v3_pharmaceuticalsViDoRe V3 : Pharmaceuticals
This dataset, Pharmaceutical, is a corpus of slides from the FDA, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with human-verified relevant pages… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_pharmaceuticals.vidore_v3_pharmaceuticals_mteb_format
Vidore3PharmaceuticalsRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions.
Task category
t2i
Domains
Academic
Reference
https://huggingface.co/blog/QuentinJG/introducing-vidore-v3
Source datasets:
vidore/vidore_v3_pharmaceuticals
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_pharmaceuticals_mteb_format.pharmaconerPharmaCoNER: Pharmacological Substances, Compounds and Proteins Named Entity Recognition track
This dataset is designed for the PharmaCoNER task, sponsored by Plan de Impulso de las Tecnologías del Lenguaje (Plan TL).
It is a manually classified collection of clinical case studies derived from the Spanish Clinical Case Corpus (SPACCC), an
open access electronic library that gathers Spanish medical publications from SciELO (Scientific Electronic Library Online).
The annotation of the entire set of entity mentions was carried out by medicinal chemistry experts
and it includes the following 4 entity types: NORMALIZABLES, NO_NORMALIZABLES, PROTEINAS and UNCLEAR.
The PharmaCoNER corpus contains a total of 396,988 words and 1,000 clinical cases that have been randomly sampled into 3 subsets.
The training set contains 500 clinical cases, while the development and test sets contain 250 clinical cases each.
In terms of training examples, this translates to a total of 8074, 3764 and 3931 annotated sentences in each set.
The original dataset was distributed in Brat format (https://brat.nlplab.org/standoff.html).
For further information, please visit https://temu.bsc.es/pharmaconer/ or send an email to encargo-pln-life@bsc.espharma-bench-mlmPharmaShip
PharmaShip: An Entity-Centric, Reading-Order-Supervised Benchmark for Chinese Pharmaceutical Shipping Documents
🔗 Paper: https://arxiv.org/abs/2512.23714
🔗 Github: https://github.com/KevinYuLei/PharmaShip
Description
PharmaShip is a real-world Chinese dataset of scanned pharmaceutical shipping documents designed to stress-test pre-trained text-layout models under noisy OCR and heterogeneous templates.
It covers three complementary tasks:
Sequence Entity Recognition… See the full description on the dataset page: https://huggingface.co/datasets/YuLeiKevin/PharmaShip.pharma-kb-obesity
Pharma KB — Obesity Drug Landscape
A structured pharmaceutical knowledge base covering the obesity drug pipeline — compiled from public regulatory, clinical, and scientific sources. Designed for RAG pipelines, competitive intelligence workflows, and pharma-domain LLM fine-tuning.
Overview
Indication
Obesity (ICD: E66)
Drug articles
285 (all phases, launched through discovery)
Company articles
248
Target articles
35
Full content size
~11 MB… See the full description on the dataset page: https://huggingface.co/datasets/OpenPharma/pharma-kb-obesity.2023_Pharmacist_Licensure_Examination-TCM_trackThe 2023 Chinese National Pharmacist Licensure Examination is divided into two distinct tracks: the Pharmacy track and the Traditional Chinese Medicine (TCM) Pharmacy track. The data provided here pertains to the Traditional Chinese Medicine (TCM) Pharmacy track examination. It is important to note that this dataset was collected from online sources, and there may be some discrepancies between this data and the actual examination.
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/2023_Pharmacist_Licensure_Examination-TCM_track.ICD-10_Pharmacy
Medical train/valid V1 — Viettel AI Race clinical concept extraction
What is included
joint_train.jsonl: 12,000 synthetic multi-concept clinical notes with full competition-style labels.
joint_valid_template.jsonl: 1,500 notes using held-out templates but concepts seen in train.
joint_valid_concept.jsonl: 1,500 notes using held-out templates and held-out concepts.
ner_*.jsonl: same notes, stripped to text + type + position for token/span NER.
assertion_*.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/leminhhung0101/ICD-10_Pharmacy.vdac-pharmacology-atlas
VDAC1 Pharmacology Atlas — Multi-LLM Convergence Dataset
The first publicly available multi-LLM scientific convergence dataset.
32 IRIS Gate Evo runs, 200+ synthesized claims across 5 independent AI models, 23 curated gold extractions, and cross-run convergence analysis — all from a single research program mapping the pharmacology of life's decision gate.
Paper: bioRxiv BIORXIV/2026/706165
Repository: templetwo/vdac-pharmacology-atlas
Authors: Anthony J. Vasquez Sr. (Delaware Valley… See the full description on the dataset page: https://huggingface.co/datasets/TheTempleofTwo/vdac-pharmacology-atlas.nursing-pharmacology
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/timzhou99/nursing-pharmacology.FDA_Pharmaceuticals_FAQ
FDA Pharmaceutical Q&A Dataset
Description
This dataset contains a collection of question-and-answer pairs related to pharmaceutical regulatory compliance provided by the Food and Drug Administration (FDA). It is designed to support research and development in the field of natural language processing, particularly for tasks involving information retrieval, question answering, and conversational agents within the pharmaceutical domain.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/Jaymax/FDA_Pharmaceuticals_FAQ.mined_negatives_pharma_qaindian-pharma-dataset-2026-augast
PharmaLens: 200k Medicine Catalog (Salts, Prices, Interactions and Reviews)
I spent weeks compiling and cleaning this retrieval database for a project. Instead of letting 300MB+ of structured pharmaceutical data sit idle on my hard drive, I am open-sourcing it. Use it for your RAG pipelines, chatbots, pricing tools, or whatever else you are building.
Overview
Finding clean, structured pharmaceutical datasets with commercial brand names, active salt compositions… See the full description on the dataset page: https://huggingface.co/datasets/sinhal/indian-pharma-dataset-2026-augast.pharmaceutical-procurement-pricing
Pharmaceutical Procurement & Pricing (Price Ratios, Quality Assurance, Lead Times, Savings) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: not declared - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/pharmaceutical-procurement-pricing.vn-provinces-doh-pharmacy-workforce
Vietnam DOH pharmacy workforce by qualification
Vietnam DOH pharmacy workforce by qualification. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (1008 rows)
data/provinces.csv
data/provinces.dta
data/provinces.xlsx
regions (95 rows)… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-doh-pharmacy-workforce.Stocks-Weekly-PharmaClinicalPredict
Stocks Weekly PharmaClinicalPredict
Modelled outcome, duration and economic impact for pharmaceutical clinical trials, one row per trial.
91,981 rows, 9 columns. Updated by Papers With Backtest.
Why It Matters
A biotech's value is a probability-weighted pipeline, and the probabilities are what this dataset estimates:
Event odds ahead of the readout: success_prediction is a modelled probability that a trial reaches its endpoint. Priced against the sponsor's market… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Stocks-Weekly-PharmaClinicalPredict.combined_pharma_qapharma-serialized-events
ZigoTrace Pharma — Serialized Events (synthetic)
Feature vectors extracted from a synthetic DSCSA-style serialized medicine supply
chain, generated by packages/intelligence/src/synthetic.ts in the
zigo-pharma engine and exported via
hf/generate_fixtures.mjs. Used to train and validate the diversion-detection model
(zigotrace/pharma-authenticity-model).
⚠️ Synthetic data notice
This is entirely synthetic — a seeded generator (generateChain), not real
distributor or… See the full description on the dataset page: https://huggingface.co/datasets/AsamAce/pharma-serialized-events.pharmaceutical-regulatory-capacity
Pharmaceutical Regulatory Capacity (NRA Maturity, GBT Scores, SF Prevalence) | Africa (World Health Organization)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/pharmaceutical-regulatory-capacity.praxy-gtts-pharmaceutical-audiopostmarket-surveillance-pharmacovigilance
Post-Market Surveillance & Pharmacovigilance (ADR Reporting, Causality, VigiBase) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/postmarket-surveillance-pharmacovigilance.asia-who-pharmaceutical-technicians-and-assistants
Pharmaceutical Technicians and Assistants (number) | Asia (WHO GHO)
🌏 182 observations · 28 Asia countries · 1983–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 182 observations of Pharmaceutical Technicians and Assistants (number) data across 28 Asia countries, spanning 1983–2024, covering 1 distinct indicators.
About the source
Source: WHO Global Health Observatory
Publisher: World Health Organization
License:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-who-pharmaceutical-technicians-and-assistants.pharma-preference-dataset
Pharma DPO Preference Dataset
Pharmaceutical domain preference dataset used for
Direct Preference Optimization (DPO) — Stage 3 of the pharma TinyLlama
fine-tuning pipeline.
Format
Each JSONL record contains 3 fields:
{
"prompt": "### Instruction:\nExplain the mechanism of metformin.\n\n### Response:\n",
"chosen": "Metformin primarily works by ...",
"rejected": "Metformin is a drug that ..."
}
prompt — Alpaca-style instruction prompt (same format as… See the full description on the dataset page: https://huggingface.co/datasets/ThakrePranjal/pharma-preference-dataset.Pharmacology-QAclinical-prescription-pharmacy-dispense-coherence-risk-v0.1What this repo is for
Detect when
a prescription exists
but pharmacy dispense
does not happen in time
Common breaks
stockout
verification delay
clarification needed
queue delay for discharge meds
Examples you can use
urgent anticoagulant delayed
antibiotic not dispensed due to stockout
TTO delayed so discharge stalls
You use it to flag
missed dose risk
discharge delay risk
indian-pharma-dataHiro-Pharma-RAG-Benchmark
Hiro Pharma RAG Benchmark
This private dataset repository contains multilingual biomedical RAG benchmark data associated with the paper CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine.
The benchmark is designed to evaluate whether retrieval-augmented language models can answer biomedical questions while selecting and citing useful evidence and filtering out noisy or irrelevant references.
Repository Contents
File
Language… See the full description on the dataset page: https://huggingface.co/datasets/PatSnap/Hiro-Pharma-RAG-Benchmark.pharma-rag-demopharmacy-ner-sft
Pharmacy NER — Drug Entity Extraction
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Clinical/biomedical text → drug name, dosage, frequency, route, indication
Why download this
Automate medication extraction from clinical notes, discharge summaries, or biomedical literature. Powers medication reconciliation and… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/pharmacy-ner-sft.pharmaconerPharmaCoNER: Pharmacological Substances, Compounds and Proteins Named Entity Recognition track
This dataset is designed for the PharmaCoNER task, sponsored by Plan de Impulso de las Tecnologías del Lenguaje.
It is a manually classified collection of clinical case studies derived from the Spanish Clinical Case Corpus (SPACCC), an open access electronic library that gathers Spanish medical publications from SciELO (Scientific Electronic Library Online).
The annotation of the entire set of entity mentions was carried out by medicinal chemistry experts and it includes the following 4 entity types: NORMALIZABLES, NO_NORMALIZABLES, PROTEINAS and UNCLEAR.
The PharmaCoNER corpus contains a total of 396,988 words and 1,000 clinical cases that have been randomly sampled into 3 subsets. The training set contains 500 clinical cases, while the development and test sets contain 250 clinical cases each.
For further information, please visit https://temu.bsc.es/pharmaconer/ or send an email to encargo-pln-life@bsc.es
