datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2023_Pharmacist_Licensure_Examination-TCM_trackThe 2023 Chinese National Pharmacist Licensure Examination is divided into two distinct tracks: the Pharmacy track and the Traditional Chinese Medicine (TCM) Pharmacy track. The data provided here pertains to the Traditional Chinese Medicine (TCM) Pharmacy track examination. It is important to note that this dataset was collected from online sources, and there may be some discrepancies between this data and the actual examination.
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/2023_Pharmacist_Licensure_Examination-TCM_track.indian-pharma-dataset-2026-augast
PharmaLens: 200k Medicine Catalog (Salts, Prices, Interactions and Reviews)
I spent weeks compiling and cleaning this retrieval database for a project. Instead of letting 300MB+ of structured pharmaceutical data sit idle on my hard drive, I am open-sourcing it. Use it for your RAG pipelines, chatbots, pricing tools, or whatever else you are building.
Overview
Finding clean, structured pharmaceutical datasets with commercial brand names, active salt compositions… See the full description on the dataset page: https://huggingface.co/datasets/sinhal/indian-pharma-dataset-2026-augast.pharma-serialized-events
ZigoTrace Pharma — Serialized Events (synthetic)
Feature vectors extracted from a synthetic DSCSA-style serialized medicine supply
chain, generated by packages/intelligence/src/synthetic.ts in the
zigo-pharma engine and exported via
hf/generate_fixtures.mjs. Used to train and validate the diversion-detection model
(zigotrace/pharma-authenticity-model).
⚠️ Synthetic data notice
This is entirely synthetic — a seeded generator (generateChain), not real
distributor or… See the full description on the dataset page: https://huggingface.co/datasets/AsamAce/pharma-serialized-events.personaplex-finetuning-pharma-data-sample
PersonaPlex Finetuning — Pharma Data Sample
A 10-example slice of the synthetic patient-support / medication
adherence dataset used to train
demegire/personaplex-finetune-pharma.
The on-disk layout below is exactly what the trainer in
emotion-machine-org/personaplex-finetune
consumes — use this as a template when building your own.
Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at
sample scale).
Layout
.
├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.pharma-preference-dataset
Pharma DPO Preference Dataset
Pharmaceutical domain preference dataset used for
Direct Preference Optimization (DPO) — Stage 3 of the pharma TinyLlama
fine-tuning pipeline.
Format
Each JSONL record contains 3 fields:
{
"prompt": "### Instruction:\nExplain the mechanism of metformin.\n\n### Response:\n",
"chosen": "Metformin primarily works by ...",
"rejected": "Metformin is a drug that ..."
}
prompt — Alpaca-style instruction prompt (same format as… See the full description on the dataset page: https://huggingface.co/datasets/ThakrePranjal/pharma-preference-dataset.Hiro-Pharma-RAG-Benchmark
Hiro Pharma RAG Benchmark
This private dataset repository contains multilingual biomedical RAG benchmark data associated with the paper CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine.
The benchmark is designed to evaluate whether retrieval-augmented language models can answer biomedical questions while selecting and citing useful evidence and filtering out noisy or irrelevant references.
Repository Contents
File
Language… See the full description on the dataset page: https://huggingface.co/datasets/PatSnap/Hiro-Pharma-RAG-Benchmark.pharmacy-ner-sft
Pharmacy NER — Drug Entity Extraction
Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs.
Built by AmareshHebbar | Studio Ilios / Humanova Minds
What this dataset does
Clinical/biomedical text → drug name, dosage, frequency, route, indication
Why download this
Automate medication extraction from clinical notes, discharge summaries, or biomedical literature. Powers medication reconciliation and… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/pharmacy-ner-sft.Pharmacy_Licensing_Exam_FAQ
💊 Pharmacy Licensing Exam FAQ (Nepali) — Dataset README
A 100% Nepali-language synthetic reasoning dataset built from a single row-pair of official Nepal Department of Health Services exam-result data — turned into 100 question-answer pairs covering counts, percentages, ratios, statistics, and "what-if" arithmetic.
🔖 TL;DR (At a Glance)
What
Answer
Total records
100
File size
~198 KB
Language / script
Nepali (ne / npi) — Devanagari
Format… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Pharmacy_Licensing_Exam_FAQ.confusable-pharma-retrieval
Confusable Pharmaceutical Product Names: A Retrieval Benchmark
Dataset Description
A retrieval benchmark for evaluating systems' ability to distinguish between pharmaceutical products with similar-sounding names. This addresses a critical challenge in pharmaceutical RAG systems where confusable drug names (e.g., "Abacavir" vs. "Abametapir") can lead to retrieval errors with regulatory and safety implications.
Summary
931 validated exclusive queries from FDA… See the full description on the dataset page: https://huggingface.co/datasets/i-was-here/confusable-pharma-retrieval.K-Paths-inductive-reasoning-pharmaDB
🔗 This dataset is part of the study:
K-Paths: Reasoning over Graph Paths for Drug Repurposing and Drug Interaction Prediction
📖 Read the Paper
💾 GitHub Repository
PharmacotherapyDB: Inductive Reasoning Dataset
PharmacotherapyDB is a drug repurposing dataset containing drug–disease treatment relations in three categories (disease-modifying, palliates, or non-indication).
Each entry includes a drug and a disease, an interaction label, drug, disease descriptions, and… See the full description on the dataset page: https://huggingface.co/datasets/Tassy24/K-Paths-inductive-reasoning-pharmaDB.2023_Pharmacist_Licensure_Examination-Pharmacy_trackThe 2023 Chinese National Pharmacist Licensure Examination is divided into two distinct tracks: the Pharmacy track and the Traditional Chinese Medicine (TCM) Pharmacy track. The data provided here pertains to the Pharmacy track examination. It is important to note that this dataset was collected from online sources, and there may be some discrepancies between this data and the actual examination.
Repository: https://github.com/FreedomIntelligence/HuatuoGPT-II
pharmacy-technician-supervision-ratios
State pharmacy technician-to-pharmacist supervision ratios
Canonical, always-current version: https://referencesource.org/pharmacy-technician-supervision-ratios/
Machine-readable: https://referencesource.org/pharmacy-technician-supervision-ratios/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-16
Stale after: 2027-08-16 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 21
The maximum number of… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/pharmacy-technician-supervision-ratios.pharma-instruction-dataset
Pharma Instruction Dataset
Pharmaceutical domain instruction-tuning dataset in Alpaca format
(instruction, input, output fields).
Used to fine-tune TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T
as Stage 2 (instruction tuning) on top of the domain-adapted Stage 1 model.
Format
Each JSONL record contains:
{
"instruction": "Explain the primary mechanism of action of metformin.",
"input": "",
"output": "Metformin works primarily by ..."
}
The input field… See the full description on the dataset page: https://huggingface.co/datasets/ThakrePranjal/pharma-instruction-dataset.pharma-digital-marketing-dataset
Pharma Digital Marketing Dataset
Summary
High-level, non-diagnostic content patterns for pharmaceutical digital marketing and AI visibility: disease education framing, access and policy narratives, HCP-neutral explainers, and safe handling of regulated claims. Each row includes a compliance_zone label; not legal or medical advice.
Hub target: nebulatech/pharma-digital-marketing-dataset
Terminology
AI SEO — Optimizing owned content and structured data so AI… See the full description on the dataset page: https://huggingface.co/datasets/nebulatech/pharma-digital-marketing-dataset.pharmacy_rx_questions
Pharmacy_Rx_Questions (Synthetic B2B Dataset Preview)
Add me on Discord: xomohappy for access support, delivery questions, or product questions about this premade commercial dataset.
This is a premium, privacy-compliant, industry-safe synthetic dataset simulating Pharmacy Prescription Inquiries & Advisory Logs for B2B applications.
About this Dataset
This dataset is generated programmatically using large language models combined with a strict data curation and… See the full description on the dataset page: https://huggingface.co/datasets/HaseebDev/pharmacy_rx_questions.pharma-preference-dataset-unsloth
Pharma DPO Preference Dataset — Unsloth Pipeline
Preference dataset in DPO format (prompt / chosen / rejected) used for
Stage 3 DPO training in the Unsloth 3-stage pharma fine-tuning pipeline.
Format
{
"prompt": "### Instruction:\nExplain the mechanism of metformin.\n\n### Response:",
"chosen": "Metformin primarily acts by activating AMPK...",
"rejected": "Metformin mainly works by increasing insulin secretion..."
}
Stats
Total rows: 48… See the full description on the dataset page: https://huggingface.co/datasets/ThakrePranjal/pharma-preference-dataset-unsloth.pharmacoeconomic-evidence-extraction-dataset
Pharmacoeconomic Evidence Extraction Dataset
License
This dataset is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. You may use, share, and adapt the dataset provided that appropriate credit is given to the dataset authors.
For the full license terms, see the CC BY 4.0 license.
Overview
This dataset contains 250 expert-annotated records for research on automated extraction of structured pharmacoeconomic and… See the full description on the dataset page: https://huggingface.co/datasets/MJ16/pharmacoeconomic-evidence-extraction-dataset.linking-pharmacybenchmark-vidore-v3-pharmaceuticalsPharmacy-Prescription-Text-Extraction-Dataset
Pharmacy Prescription Text Extraction Dataset
In the field of healthcare, digitizing pharmacy prescriptions is a crucial direction to improve medication management efficiency and reduce human errors. However, current solutions have significant limitations in text extraction accuracy and handwritten text recognition capabilities, affecting the practicality of electronic prescription systems. The construction of this dataset aims to address these issues by providing high-quality… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/Pharmacy-Prescription-Text-Extraction-Dataset.pharma-instruction-dataset-unsloth
Pharma Instruction Dataset — Unsloth Pipeline
Instruction-tuning dataset in Alpaca format used for Stage 2 SFT
in the Unsloth 3-stage pharma fine-tuning pipeline.
Format
{
"instruction": "Explain the primary mechanism of action of metformin.",
"input": "",
"output": "Metformin primarily acts by activating AMP-activated protein kinase (AMPK)..."
}
Training prompt template
### Instruction:
<instruction>
### Input: (omitted if empty)… See the full description on the dataset page: https://huggingface.co/datasets/ThakrePranjal/pharma-instruction-dataset-unsloth.Pharmacy_Identity_Synthetic_QA
Eczacılık Soru-Cevap Veri Seti (Turkish Pharmacy Synthetic QA)
Veri Seti Özeti
Bu veri seti, Türkçe konuşan bir eczacılık/ilaç bilgisi asistanının supervised fine-tuning (SFT)
ile eğitilmesi amacıyla hazırlanmıştır. Toplam 1030 konuşma içerir:
1000 alan bilgisi (domain) örneği — gerçek, hakemli eczacılık/farmasötik bilim makalelerinin
özetlerinden (ÖZ / Amaç / Gereç ve Yöntem / Sonuç ve Tartışma) üretilmiş soru-cevap çiftleri.
30 kimlik (identity/persona) örneği… See the full description on the dataset page: https://huggingface.co/datasets/menesnas/Pharmacy_Identity_Synthetic_QA.trainval
