datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pharma-kb-obesity
Pharma KB — Obesity Drug Landscape
A structured pharmaceutical knowledge base covering the obesity drug pipeline — compiled from public regulatory, clinical, and scientific sources. Designed for RAG pipelines, competitive intelligence workflows, and pharma-domain LLM fine-tuning.
Overview
Indication
Obesity (ICD: E66)
Drug articles
285 (all phases, launched through discovery)
Company articles
248
Target articles
35
Full content size
~11 MB… See the full description on the dataset page: https://huggingface.co/datasets/OpenPharma/pharma-kb-obesity.indian-pharma-dataset-2026-augast
PharmaLens: 200k Medicine Catalog (Salts, Prices, Interactions and Reviews)
I spent weeks compiling and cleaning this retrieval database for a project. Instead of letting 300MB+ of structured pharmaceutical data sit idle on my hard drive, I am open-sourcing it. Use it for your RAG pipelines, chatbots, pricing tools, or whatever else you are building.
Overview
Finding clean, structured pharmaceutical datasets with commercial brand names, active salt compositions… See the full description on the dataset page: https://huggingface.co/datasets/sinhal/indian-pharma-dataset-2026-augast.Hiro-Pharma-RAG-Benchmark
Hiro Pharma RAG Benchmark
This private dataset repository contains multilingual biomedical RAG benchmark data associated with the paper CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine.
The benchmark is designed to evaluate whether retrieval-augmented language models can answer biomedical questions while selecting and citing useful evidence and filtering out noisy or irrelevant references.
Repository Contents
File
Language… See the full description on the dataset page: https://huggingface.co/datasets/PatSnap/Hiro-Pharma-RAG-Benchmark.K-Paths-inductive-reasoning-pharmaDB
🔗 This dataset is part of the study:
K-Paths: Reasoning over Graph Paths for Drug Repurposing and Drug Interaction Prediction
📖 Read the Paper
💾 GitHub Repository
PharmacotherapyDB: Inductive Reasoning Dataset
PharmacotherapyDB is a drug repurposing dataset containing drug–disease treatment relations in three categories (disease-modifying, palliates, or non-indication).
Each entry includes a drug and a disease, an interaction label, drug, disease descriptions, and… See the full description on the dataset page: https://huggingface.co/datasets/Tassy24/K-Paths-inductive-reasoning-pharmaDB.pharmacology-1k
Pharmacology Medical Dataset — 1,000 Record Free Sample
Enterprise-grade synthetic medical data. Zero PHI. HIPAA-Aligned.
Quality Metrics
Metric
Score
Industry Benchmark
Trinity Consensus Score (TAS)
98.0%
85-92% typical
Trinity Assurance Score (TAS)
0.97
0.75-0.85 typical
Macro F1
0.97
0.80-0.90 typical
PHI Present
None
--
Generation Method
3-LLM Trinity Ensemble
Single model typical
What's Included (Free)
1,000… See the full description on the dataset page: https://huggingface.co/datasets/WitnessDataFactory/pharmacology-1k.Pharmacy_Identity_Synthetic_QA
Eczacılık Soru-Cevap Veri Seti (Turkish Pharmacy Synthetic QA)
Veri Seti Özeti
Bu veri seti, Türkçe konuşan bir eczacılık/ilaç bilgisi asistanının supervised fine-tuning (SFT)
ile eğitilmesi amacıyla hazırlanmıştır. Toplam 1030 konuşma içerir:
1000 alan bilgisi (domain) örneği — gerçek, hakemli eczacılık/farmasötik bilim makalelerinin
özetlerinden (ÖZ / Amaç / Gereç ve Yöntem / Sonuç ve Tartışma) üretilmiş soru-cevap çiftleri.
30 kimlik (identity/persona) örneği… See the full description on the dataset page: https://huggingface.co/datasets/menesnas/Pharmacy_Identity_Synthetic_QA.
