pharma
OpenMed-NER-PharmaDetect-SuperClinical-434MOpenMed-NER-PharmaDetect-BigMed-278MOpenMed-NER-PharmaDetect-ModernClinical-149MOpenMed-NER-PharmaDetect-BioPatient-108MOpenMed-NER-PharmaDetect-BigMed-560MOpenMed-NER-PharmaDetect-SuperMedical-125MOpenMed-NER-PharmaDetect-SuperMedical-355MOpenMed-NER-PharmaDetect-ElectraMed-560M
Datasets
All datasets matching “pharma”vidore_v3_pharmaceuticalsViDoRe V3 : Pharmaceuticals
This dataset, Pharmaceutical, is a corpus of slides from the FDA, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with human-verified relevant pages… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_pharmaceuticals.vidore_v3_pharmaceuticals_mteb_format
Vidore3PharmaceuticalsRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions.
Task category
t2i
Domains
Academic
Reference
https://huggingface.co/blog/QuentinJG/introducing-vidore-v3
Source datasets:
vidore/vidore_v3_pharmaceuticals
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_pharmaceuticals_mteb_format.pharmaconerPharmaCoNER: Pharmacological Substances, Compounds and Proteins Named Entity Recognition track
This dataset is designed for the PharmaCoNER task, sponsored by Plan de Impulso de las Tecnologías del Lenguaje (Plan TL).
It is a manually classified collection of clinical case studies derived from the Spanish Clinical Case Corpus (SPACCC), an
open access electronic library that gathers Spanish medical publications from SciELO (Scientific Electronic Library Online).
The annotation of the entire set of entity mentions was carried out by medicinal chemistry experts
and it includes the following 4 entity types: NORMALIZABLES, NO_NORMALIZABLES, PROTEINAS and UNCLEAR.
The PharmaCoNER corpus contains a total of 396,988 words and 1,000 clinical cases that have been randomly sampled into 3 subsets.
The training set contains 500 clinical cases, while the development and test sets contain 250 clinical cases each.
In terms of training examples, this translates to a total of 8074, 3764 and 3931 annotated sentences in each set.
The original dataset was distributed in Brat format (https://brat.nlplab.org/standoff.html).
For further information, please visit https://temu.bsc.es/pharmaconer/ or send an email to encargo-pln-life@bsc.espharma-bench-mlmPharmaShip
PharmaShip: An Entity-Centric, Reading-Order-Supervised Benchmark for Chinese Pharmaceutical Shipping Documents
🔗 Paper: https://arxiv.org/abs/2512.23714
🔗 Github: https://github.com/KevinYuLei/PharmaShip
Description
PharmaShip is a real-world Chinese dataset of scanned pharmaceutical shipping documents designed to stress-test pre-trained text-layout models under noisy OCR and heterogeneous templates.
It covers three complementary tasks:
Sequence Entity Recognition… See the full description on the dataset page: https://huggingface.co/datasets/YuLeiKevin/PharmaShip.pharma-kb-obesity
Pharma KB — Obesity Drug Landscape
A structured pharmaceutical knowledge base covering the obesity drug pipeline — compiled from public regulatory, clinical, and scientific sources. Designed for RAG pipelines, competitive intelligence workflows, and pharma-domain LLM fine-tuning.
Overview
Indication
Obesity (ICD: E66)
Drug articles
285 (all phases, launched through discovery)
Company articles
248
Target articles
35
Full content size
~11 MB… See the full description on the dataset page: https://huggingface.co/datasets/OpenPharma/pharma-kb-obesity.
