datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
drug-reviewsUCI_drug_reviews
Data Description
This data comes from the UC Irvine Machine Learning Repository. It has been preprocessed to only contain reviews at least 13 or more words in length. The raw data for this
specific dataset can be found here. The base UCI ML url can be found
here.
Drug-Review-DatasetFR_Drugs_tiret_dictation_augmented
FR_Drugs_tiret_dictation_augmented
Dictee de listes de medicaments au format "tiret" (patterns du dataset
FR_Drugs_dictation_with_dashes_pattern, molecules completees depuis les sources
autocorrect/medical), synthetisee en TTS Coqui XTTS v2 (600 voix clonees) puis
augmentee acoustiquement.
element
valeur
extraits
141960
moteur
Coqui XTTS v2, 600 voix clonees (16 kHz mono)
base
audio/coqui/<shard>/
parasite seul (65%)
audio_babble_only/
echo seul (25%)… See the full description on the dataset page: https://huggingface.co/datasets/PraxySante/FR_Drugs_tiret_dictation_augmented.DrugCLIP_data
🧬 DrugCLIP data repository
This repository hosts benchmark datasets, pre-computed molecular embeddings, pretrained model weights, and supporting files used in the DrugCLIP project. It also includes data and models used for wet lab validation experiments.
📁 Repository Contents
1. DUD-E.zip
Full dataset for the DUD-E benchmark.
Includes ligand and target files for all targets.
2. LIT-PCBA.zip
Full dataset for the LIT-PCBA benchmark.
Includes… See the full description on the dataset page: https://huggingface.co/datasets/THU-ATOM/DrugCLIP_data.uci-drug-review-cleaneddrugReviewsagentic-drug-discovery-system
Agentic Drug Discovery System
This card describes the public 0.3.0.dev3 Agentic Drug Discovery System mirror.
Scope. The proposed eight-stage, long-horizon agentic drug discovery system remains a research scaffold rather than a completed public platform. Seven of eight planned atlases have no standalone public data, and the demonstrated continuous multi-stage program currently covers one disease/target slice traversed retrospectively.
It contains the executable control plane… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/agentic-drug-discovery-system.UCI_drugDrugBankdrug-seq-u2os-novartisI AM NOT AFFILIATED WITH NOVARTIS IN ANY WAY; THIS IS SIMPLY AN UPLOAD OF THEIR DATASET, "NOVARTIS/DRUG-SEQ U2OS MOABOX DATASET."
Novartis DRUG-seq U2OS MoABox Dataset
This dataset profiles transcriptomic responses of the U-2 OS human osteosarcoma cell line to a broad collection of small molecule perturbations. It contains 49,392 observations spanning 3,742 unique compounds tested at 4 distinct dosages + 0.0, each annotated with their respective mechanisms of action (MoA).
Each… See the full description on the dataset page: https://huggingface.co/datasets/TitouanCh/drug-seq-u2os-novartis.encoded_drug_reviewstw-drug-labels-vision
Dataset Card for tw-drug-labels-vision
💊 tw-drug-labels-vision 是一份涵蓋臺灣食品藥物管理署(TFDA)核發之 44,663 筆藥品仿單/外盒 的繁體中文多模態資料集。每一筆紀錄同時包含 PDF 全部頁面的渲染圖(WebP 多頁)以及一份依統一 17 欄 JSON Schema 抽取自原始藥品標示文件的結構化資料,可直接用於語言模型微調、視覺語言模型訓練、文件問答、藥品知識檢索、繁體中文醫藥 NLP 任務之素材。
Dataset Details
Dataset Description
本資料集源自臺灣 TFDA 公開的藥品許可證查詢系統。每筆紀錄對應一份藥品文件(仿單或外盒),原始為 PDF 圖檔形式。處理流程分為三階段:
下載:依據 20251222政府開放資料集_仿單與藥品外盒_66032.xlsx 中的 PDF URL,下載原始檔。
頁面渲染:將 PDF 各頁渲染為 WebP 圖檔,封裝在 images 欄位中。
OCR +… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-drug-labels-vision.drug-list
Dataset Card for Every Cure Drug List
Dataset Summary
The Every Cure Drug List is a manually curated list of drug entities used by the MATRIX project for drug repurposing predictions.
For reference, the list contains ~1,800 drugs with metadata including approval status, drug class flags, therapeutic annotations, and ATC classifications.
For more information see here.
Source Data
Attribution
faers_bronze_drugdrugtargetbench-assets
DrugTargetBench assets
Immutable per-world payloads for the
DrugTargetBench Harbor
benchmark. You do not download this repository by hand: Harbor resolves the
archives a task needs and verifies them against the frozen manifest before
the agent starts.
uv tool install harbor
harbor run -d drugtargetbench/drugtargetbench@v1.0 -a <agent> -m <model>
Layout
worlds/<world-id>/tables.tar omics, tabular, data dictionary
worlds/<world-id>/signals.tar… See the full description on the dataset page: https://huggingface.co/datasets/sammargolis/drugtargetbench-assets.drug-reviews
Dataset Details
1.Dataset Loading:
Initially, we load the Drug Review Dataset from the UC Irvine Machine Learning Repository. This dataset contains patient reviews of different drugs, along with the medical condition being treated and the patients' satisfaction ratings.
2.Data Preprocessing:
The dataset is preprocessed to ensure data integrity and consistency. We handle missing values and ensure that each patient ID is unique across the dataset.
3.Text… See the full description on the dataset page: https://huggingface.co/datasets/Mouwiya/drug-reviews.drug_label_approved_openfda
KEMIRIX OpenFDA Clinical Drug Dataset
Built for KEMIRIX — Africa's first Clinical Decision Support AI
Developer: Emmanuel Bain Oduwo | TechFryz Ltd. | Nairobi, Kenya
Generated: May 2026
Configurations
clean (default): instruction + output only, fully cleaned, ready for fine-tuning Kemirix
raw: full metadata schema, original generated data
Usage
from datasets import load_dataset
# Clean data for training Kemirix
ds =… See the full description on the dataset page: https://huggingface.co/datasets/Oduwo/drug_label_approved_openfda.drugprotThe DrugProt corpus consists of a) expert-labelled chemical and gene mentions, and (b) all binary relationships between them corresponding to a specific set of biologically relevant relation types.drug_reviewsStack-DrugPBMC-Example
License
CC-BY-4.0
Source
This dataset is derived from the Open Problems – Single-Cell Perturbations Kaggle competition, sponsored by Cellarity, Inc. and organized by Open Problems in Single-Cell Analysis.
The original competition data is licensed under CC-BY-4.0.
Citation
If you use this dataset, please cite:
Luecken et al., 2025: Defining and benchmarking open problems in single-cell analysis. Nature Biotechnology.
drugchat_liang_zhang_et_al
Dataset Details
Dataset Description
Instruction tuning dataset used for the LLM component of DrugChat.
10,834 compounds (3,8962 from ChEMBL and 6,942 from PubChem) containing
descriptive drug information were collected. 143,517 questions were generated
using the molecules' classification, properties and descriptions from ChEBI, LOTUS & YMDB.
Curated by:
License: BSD-3-Clause
Dataset Sources
corresponding publication
rep & data source
Citation… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/drugchat_liang_zhang_et_al.drug-target-activity
Introduction
This dataset containing measurements of drug-target interactions is provided by EvE Bio, highlighted in an exciting Hugging Face blog post. It is actively being generated with a quantitative screening process, and data for new targets is added every other month. For each target, one or more types of activity (agonism, antagonism, etc.) are measured
for every drug in a 1,397 member compound library that primarily represents FDA approved small molecule drugs. Results… See the full description on the dataset page: https://huggingface.co/datasets/eve-bio/drug-target-activity.DrugbankRawParquetgeom_drugs
GEOM: Molecular Conformations (Drugs Subset)
Note: This is a mirrored and specifically preprocessed version of the GEOM dataset (Drugs subset), originally created by Simon Axelrod and Rafael Gómez-Bombarelli. All credit for the original conformational sampling and DFT calculations goes to the original authors. This repository exists to guarantee availability and exact reproducibility for downstream machine learning projects.
Dataset Description
The Geometric Ensemble… See the full description on the dataset page: https://huggingface.co/datasets/raulsofia/geom_drugs.DrugDiscoveryBench-Preview
DrugDiscoveryBench (Preview)
DrugDiscoveryBench is a benchmark of 82 expert-authored, execution-grounded tasks spanning the early
drug-discovery and life-sciences workflow (target identification & genetics, database screening,
patent mining, cheminformatics, structural reasoning, SAR & affinity, molecular biology). Each task
asks an agent to carry out a multi-step biomedical investigation and produce a terse final answer.
This is the Preview release: task prompts and metadata… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/DrugDiscoveryBench-Preview.drug-reviewsDrugbankVocabularyTDC-DrugADMETTDC_pampa_approved_drugs
