datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
drug-seq-u2os-novartisI AM NOT AFFILIATED WITH NOVARTIS IN ANY WAY; THIS IS SIMPLY AN UPLOAD OF THEIR DATASET, "NOVARTIS/DRUG-SEQ U2OS MOABOX DATASET."
Novartis DRUG-seq U2OS MoABox Dataset
This dataset profiles transcriptomic responses of the U-2 OS human osteosarcoma cell line to a broad collection of small molecule perturbations. It contains 49,392 observations spanning 3,742 unique compounds tested at 4 distinct dosages + 0.0, each annotated with their respective mechanisms of action (MoA).
Each… See the full description on the dataset page: https://huggingface.co/datasets/TitouanCh/drug-seq-u2os-novartis.geom_drugs
GEOM: Molecular Conformations (Drugs Subset)
Note: This is a mirrored and specifically preprocessed version of the GEOM dataset (Drugs subset), originally created by Simon Axelrod and Rafael Gómez-Bombarelli. All credit for the original conformational sampling and DFT calculations goes to the original authors. This repository exists to guarantee availability and exact reproducibility for downstream machine learning projects.
Dataset Description
The Geometric Ensemble… See the full description on the dataset page: https://huggingface.co/datasets/raulsofia/geom_drugs.Misssing_FR_TTS_ICD_Snomed_drugs_0806
Missing FR TTS ICD Snomed drugs 0506
Dataset TTS pour des termes medicaux manquants. Genere le 2026-06-06.
Contenu
53192 fichiers audio
25295 textes uniques
Edge TTS: 25277 fichiers (voix Denise, Microsoft)
Coqui XTTS: 0 fichiers (2 voix/texte, ~10% des textes)
Source: synthese sur 3x V100-32GB
Format
Les fichiers audio sont dans archives/*.tar.gz.
Le fichier metadata.jsonl est a la racine du dataset.
Stats
Edge TTS: 25277
Coqui… See the full description on the dataset page: https://huggingface.co/datasets/PraxySante/Misssing_FR_TTS_ICD_Snomed_drugs_0806.drugsComTest_rawWith a dataset of over 2000 drugs for varying health situations, over 200,000 observations, 7 attributes, and tens of thousands of texts by users of their experience; categorizing these texts will be an extremely difficult task without an efficient algorithm for resolving the problem.
https://www.kaggle.com/
safe-drugs
SAFE
Sequential Attachment-based Fragment Embedding (SAFE) is a novel molecular line notation that represents molecules as an unordered sequence of fragment blocks to improve molecule design using generative models.
This is the drugs dataset used for benchmarking.
Find the details and how to use at SAFE in the repo https://github.com/datamol-io/safe or the paper https://arxiv.org/pdf/2310.10773.pdf.
drug-screening-qa
Workplace Drug Screening Q&A
A question-answering dataset for workplace drug screening — federal drug testing
regulations, specimen collection, laboratory methodology, testing panels, and Medical
Review Officer (MRO) procedures. Built to fine-tune a small instruction model into a
domain assistant for HR professionals, employers, occupational health staff, and MROs.
Contents
File
Rows
Purpose
train.jsonl
639
Training split
val.jsonl
71
Validation split… See the full description on the dataset page: https://huggingface.co/datasets/oikyoni/drug-screening-qa.drug-safety-faersFDA-Approved-Drugs
FDA-Approved-Drugs
Compendium of FDA-approved drugs (including withdrawn) compiled from ChEMBL, DrugCentral, and Thera-SAbDab.
Splits
small_molecule: one row per molecule. smiles and selfies populated; sequence empty.
single_chain: one row per single-chain protein drug (peptides, hormones, scFv, single-chain Fc-fusions).
multi_chain: one row per chain of a multi-chain biologic. Antibodies are decomposed into variable regions only (VH/VL) when extractable via… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/FDA-Approved-Drugs.drug_slangs
Drug Slang Terms with Nationality Mapping
Description
A curated multilingual dataset of drug‑related slang terms, mapped to their language, nationality, region, category, and confidence score.Terms were collected from public sources, social media, and linguistic research.
Files
drug_slang_terms.csv – CSV format (easy to open in Excel/Pandas)
drug_slang_terms.json – JSON format (structured, ready for web apps)
Format
Field
Type… See the full description on the dataset page: https://huggingface.co/datasets/adi515/drug_slangs.drugs-75kDrugs-75K Dataset
A dataset of 75,099 drug-like molecules with 558,002 conformers total.
Statistics
Num. Molecules: 75,099
Num. Conformers: 558,002
Avg. Heavy Atoms: 30.56
Avg. Rotatable Bonds: 7.53
Num. Targets: 3 (ip, ea, chi)
Atomic Species: H, C, N, O, F, Si, P, S, Cl
Format
Each example contains:
name: Unique molecule identifier
smiles: SMILES string representation
num_conformers: Number of conformers for this molecule
conformers: List of mol block strings (one… See the full description on the dataset page: https://huggingface.co/datasets/eamag/drugs-75k.drugs_reviews_datasetdrugs-composition-indonesian-donut
Dataset Card for "drugs-composition-indonesian-donut"
Generate Custom Data
Please visit https://huggingface.co/spaces/jonathanjordan21/donut-labelling for the interface to generate custom data.
The data format is (.zip). Images and Labels are stored in separated .zip files.
[More Information needed](https://github.com/huggingface/datasets/blob/main/CONTRIBUTING.md#how-to-contribute-to-the-dataset-cards
fda-approved-drugsDrugs-NER-Data
Named Entity Recognition (NER) Dataset
Overview
This dataset is designed for Named Entity Recognition (NER) tasks. It contains annotated text data with entities classified into various categories. The dataset is split into training, validation, and test sets, and is suitable for training and evaluating NER models.
Dataset Description
Features
The dataset includes the following features:
tokens: A sequence of strings representing the words or tokens… See the full description on the dataset page: https://huggingface.co/datasets/pnr-svc/Drugs-NER-Data.llama_drugs_datasetdrugs-and-substances-wiktionary-pronunciations
TigreGotico/drugs-and-substances-wiktionary-pronunciations
Rich entity dataset scraped by metadatarr
scraper wiktionary_pronunciations.
Part of the drugs-and-substances collection.
Rows: 1,200
Fields
term
language
ipa
wikitext_excerpt
wiktionary_url
source_wiktionary
Source
Generated by scrapers/wiktionary_pronunciations.py. See the metadatarr repo for the full
pipeline and scraper source code.
drugsfullafrica-who-people-who-inject-drugs-population-size-estimates
Africa — WHO GHO: People who inject drugs: Population size estimates (number) | Africa (World Health Organization)
Size category: n<1K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-who-people-who-inject-drugs-population-size-estimates.phi_drugs_datasetsmall_molecule_drugsdrugsOncology-Cancer-Approved-Drugsasia-owid-deaths-alcohol-drugs-by-sex-who
Deaths Alcohol Drugs By Sex Who | Asia (Our World in Data)
🌏 1,034 observations · 47 Asia countries · 2000–2021 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 1,034 observations of Deaths Alcohol Drugs By Sex Who data across 47 Asia countries, spanning 2000–2021.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Deaths Alcohol Drugs By Sex Who
Geographic coverage… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-deaths-alcohol-drugs-by-sex-who.asia-owid-cost-drugs-extensive-resistant-tb
Cost Drugs Extensive Resistant Tb | Asia (Our World in Data)
🌏 179 observations · 31 Asia countries · 2017–2024 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 179 observations of Cost Drugs Extensive Resistant Tb data across 31 Asia countries, spanning 2017–2024.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Cost Drugs Extensive Resistant Tb
Geographic coverage… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-cost-drugs-extensive-resistant-tb.common_drugsdrugstotal_synthesis_of_new_drugs
Total Synthesis of New Drugs
This dataset contains basic information and raw synthesis route data for 60 drugs, along with 600 multimodal reasoning supervision (CoT) fine-tuning data.
The experimental and computational work in this dataset run on the Huawei Cloud AI Compute Service. We appreciate the stable compute supply from this platform.
新药化学全合成路线
该数据集包含 60 种药物的基本信息与合成路线原始数据,以及 600 条多模态思维监督微调数据。
本数据集的实验与计算工作依托于华为昇腾AI云服务平台完成,特此对其提供的稳定算力支持表示感谢。
drug-synonym-sid-mappingner-drugs
Dataset Card for ner-drugs
This dataset was originally part of this tutorial. The goal of the
dataset is to find references to drugs in Reddit discussions.
exclusive-pediatric-drugs
Exclusive Pediatric Drugs
Description
Exclusively Pediatric (EP) designation allows for a minimum rebate percentage of 17.1 percent of Average Manufacturer Price (AMP) for single source and innovator multiple source drugs.
Dataset Details
Publisher: Centers for Medicare & Medicaid Services
Last Modified: 2025-05-05
Contact: Medicaid.gov (Medicaid.gov@cms.hhs.gov)
Source
Original data can be found at: https://healthdata.gov/d/5tbn-ibkm
Usage… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/exclusive-pediatric-drugs.
