datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hugging-Ligand-embeddings
HuggingLigand Dataset
Overview
HuggingLigand is a deep learning pipeline developed to predict the binding affinity between proteins and ligands. This prediction task is essential in fields such as drug discovery, biophysics, and computational biology, where determining how strongly a small molecule ligand binds to a protein target is a key step in understanding molecular interactions and prioritizing drug candidates.
The dataset provides precomputed embeddings for… See the full description on the dataset page: https://huggingface.co/datasets/RSE-Group11/Hugging-Ligand-embeddings.family-target-interaction-ligand-gtopdb
Dataset Card for dwb2023/family-target-interaction-ligand-gtopdb
Dataset Details
Dataset Description
Dataset created using the IUPHAR database from the Guide to Pharmacology website
Dataset Sources
Repository: https://www.guidetopharmacology.org/DATA/public_iuphardb_v2024.2.zip
Paper: https://academic.oup.com/nar/article/52/D1/D1438/7332061?login=false
Dataset Creation
Source Data
Query used to create the dataset:
SELECT… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/family-target-interaction-ligand-gtopdb.LigandsSupraDB-LigandScore
SupraDB-LigandScore
What it is
SupraBench/SupraDB-LigandScore is a SupraBench feature dataset generated from the
SupraEngineering compute pipeline. Pipeline position: Phase 1 whole-pool ligand-only scoring by score_nodock; Phase 2 publish split.
The table is designed to join with the Phase 0 SupraDB-GEOM identity table and
the other feature datasets through inchikey.
Schema
Column
Dtype
Units
Meaning
inchikey
string
none
Primary InChIKey… See the full description on the dataset page: https://huggingface.co/datasets/SupraBench/SupraDB-LigandScore.nano-syn-nanoparticle-surface-ligand-decoupling-v0.1Goal
Detect ligand decoupling.
This is the failure mode where:
Ligand metricsstop predictingfunctional stability and targeting.
A batch can look acceptableby ligand density alonewhile zeta potential, size, and bindingdrift into failure.
Inputs
ligand density
zeta potential
hydrodynamic diameter
aggregation rate
in vitro binding efficiency
stability window
Required outputs
ligand_coherence_score
decoupling_flag
decoupling_type
targeting_failure_probability
aggregation_risk… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/nano-syn-nanoparticle-surface-ligand-decoupling-v0.1.family_target_ligand_gtopdb
Dataset Card for dwb2023/family_target_ligand_gtopdb
Dataset Details
Dataset Description
Dataset created using the IUPHAR database from the Guide to Pharmacology website
Dataset Sources
Repository: https://www.guidetopharmacology.org/DATA/public_iuphardb_v2024.2.zip
Paper [optional]: https://academic.oup.com/nar/article/52/D1/D1438/7332061?login=false
Dataset Creation
Source Data
Query used to create the dataset:… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/family_target_ligand_gtopdb.DrugMap_Ligandability
Cysteine Structure Database
The [Cysteine Structure Database] is a dataset compiled of strucutral data for 6515 cysteine sites in hundreds of proteins.
This dataset was published in Cell
and is also available at the official DrugMap Github repo.
For each cysteine site, this database includes numerical values for Solvent Accessible Surface Area (SASA), Cysteine Depth, etc.
Additionally, each cysteine site has a probe engagement score derived from isotopic tandem orthogonal… See the full description on the dataset page: https://huggingface.co/datasets/ymanasa2000/DrugMap_Ligandability.
