binding
Datasets
All datasets matching “binding”pane-binding-functions-attributionbinding_affinityA dataset to fine-tune language models on protein-ligand binding affinity prediction.Antibody_Binding_Benchmark_DatasetWe introduce AbBiBench (Antibody Binding Benchmarking), a benchmarking framework for optimizing antibody binding affinity. This dataset contains the sequences of mutants, experimentally measured affinity values, and the structures of antigen-antibody complexes.
Mutant structure files for AbBiBench dataset is available at https://zenodo.org/records/16557372
biomap-research-metal_ion_binding
metal_ion_binding
Sourced from biomap-research/metal_ion_binding and prepared for Hugging Face datasets usage.
Data files
Parquet files are stored under data/ using Hugging Face split naming conventions
(train-*, validation-*, test-*).
Preparation
Preprocess mode: minimal.
Seed: 1957723.
No max sequence length filter was applied.
Renamed source columns: label -> targets, seq -> sequence.
Columns: id, sequence, targets, split.
Validation handling:… See the full description on the dataset page: https://huggingface.co/datasets/swhitfield/biomap-research-metal_ion_binding.metal_ion_binding
Dataset Card for Metal Ion Binding Dataset
Dataset Summary
Metal ion binding sites within proteins play a crucial role across a spectrum of processes, spanning from physiological to pathological, toxicological, pharmaceutical, and diagnostic. Consequently, the development of precise and efficient methods to identify and characterize these metal ion binding sites in proteins has become an imperative and intricate task for bioinformatics and structural biology.… See the full description on the dataset page: https://huggingface.co/datasets/biomap-research/metal_ion_binding.binding_sites_random_split_by_family_550KThis dataset is obtained from a UniProt search
for protein sequences with family and binding site annotations. The dataset includes unreviewed (TrEMBL) protein sequences as well as
reviewed sequences. We refined the dataset by only including sequences with an annotation score of 4. We sorted and split by family, where
random families were selected for the test dataset until approximately 20% of the protein sequences were separated out for test data.
We excluded any sequences with <, >, or ?… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/binding_sites_random_split_by_family_550K.
