CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01THU-ATOM /PDBbind_test0 likes16k downloads9mo agoHugging Face02THU-ATOM /PDBbind0 likes6k downloads11mo agoHugging Face03LiteFold /PDB PDB mmCIF Entry Index The Protein Data Bank is the single global archive of experimentally-determined 3D structures of biological macromolecules, established in 1971 and now holding well over 230,000 entries. It stores atomic coordinates for proteins, nucleic acids, and their complexes determined by X-ray crystallography, cryo-EM, NMR, micro-electron diffraction, and integrative methods, along with the underlying experimental data (structure factors, EM maps, NMR restraints) and… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/PDB.tabular10K<n<100K2 likes3.9k downloads4mo agoHugging Face04rouskinlab /PDB Data types sequence: 355 datapoints structure: 355 datapoints text10K<n<100K0 likes1.8k downloads2y agoHugging Face05houlab /pdb-dbtabular1K<n<10K0 likes626 downloads25d agoHugging Face06jglaser /pdb_protein_ligand_complexes How to use the data sets This dataset contains about 36,000 unique pairs of protein sequences and ligand SMILES, and the coordinates of their complexes from the PDB. SMILES are assumed to be tokenized by the regex from P. Schwaller. Ligand selection criteria Only ligands that have at least 3 atoms, a molecular weight >= 100 Da, and which are not among the 280 most common ligands in the PDB (this includes common additives like PEG, ADP, ..) are considered. Use… See the full description on the dataset page: https://huggingface.co/datasets/jglaser/pdb_protein_ligand_complexes.7 likes381 downloads4y agoHugging Face07biodatasets /pdbtext100K<n<1M0 likes321 downloads2y agoHugging Face08PDBEurope /protein_structure_NER_model_v3.1 Overview This data was used to train model: https://huggingface.co/PDBEurope/BiomedNLP-PubMedBERT-ProteinStructure-NER-v3.1 There are 20 different entity types in this dataset: "bond_interaction", "chemical", "complex_assembly", "evidence", "experimental_method", "gene", "mutant", "oligomeric_state", "protein", "protein_state", "protein_type", "ptm", "residue_name", "residue_name_number","residue_number", "residue_range", "site", "species", "structure_element", "taxonomy_domain"… See the full description on the dataset page: https://huggingface.co/datasets/PDBEurope/protein_structure_NER_model_v3.1.0 likes290 downloads2y agoHugging Face09PDBEurope /protein_structure_NER_model_v2.1 Overview This data was used to train model: https://huggingface.co/PDBEurope/BiomedNLP-PubMedBERT-ProteinStructure-NER-v2.1 There are 20 different entity types in this dataset: "bond_interaction", "chemical", "complex_assembly", "evidence", "experimental_method", "gene", "mutant", "oligomeric_state", "protein", "protein_state", "protein_type", "ptm", "residue_name", "residue_name_number","residue_number", "residue_range", "site", "species", "structure_element", "taxonomy_domain"… See the full description on the dataset page: https://huggingface.co/datasets/PDBEurope/protein_structure_NER_model_v2.1.0 likes283 downloads2y agoHugging Face10alegendaryfish /CDDB-PDB-Protein-50-512 CDDB–PDB Protein Structures, 50–512 Residues Curated sequences, observed atomic coordinates, physical side-chain torsions, observation masks, and separately filtered intrinsic-backbone labels for protein generation and conditional modeling. Training PDB cutoff: 31 December 2023, inclusive, using the entry's initial public release date. The source snapshot was collected on 11 September 2026. These dates serve different purposes: historical entries use their audited coordinates… See the full description on the dataset page: https://huggingface.co/datasets/alegendaryfish/CDDB-PDB-Protein-50-512.tabular1M<n<10M0 likes262 downloads10d agoHugging Face11photonmz /pdbbindpp-2020 PDBbind 2020 Refined Set Curated collection of 5,316 high-quality protein-ligand complexes with experimentally measured binding affinity data from PDBbind v.2020 refined set. Structure 5,316 directories named by PDB ID Each contains: protein.pdb, pocket.pdb, ligand.mol2, ligand.sdf Total size: ~2.7 GB Citation Liu, Z.H. et al. Acc. Chem. Res. 2017, 50, 302-309. other1K<n<10K2 likes203 downloads1y agoHugging Face12Synthyra /PDB-Monomeric-Structure-ESMFold2 PDB-Monomeric-Structure-ESMFold2 Monomeric, protein-only PDB structure dataset for minimum ESMFold2-style training. Each row is one eligible single-chain biological assembly with a canonical amino-acid sequence input and all-atom protein labels in atom37. Labels atom37_positions: residue x 37 x 3 coordinates, with zeros for missing atoms. atom37_mask: residue x 37 resolved-atom mask. aatype, residue_index, auth_seq_id, insertion_code, residue_name, ca_mask.… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/PDB-Monomeric-Structure-ESMFold2.tabular100K<n<1M0 likes157 downloads3mo agoHugging Face13multimolecule /pdb-rna_secondary_structure pdb-rna_secondary_structure [!IMPORTANT]The pdb-rna_secondary_structure dataset is in beta test. This dataset card may not accurately reflects the data content. The data content and this dataset card may subject to change. Please contact the MultiMolecule team on GitHub issues should you have any feedback. [!CAUTION] This dataset is converted from the dataset released by the authors of SPOT-RNA. The MultiMolecule is aware of a potential issue in data quality. We are working on… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/pdb-rna_secondary_structure.text-generation100K<n<1M0 likes156 downloads1y agoHugging Face14PDBEurope /protein_structure_NER_model_v1.4 Overview This data was used to train model: https://huggingface.co/PDBEurope/BiomedNLP-PubMedBERT-ProteinStructure-NER-v1.4 There are 19 different entity types in this dataset: "chemical", "complex_assembly", "evidence", "experimental_method", "gene", "mutant", "oligomeric_state", "protein", "protein_state", "protein_type", "ptm", "residue_name", "residue_name_number","residue_number", "residue_range", "site", "species", "structure_element", "taxonomy_domain" The data prepared as… See the full description on the dataset page: https://huggingface.co/datasets/PDBEurope/protein_structure_NER_model_v1.4.0 likes151 downloads2y agoHugging Face15jglaser /pdbbind_complexesA dataset to fine-tune language models on protein-ligand binding affinity and contact prediction.text10K<n<100K1 likes141 downloads4y agoHugging Face16hjchang /PDB_primary_citation primary_citation_with_pmcid.jsonl This dataset links PDB protein structures with their corresponding primary ciatation text content. Format Each line is a JSON object: { "protein_name": "9HCG", "structure_title": "Mouse mitoribosome large subunit assembly intermediate bound to NSUN4, MTERF4, and mt-RNAs", "main_text": "..." } text10K<n<100K0 likes123 downloads11mo agoHugging Face17PDBEurope /protein_structure_NER_independent_val_set Overview This data was used to evaluate the two models below to decide whether convergence was reached. https://huggingface.co/PDBEurope/BiomedNLP-PubMedBERT-ProteinStructure-NER-v2.1 https://huggingface.co/PDBEurope/BiomedNLP-PubMedBERT-ProteinStructure-NER-v3.1 There are 20 different entity types in this dataset: "bond_interaction", "chemical", "complex_assembly", "evidence", "experimental_method", "gene", "mutant", "oligomeric_state", "protein", "protein_state", "protein_type"… See the full description on the dataset page: https://huggingface.co/datasets/PDBEurope/protein_structure_NER_independent_val_set.textn<1K0 likes118 downloads2y agoHugging Face18baber /pdbookstext1M<n<10M0 likes116 downloads3y agoHugging Face19DaInternet12 /pdbbind_affinitiestabular10K<n<100K0 likes114 downloads2y agoHugging Face20ronig /pdb_sequences PDB Sequences This dataset contains 780,163 protein sequences from the RCCB Protein Data Bank text100K<n<1M0 likes111 downloads3y agoHugging Face21MaxwellBauer /Archive_CollaGNN_PDB_txttext10K<n<100K0 likes108 downloads1y agoHugging Face22airkingbd /pdb_swissprottabular100K<n<1M3 likes104 downloads1y agoHugging Face23Synthyra /PDB-Chain-Complex-Benchmark-Rigor PDB-Chain-Complex-Benchmark-Rigor Rigor rebuild of Synthyra/PDB-Chain-Complex-Benchmark with split assignments recomputed from the published chain and complex parquet artifacts. Split Policy Splits are assigned by connected components over exact sequence, sequence hash, 30% sequence cluster, structure cluster, source split component, same-PDB asymmetric-unit membership, chain assembly membership, and biological assembly co-membership from the complex rows. The… See the full description on the dataset page: https://huggingface.co/datasets/Synthyra/PDB-Chain-Complex-Benchmark-Rigor.tabular1M<n<10M0 likes102 downloads3mo agoHugging Face24djh992 /pdbbind_complex_GB2022 To generate the dataset Register for an account at https://www.pdbbind.org.cn/, confirm the validation email, then login and download the Index files (1) the general protein-ligand complexes (2) the refined protein-ligand complexes (3) Extract those files in pdbbind_complex_GB2022/data Run the script pdbbind.py in a compute job on an MPI-enabled cluster (e.g., mpirun -n 64 pdbbind.py). Output will be tar files in train/, val/ and test/ folders, following the split direction which is… See the full description on the dataset page: https://huggingface.co/datasets/djh992/pdbbind_complex_GB2022.text10K<n<100K1 likes98 downloads4y agoHugging Face25TerminatorJ /PDB_Ribonanzanettextn<1K0 likes98 downloads2y agoHugging Face26Annie37 /pdb0 likes92 downloads9d agoHugging Face27nithinc1 /pdb-datasets Datasets for Cryo-EM synthetic training Existing datasets: scope: collection of PDBs categorized from scope tim-rossman: collection of tim and rossman PDBs Dataset Structure: {dataset}/ - pdbs: all pdbs in the dataset - dataframes: different types of datasets (how many families, which pdbs selected, etc.). All paths are relative textn<1K0 likes82 downloads9mo agoHugging Face28PDBEurope /protein_structure_NER_model_v1.2 Overview This data was used to train model: https://huggingface.co/PDBEurope/BiomedNLP-PubMedBERT-ProteinStructure-NER-v1.2 There are 19 different entity types in this dataset: "chemical", "complex_assembly", "evidence", "experimental_method", "gene", "mutant", "oligomeric_state", "protein", "protein_state", "protein_type", "ptm", "residue_name", "residue_name_number","residue_number", "residue_range", "site", "species", "structure_element", "taxonomy_domain" The data prepared as… See the full description on the dataset page: https://huggingface.co/datasets/PDBEurope/protein_structure_NER_model_v1.2.0 likes71 downloads2y agoHugging Face29Precise-Debugging-Benchmarking /PDB-Single-Hard PDB-Single-Hard: Precise Debugging Benchmarking — hard single-line bug subset 📄 Paper  ·  💻 Code  ·  🌐 Project page  ·  🏆 Leaderboard PDB-Single-Hard is the hard single-line bug subset of the PDB (Precise Debugging Benchmarking) evaluation suite. Every example pairs a ground-truth program with a synthesized buggy version plus a line-level edit script (gt_diff) that encodes the minimal correct fix. Source datasets: BigCodeBench + LiveCodeBench Sibling datasets: PDB-Single ·… See the full description on the dataset page: https://huggingface.co/datasets/Precise-Debugging-Benchmarking/PDB-Single-Hard.tabulartext-generation1K<n<10K0 likes61 downloads5mo agoHugging Face30x5tne /pdbs-v1-resultsi benchmark my es, Pivot-Directional Basin Search, and i think it's better (dim: 10) Objective old pdbs pdbs-linear pdbs-quadratic PDBS-v1 sphere (mean) 3.2e-9 0.066 0.0025 2.6e-11 rosenbrock (mean) 0.36 22.8 94.1 (unstable) 0.50 rastrigin (mean) 40.0 (stuck) 40.0 (stuck) 1.87 16.1 ackley (mean) 1.76 6.59 (stuck) 0.073 6.8e-6 griewank (mean) 0.0070 0.0142 0.0048 0.0059 (dim:10, 20, 50) Objective Dim CMA-ES (mean fitness) CMA-ES time (s) PDBS-v1 (mean fitness)… See the full description on the dataset page: https://huggingface.co/datasets/x5tne/pdbs-v1-results.0 likes59 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.