Protein
OpenMed-NER-ProteinDetect-SuperClinical-141MOpenMed-NER-ProteinDetect-SnowMed-568MOpenMed-NER-ProteinDetect-TinyMed-135MOpenMed-NER-ProteinDetect-BioMed-109MOpenMed-NER-ProteinDetect-ElectraMed-560MOpenMed-NER-ProteinDetect-EuroMed-212MOpenMed-NER-ProteinDetect-BioClinical-108MOpenMed-NER-ProteinDetect-BioMed-335M
Datasets
All datasets matching “Protein”claude-protein-binder-design
Claude protein binder design — data release v1.0
1,440 de novo miniprotein binders (50 to 120 residues) against 16 targets, designed by two Claude models operating as autonomous protein-design agents (Mythos Preview, 900 designs; Opus 4.8, 540 designs) and characterized at two contract research organizations, Adaptyv Bio (cell-free expression; SPR/BLI kinetics with the design immobilized) and Twist Bioscience (Fc-fusion expression; capture SPR with a six-point antigen… See the full description on the dataset page: https://huggingface.co/datasets/Anthropic/claude-protein-binder-design.Protein-FN
Do Protein Transformers Have Biological Intelligence?
Figure 1: Illustration of two key protein motifs, i.e., His94-His96-His119and Ser29-His107-Tyr194, identified by our approach.
Deep neural networks, particularly Transformers, have been widely adopted for predicting the functional properties of proteins. In this work, we focus on exploring whether Protein Transformers can capture biological intelligence among protein sequences. To achieve our goal, we first introduce a protein… See the full description on the dataset page: https://huggingface.co/datasets/Protein-FN/Protein-FN.protein-docs
Protein Documents (Parquet)
Structured text documents encoding protein residue sequences and 3D contact maps from AlphaFold Database v4 predicted structures, stored as Parquet files. Each row is one protein document with metadata.
Source structures: timodonnell/afdb-24M and timodonnell/afdb-1.6M
Document Schemes
Each subdirectory contains documents generated with a different scheme. All schemes share leakage-resistant train/val/test splits based on structural… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/protein-docs.initial-dynamic-proteins
Initial 2,000-protein dataset
This is the canonical local root for the first complete router dataset: 1,000
single-dominant structured-state proteins (label 0) and 1,000 dynamic or
heterogeneous-state proteins (label 1). The fixed split is 1,400 train, 300
validation, and 300 test proteins.
Place Colab's completed ESMFold result files (<sequence_sha256>.npz) in
esmfold_results/. Then import them with:
uv run python scripts/esmfold_dataset.py import
The importer validates every… See the full description on the dataset page: https://huggingface.co/datasets/archiitecture/initial-dynamic-proteins.protein_data_testsplit 1, 2 -> for sequences
split 3, 4 -> for residues
Proteins
