CoolFace
Datasetpublic

monsoon-nlp/protein-pairs-uniprot-swissprot

Protein Pairs and Similarity Selected protein similarities within training, test, and validation sets. Each protein gets two similarities selected at random and (usually) proteins within the top and bottom quintiles for similarity. The protein is represented by its UniProt ID and its amino acid sequence (using IUPAC-IUB codes where each amino acid maps to a letter of the alphabet, see: https://en.wikipedia.org/wiki/FASTA_format ). The distance column is cosine distance… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/protein-pairs-uniprot-swissprot.

sourceHugging Faceccupdated 3y agoView on Hugging Face
2likes73downloads
6 commits on main
5b357af3y ago

Update README.md

monsoon-nlp
628228d3y ago

Upload fullseq_train.csv

monsoon-nlp
233e20a3y ago

Update README.md

monsoon-nlp
e52e9c03y ago

Update README.md

monsoon-nlp
e44efc53y ago

add test and validation

monsoon-nlp
1de108f3y ago

initial commit

monsoon-nlp