CoolFace
Datasetpublic

yarongef/human_proteome_doublets

Dataset Description Out of 20,577 human proteins (from UniProt human proteome), sequences shorter than 20 amino acids or longer than 512 amino acids were removed, resulting in a set of 12,703 proteins. The uShuffle algorithm (python pacakge) was then used to shuffle these protein sequences while maintaining their doublet distribution. The very few sequences for which uShuffle failed to create a shuffled version were eliminated. Afterwards, h-CD-HIT algorithm (web server) was… See the full description on the dataset page: https://huggingface.co/datasets/yarongef/human_proteome_doublets.

sourceHugging Facemitupdated 4y agoView on Hugging Face
0likes18downloads
9 commits on main
4e2d0dd4y ago

Update README.md

yarongef
0a23dcd4y ago

Update README.md

yarongef
1319cb34y ago

Delete one_protein_sequence_per_gene_20577.fasta

yarongef
21911024y ago

Upload one_protein_sequence_per_gene_20577.fasta with git-lfs

yarongef
e7db1fa4y ago

Update README.md

yarongef
84e960c4y ago

Update README.md

yarongef
15a8e924y ago

Upload doublets_test_set.csv

yarongef
1eeb0b04y ago

Upload doublets_training_set.csv

yarongef
8bb15a24y ago

initial commit

yarongef