yarongef/human_proteome_doublets
Dataset Description Out of 20,577 human proteins (from UniProt human proteome), sequences shorter than 20 amino acids or longer than 512 amino acids were removed, resulting in a set of 12,703 proteins. The uShuffle algorithm (python pacakge) was then used to shuffle these protein sequences while maintaining their doublet distribution. The very few sequences for which uShuffle failed to create a shuffled version were eliminated. Afterwards, h-CD-HIT algorithm (web server) was… See the full description on the dataset page: https://huggingface.co/datasets/yarongef/human_proteome_doublets.
Update README.md
Update README.md
Delete one_protein_sequence_per_gene_20577.fasta
Upload one_protein_sequence_per_gene_20577.fasta with git-lfs
Update README.md
Update README.md
Upload doublets_test_set.csv
Upload doublets_training_set.csv
initial commit
