datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
doublets
NuBerea Doublets and Parallels
Doublet (repeated-narrative) and parallel tracking across the Old and New Testaments. A doublet is the classic phenomenon of source and redaction criticism: the same story, saying, or psalm appearing more than once — J/E/P source strands narrating one event twice, shared psalms within the Psalter, synoptic gospel parallels, and typological Old-Testament-to-New-Testament echoes. Each locus gathers its parallel members, classifies the kind of… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/doublets.human_proteome_doublets
Dataset Description
Out of 20,577 human proteins (from UniProt human proteome), sequences shorter than 20 amino acids or longer than 512 amino acids were removed, resulting in a set of 12,703 proteins. The uShuffle algorithm (python pacakge) was then used to shuffle these protein sequences while maintaining their doublet distribution. The very few sequences for which uShuffle failed to create a shuffled version were eliminated.
Afterwards, h-CD-HIT algorithm (web server) was used… See the full description on the dataset page: https://huggingface.co/datasets/yarongef/human_proteome_doublets.
