protein2text
protein2text-swissprot-split
Protein2Text Dataset.
This dataset contains the aggregated protein annotations for all SwissProt proteins
as well as GSMC (microprotein database) and nmpfamsdb (metagenomic database).
Swiss-Prot proteins are the only proteins that have a train/text split.
738,769 records across three sources
562,258 Swiss-Prot entries, 517,257 train / 45,001 test
4,796 homology-isolated proteins tagged as the dark-protein eval set (these are a subset of the Swiss-Prot test set)… See the full description on the dataset page: https://huggingface.co/datasets/tumorailab/protein2text-swissprot-split.Protein2Text-QA
Protein2Text-QA Dataset
The Protein2Text-QA dataset is designed to generate human-readable explanations for protein functions based on protein sequences. It consists of question-answer (QA) pairs generated from PubMed Central (PMC) articles using LLaMA3.1-8B-Instruct. The dataset is structured into different subsets tailored for pretraining, fine-tuning, and evaluation.
Dataset Overview
Size: ~210,000 QA pairs
Source: UniProt (pretraining), PubMed Central (PMC) (QA… See the full description on the dataset page: https://huggingface.co/datasets/tumorailab/Protein2Text-QA.
