CoolFace
Datasetpublic

monsoon-nlp/greenbeing-proteins

GreenBeing Proteins dataset Proteins from UniProtKB (knowledge base), from select food crops and related species. Amino acid sequences use IUPAC-IUB codes where letters A-Z map to amino acids. Usage (due to different schema on splits): load_dataset("monsoon-nlp/greenbeing-proteins", "pretraining", split="pretraining") XML source from https://www.uniprot.org/help/downloads CoLab notebook: https://colab.research.google.com/drive/1M6sO0Ws6i5z9VUXIXopiOqo1OkQ7K-1g?usp=sharing… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/greenbeing-proteins.

sourceHugging Faceccupdated 2y agoView on Hugging Face
3likes52downloads
settings

This repository belongs to monsoon-nlp on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namegreenbeing-proteins
visibilitypublic
licencecc
gatedno
ownermonsoon-nlp
Account settings
monsoon-nlp/greenbeing-proteins · CoolFace