CoolFace
Datasetpublic

monsoon-nlp/greenbeing-proteins

GreenBeing Proteins dataset Proteins from UniProtKB (knowledge base), from select food crops and related species. Amino acid sequences use IUPAC-IUB codes where letters A-Z map to amino acids. Usage (due to different schema on splits): load_dataset("monsoon-nlp/greenbeing-proteins", "pretraining", split="pretraining") XML source from https://www.uniprot.org/help/downloads CoLab notebook: https://colab.research.google.com/drive/1M6sO0Ws6i5z9VUXIXopiOqo1OkQ7K-1g?usp=sharing… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/greenbeing-proteins.

sourceHugging Faceccupdated 2y agoView on Hugging Face
3likes53downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
monsoon-nlp/greenbeing-proteins · CoolFace