MichelNivard/proteinLM-mixed-pretraining-v1
Pretraining mix for Protein language models In order to construct a solid pre-training data mixture for protein language models we sample a mix of proteins from 3 sources: MG_Prot50 (https://huggingface.co/datasets/tattabio/OMG_prot50): meta-genomic proteins created by clustering the Open MetaGenomic dataset (OMG) at 50% sequence identity. UniRef50: UniProt proteins from across all species clustered to 50% sequences identity, downloaded form UniProt on 31th of March 2025… See the full description on the dataset page: https://huggingface.co/datasets/MichelNivard/proteinLM-mixed-pretraining-v1.
Upload train_10.tsv with huggingface_hub
Upload train_9.tsv with huggingface_hub
Update README.md
Update README.md
Upload train_8.tsv with huggingface_hub
Update README.md
Update README.md
Update README.md
Upload train_7.tsv with huggingface_hub
Update README.md
Update README.md
Update README.md
Upload train_6.tsv with huggingface_hub
Upload train_5.tsv with huggingface_hub
Upload train_4.tsv with huggingface_hub
Upload train_3.tsv with huggingface_hub
Upload train_2.tsv with huggingface_hub
Upload train_1.tsv with huggingface_hub
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Create README.md
initial commit
