CoolFace
20 results

esm

LuminScience /LuminBench-Nano-ESMC LuminBench Nano ESMC Full Open Reservoir v2 This is the complete decontaminated 70%-identity representative reservoir for Lumin-Science/LuminBench-Nano-ESMC. It is organized as immutable, SHA-ordered Parquet shards so each run can download only the smallest deterministic prefix required by its training budget. License and source terms Lumin Science's original database selection, arrangement, decontamination ledger, packing, and metadata are offered under CC BY-SA… See the full description on the dataset page: https://huggingface.co/datasets/LuminScience/LuminBench-Nano-ESMC.100M<n<1B1 likes1.2k downloads13d agoHugging FaceEsmaeil9ss /Tickdatatext100K<n<1M2 likes1.2k downloads2h agoHugging Facenvidia /esm2_uniref_pretraining_data ESM-2 Uniref Pretraining Data Dataset Description: UniRef, or UniProt Reference Clusters, are databases of clustered protein sequences from the UniProt Knowledgebase (UniProtKB) that group similar sequences to reduce redundancy and make data easier to work with for biological research. It offers different levels of clustering (UniRef100, UniRef90, and UniRef50) based on sequence identity, with each cluster containing a representative sequence, a count of member proteins… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/esm2_uniref_pretraining_data.textfill-mask100M<n<1B9 likes980 downloads1y agoHugging Facefredzzp /esm-teddymer-pseudodimers ESM-Teddymer pseudo-dimers 60,177,402 intra-chain domain pairs ("pseudo-dimers") cut out of ESM Metagenomic Atlas monomers at Chainsaw/TED domain boundaries. Each row is a target domain and a binder domain that were adjacent in one real folded chain, so the pair comes with a real interface without anyone having to dock anything. Built to train target-conditioned binder-design models. All-atom structures for both chains ship alongside as foldcomp. What is in here… See the full description on the dataset page: https://huggingface.co/datasets/fredzzp/esm-teddymer-pseudodimers.tabularother100M<n<1B0 likes362 downloads23d agoHugging FaceESmike /city_temperature_anomaliestabular1M<n<10M2 likes342 downloads10mo agoHugging Facebiohub /ESMC-SAE-Features ESMC Sparse Autoencoder Features Table This dataset contains a Parquet table of the 16,384 features from the ESMC-6B-sae-layer60-k64-codebook16384, that was used for analysis in the ESMC paper and to construct the ESM Atlas. This table provides descriptions of the precomputed features that can be activated through the spotlight SAE model, assisting users for downstream interpretation of the insights revealed by ESMC. Download the table here. The features descriptions are in the… See the full description on the dataset page: https://huggingface.co/datasets/biohub/ESMC-SAE-Features.tabular10K<n<100K5 likes324 downloads4mo agoHugging Face