Flogrammer/Mol-JEPA-dataset
Mol-JEPA Dataset Multimodal molecular dataset used to train Mol-JEPA (a multimodal Joint Embedding Predictive Architecture for molecules). Each row of metadata.csv describes one molecule (SMILES + InChIKey + source dataset + labels) and points to precomputed per-modality embedding/target files stored as NumPy arrays. Modalities These are the modalities included (note that not every modality is available for every row - there is quite some sparsity). For detailed… See the full description on the dataset page: https://huggingface.co/datasets/Flogrammer/Mol-JEPA-dataset.
01.2k
