karin
Datasets
All datasets matching “karin”ATOMICA
Dataset Card: Atomica Molecular Interactions | Sequence Data
Summary
A dataset of atomic‑scale molecular interaction interfaces, ready for ML workflows. Each row contains up to five interacting sequences (or SMILES), their modalities, an example ID, and the original data split.
Citation
@article{
Fang2025ATOMICA,
author = {Fang, Ada and Zhang, Zaixi and Zhou, Andrew and Zitnik, Marinka},
title = {ATOMICA: Learning Universal Representations of… See the full description on the dataset page: https://huggingface.co/datasets/karina-zadorozhny/ATOMICA.M320M
M^3-20M: A Large-Scale Multi-Modal Molecule Dataset for AI-driven Drug Design and Discovery
This repository hosts the M^3-20M multi-modal molecular dataset as a parquet table. Don't forget to check out the original HuggingFace repository for this dataset: https://huggingface.co/datasets/Alex99Gsy/M-3_Multi-Modal-Molecule .
M3-20M is a large-scale multi-modal molecular dataset with over 20 million molecules, integrating SMILES, molecular graphs, 3D structures, physicochemical… See the full description on the dataset page: https://huggingface.co/datasets/karina-zadorozhny/M320M.PeptideAtlas
Dataset Card for Human PeptideAtlas 2025-01 Peptides
Dataset Summary
Random split of the distinct peptide sequences released in the Peptide sequences in FASTA format file (125 MB) from the Human PeptideAtlas 2025-01 build.Each record is just the amino-acid sequence (uppercase 20-AA alphabet + “U”, “O”). No headers, spectra, or metadata are included.
Reference
Desiere et al., "The PeptideAtlas Project", Nucleic Acids Research, 2006, 34, D655-D658
karin_bluearchive
Dataset of karin/角楯カリン/花凛 (Blue Archive)
This is the dataset of karin/角楯カリン/花凛 (Blue Archive), containing 500 images and their tags.
The core tags of this character are black_hair, dark-skinned_female, dark_skin, long_hair, yellow_eyes, breasts, large_breasts, halo, very_long_hair, bow, ponytail, blue_bow, animal_ears, rabbit_ears, fake_animal_ears, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...), the auto-crawling system is… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/karin_bluearchive.slovo-full-landmarks
Slovo landmarks base
Base landmarks dataset built from full Slovo videos with MediaPipe Tasks HolisticLandmarker.
moleculeace
MoleculeACE Dataset
Overview
The MoleculeACE (Molecule Activity Cliff Estimation) dataset contains bioactivity data for 30 different ChEMBL targets with corresponding protein sequences. This dataset is designed to assess how well molecular machine learning models can handle activity cliffs - pairs of molecules that are structurally similar but show large differences in biological activity.
Dataset Structure
The dataset is organized with each ChEMBL target as a… See the full description on the dataset page: https://huggingface.co/datasets/karina-zadorozhny/moleculeace.
