markush
Datasets
All datasets matching “markush”MarkushGrapher-Datasets
This repository contains datasets introduced in MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures.
Training:
MarkushGrapher-Synthetic-Training: This set contains synthetic Markush structures used for training MarkushGrapher. Samples are synthetically generated using the following steps: (1) SMILES to CXSMILES conversion using RDKit; (2) CXSMILES rendering using CDK; (3) text description generation using templates; and (4) text description augmentation with LLM.… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/MarkushGrapher-Datasets.MarkushGrapher-2-Datasets
MarkushGrapher 2 Datasets
Datasets for training and evaluating MarkushGrapher 2, a model for converting patent Markush structure images into CXSMILES representations.
Dataset Subsets
Subset
Train
Test
Description
OCR
uspto-mol-m-54k-new
54,785
200
USPTO-MOL-M Markush samples
ChemicalOCR predictions
uspto-markush
—
74
USPTO Markush structures benchmark
Ground Truth OCR
m2s
—
103
Mol2Smiles (M2S) benchmark
Ground Truth OCR
IP5-markush
—
878
IP5 Markush… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/MarkushGrapher-2-Datasets.LLM-Jailbreak-Classifier
Dataset used to train various classifiers for LLM jailbreaks
Data with classification set to jailbreak is potential offensive / malicious (duh!)
Datasets used and cleaned:
Open-Orca/OpenOrca
ShawnMenz/DAN_jailbreak
EddyLuo/JailBreakV_28K
ShawnMenz/jailbreak_sft_rm_ds
https://raw.githubusercontent.com/verazuo/jailbreak_llms/main/data/jailbreak_prompts.csv
Next Steps
Enrich dataset with snythetic data (LLM generated) to improve classification
Generation… See the full description on the dataset page: https://huggingface.co/datasets/markush1/LLM-Jailbreak-Classifier.MarkushGrapher-2-Datasets
MarkushGrapher 2 Datasets
Datasets for training and evaluating MarkushGrapher 2, a model for converting patent Markush structure images into CXSMILES representations.
Dataset Subsets
Subset
Train
Test
Description
OCR
uspto-mol-m-54k-new
54,785
200
USPTO-MOL-M Markush samples
ChemicalOCR predictions
uspto-markush
—
74
USPTO Markush structures benchmark
Ground Truth OCR
m2s
—
103
Mol2Smiles (M2S) benchmark
Ground Truth OCR
IP5-markush
—
878
IP5 Markush… See the full description on the dataset page: https://huggingface.co/datasets/Mirinterplay/MarkushGrapher-2-Datasets.LLM-Jailbreak-Classifier-Largeharm_train
