datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MarkushGrapher-Datasets
This repository contains datasets introduced in MarkushGrapher: Joint Visual and Textual Recognition of Markush Structures.
Training:
MarkushGrapher-Synthetic-Training: This set contains synthetic Markush structures used for training MarkushGrapher. Samples are synthetically generated using the following steps: (1) SMILES to CXSMILES conversion using RDKit; (2) CXSMILES rendering using CDK; (3) text description generation using templates; and (4) text description augmentation with LLM.… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/MarkushGrapher-Datasets.MarkushGrapher-2-Datasets
MarkushGrapher 2 Datasets
Datasets for training and evaluating MarkushGrapher 2, a model for converting patent Markush structure images into CXSMILES representations.
Dataset Subsets
Subset
Train
Test
Description
OCR
uspto-mol-m-54k-new
54,785
200
USPTO-MOL-M Markush samples
ChemicalOCR predictions
uspto-markush
—
74
USPTO Markush structures benchmark
Ground Truth OCR
m2s
—
103
Mol2Smiles (M2S) benchmark
Ground Truth OCR
IP5-markush
—
878
IP5 Markush… See the full description on the dataset page: https://huggingface.co/datasets/docling-project/MarkushGrapher-2-Datasets.LLM-Jailbreak-Classifier
Dataset used to train various classifiers for LLM jailbreaks
Data with classification set to jailbreak is potential offensive / malicious (duh!)
Datasets used and cleaned:
Open-Orca/OpenOrca
ShawnMenz/DAN_jailbreak
EddyLuo/JailBreakV_28K
ShawnMenz/jailbreak_sft_rm_ds
https://raw.githubusercontent.com/verazuo/jailbreak_llms/main/data/jailbreak_prompts.csv
Next Steps
Enrich dataset with snythetic data (LLM generated) to improve classification
Generation… See the full description on the dataset page: https://huggingface.co/datasets/markush1/LLM-Jailbreak-Classifier.MarkushGrapher-2-Datasets
MarkushGrapher 2 Datasets
Datasets for training and evaluating MarkushGrapher 2, a model for converting patent Markush structure images into CXSMILES representations.
Dataset Subsets
Subset
Train
Test
Description
OCR
uspto-mol-m-54k-new
54,785
200
USPTO-MOL-M Markush samples
ChemicalOCR predictions
uspto-markush
—
74
USPTO Markush structures benchmark
Ground Truth OCR
m2s
—
103
Mol2Smiles (M2S) benchmark
Ground Truth OCR
IP5-markush
—
878
IP5 Markush… See the full description on the dataset page: https://huggingface.co/datasets/Mirinterplay/MarkushGrapher-2-Datasets.LLM-Jailbreak-Classifier-Largeharm_trainextracting_phylogeniesadversarial-insurance-questions
Dataset Card for Adversarial Financial Questions
This dataset is intended to be used for red teaming AI systems in the insurance services context, to verify that no actual financial advice is given through an AI model.
Dataset Details
Dataset Sources
Dataset questions, clusters and category names are artificially generated.
Uses
Red Teaming of AI systems to validate their alignment, not to give actual insurance advice.
Out-of-Scope Use… See the full description on the dataset page: https://huggingface.co/datasets/markush1/adversarial-insurance-questions.adversarial-banking-questions
Dataset Card for Adversarial Financial Questions
This dataset is intended to be used for red teaming AI systems in the financial services context, to verify that no actual financial advice is given through an AI model.
Dataset Details
Dataset Description
Curated by: Markus, Hupfauer
Language(s) (NLP): English
Dataset Sources
Dataset questions, clusters and category names are artificially generated.
Uses
Red Teaming of AI systems to… See the full description on the dataset page: https://huggingface.co/datasets/markush1/adversarial-banking-questions.
