ocsr
Datasets
All datasets matching “ocsr”OCSR-Benchmarks
OCSR Benchmarks
A collection of ten benchmark datasets for Optical Chemical Structure Recognition (OCSR) —
the task of converting chemical structure diagram images into machine-readable SMILES strings.
These benchmarks were used to evaluate the COMO model
(Closed-Loop Optical Molecule Recognition).
Subsets
Config
Split
Size
Domain
CLEF
test
992
Real
JPO
test
449
Real
UOB
test
5,740
Real
USPTO
test
5719
Real
USPTO-10K
test
9,999
Real
Staker
test
50,000… See the full description on the dataset page: https://huggingface.co/datasets/Keylab/OCSR-Benchmarks.OCSR-DeepSMILESCLEF_OCSR_benchmark
Dataset Card for CLEF-IP 2012 Structure Recognition Test Set
This dataset is the official test set from the "Chemical Structure Recognition in Images" task held as part of the CLEF-IP 2012 workshop. It contains 992 images of chemical structures extracted from US patents, each paired with a ground-truth MOL file. This Hugging Face version has been augmented with canonical SMILES, InChI, and SELFIES strings to provide a comprehensive resource for evaluating image-to-structure models… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/CLEF_OCSR_benchmark.Colored_Background_OCSR_benchmark
Dataset Card for Alpha-Extractor OCSR Benchmark
This dataset is a small, challenging benchmark for Optical Chemical Structure Recognition (OCSR), featuring 200 images. The data is a subset of the test data created for the "αExtractor" system, a tool specifically designed to perform well on difficult images, including those with backgrounds, low quality, and even hand-drawn structures. It serves as an excellent test of model robustness.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/Colored_Background_OCSR_benchmark.UOB_OCSR_benchmark
Dataset Card for UOB OCSR Benchmark
This dataset is a benchmark for Optical Chemical Structure Recognition (OCSR), containing 5,740 images of chemical structures. Originally created by the University of Birmingham, each image is paired with its corresponding MOL file. This version has been augmented with canonical SMILES, InChI, and SELFIES strings to provide a comprehensive resource for training and evaluating image-to-structure models.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/UOB_OCSR_benchmark.JPO_OCSR_benchmark
Dataset Card for Japanese Patent Office OCSR Benchmark
This dataset is a challenging benchmark for Optical Chemical Structure Recognition (OCSR), consisting of 450 images of organic molecules sourced from the Japanese Patent Office (JPO). It is a subset of the larger Chem-Infty dataset and is notable for its inclusion of noisy images, irregular features, and embedded text labels (sometimes in Japanese). This Hugging Face version includes the ground truth SDF files and adds derived… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/JPO_OCSR_benchmark.
