datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chemreason_retroSource: https://huggingface.co/datasets/liuganghuggingface/Llamole-MolQA
We parsed from original dataset.
Drug
9986 -> 5600
Material
750 -> 479
DD100
Dataset Card
Overview
DrugSeeker-mini benchmark is a streamlined evaluation dataset for end-to-end drug discovery processes, aggregating question-answering and classification tasks from multiple authoritative public data sources, totaling 91 queries that cover three major phases of drug discovery: Target Identification (TI), Hit Lead Discovery (HLD), and Lead Optimization (LO). Each query contains clear input/output descriptions, standard answers, and matching strategies… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-AI4S/DD100.
