datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
science_chemistrychemistry-knowledge
ChemBricks Knowledge
Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule?
These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer.
Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.OpenSciReasoning-Chemistry-20K
OpenSciReasoning-Chemistry-20K
Three-domain release derived from nvidia/OpenScienceReasoning-2 for
domain-specific reasoner training and cross-domain transfer experiments.
Each row preserves the stable source_row_id and has exactly one mutually
exclusive domain value: CHEMISTRY. Domain acceptance was checked from the
question and choices with two independent question-only verifiers; answer and
source-ID gates were also replayed.
The audit records list any remaining source-output… See the full description on the dataset page: https://huggingface.co/datasets/TerryJCZhang/OpenSciReasoning-Chemistry-20K.C-MHChem-Benchmark-Chinese-Middle-high-school-Chemistry-Test
Introduction
C-MHChem-Benchmark-Chinese-Middle-high-school-Chemistry-Test is a High-quality single-choice full-human-writen Benchmark of 600 entries collected from Chinese Chemistry test of middle and high schools past 25 years.
C-MHChem 是一个包含了600个高质量的全人工编写的单选题测评基准,收集自过去25年间中国各地初高中中高考测试题目。
Citation
@misc{zhang2024chemllm,
title={ChemLLM: A Chemical Large Language Model},
author={Di Zhang and Wei Liu and Qian Tan and Jingdan Chen and Hang Yan and Yuliang Yan… See the full description on the dataset page: https://huggingface.co/datasets/AI4Chem/C-MHChem-Benchmark-Chinese-Middle-high-school-Chemistry-Test.chemistry_2024-01-10_10.23.25validate_chemistry_questionUse the script generate_valid_questions.py to create an instruction set for valid questions.
python generate_valid_questions.py chemistry-by-chapter.txt valid_examples.json
Use the script generate_invalid_questions.py to create an instruction set for invalid questions.
python generate_invalid_questions.py politics.txt invalid_examples.json
Combine the two datasets.
echo -n "[" > finetune.json
cat valid_examples.json >> finetune.json
sed '$s/,$//' invalid_examples.json | cat >> finetune.json… See the full description on the dataset page: https://huggingface.co/datasets/juntaoyuan/validate_chemistry_question.chemistry-questionsbondshift-organic-chemistry
BondShift: Organic Chemistry Mechanism Dataset
10,000 ground-truth-separated records for mechanism diagnosis, misconception repair, and chemistry tutoring.
A narrow, auditable V1 dataset built from independently constructed scenario blueprints and deterministic answer keys.
TL;DR
BondShift addresses the right answer, wrong mechanism problem. It trains models to examine electron flow, formal
charge, intermediates, pathway choice, and stereochemical… See the full description on the dataset page: https://huggingface.co/datasets/prathmeshadsod/bondshift-organic-chemistry.SciTrust2-ChemistryQAorganic-chemistry-reasoning-demo
🧪 Organic Chemistry Reasoning Benchmark (OCRB-200) - DEMO
🛑 This is a DEMO version containing only 10 samples.
🚀 Want the Full Version (236+ Samples)?
The full dataset is available for purchase. It includes deep reasoning chains and trap identification for 236+ complex scenarios.
👉 [Download the Full Dataset Here] (https://7820367248654.gumroad.com/l/pkndc)
Chemistry_25k
WithinUsAI/Chemistry_25k — Master Scholars Academics (25k)
8,000 verification (TRUE/FALSE + correction)
9,000 self-contained quantitative items
8,000 definitions / micro-refreshers
Generated: 2026-01-04T05:02:58Z
chemistry-ptbr
Tradução do Camel Chemisty dataset para Portuguese (PT-BR) usando NLLB 3.3b.
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Chemistry dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 chemistry topics… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/chemistry-ptbr.first-finetuning-validate-chemistry-questionsUse the script generate_valid_questions.py to create an instruction set for valid questions.
python generate_valid_questions.py chemistry-by-chapter.txt valid_examples.json
Use the script generate_invalid_questions.py to create an instruction set for invalid questions.
python generate_invalid_questions.py politics.txt invalid_examples.json
Combine the two datasets.
echo -n "[" > finetune.json
cat valid_examples.json >> finetune.json
sed '$s/,$//' invalid_examples.json | cat >> finetune.json… See the full description on the dataset page: https://huggingface.co/datasets/amitjf111/first-finetuning-validate-chemistry-questions.mmlu-college-chemistryscience_chemistrymmlu-high-school-chemistryvalidate_chemistry_questionsRLT-physics_chemistry-expert-17kscience_chemistrychemistry-cleaned-solutionspodtech-chemistry
PodTech Chemistry — Chemistry Reasoning Question-Answer Pairs
A curated set of 101 challenging chemistry reasoning questions, each
with step-by-step reasoning and a single, one-step-verifiable answer. Questions
span organic, inorganic, and physical chemistry.
Dataset structure
Each record has the following fields:
field
type
description
id
int
Sequential identifier (1-based).
category
string
One of organic chemistry, inorganic chemistry, physical… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/podtech-chemistry.chemistry-tutor-datasetchemistry_2024-01-12_16.22.36chemistry-fine-tuning.jsonwhybook-chemistry-dataset
WhyBook Chemistry Dataset
This dataset contains synthetic NCERT-aligned chemistry tutoring records used for the WhyBook project.
Dataset Structure
Each record is a JSON object with the following keys:
concept
chapter
class
what
why
real_world
The writing style is designed for Indian school learners and focuses on:
simple English
concrete examples
chapter-grounded tutoring explanations
Format
Primary file:
data.jsonl
Each line contains one JSON object.… See the full description on the dataset page: https://huggingface.co/datasets/Stinger2311/whybook-chemistry-dataset.SciTrust2-Chemistrychemistry_arab_CoT_10K
Dataset Summary
The Arabic Chemistry Problem Dataset is a curated collection of Arabic-language chemistry exam questions.Each entry includes a question written in Modern Standard Arabic, its correct answer, a detailed explanation, and the underlying chemical concept.This dataset aims to support multilingual STEM reasoning, question answering, and educational AI applications across Arabic-speaking regions.
Each record follows a structured JSON format with the following fields:… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/chemistry_arab_CoT_10K.fine-tune-chemistrySurface_Chemistrycamel-ai_chemistry-ShareGPT
