datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chemistry
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Chemistry dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 chemistry topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/chemistry.Chinese-High-School-Chemistry-Correction-Dataset
Chinese-High-School-Chemistry-Correction-Dataset
一个面向「高中化学垂直大模型微调」的中文问答与文本生成数据集
1. 数据集缘起
为了训练一个高中化学领域的垂直大模型,我们需要大量高质量、结构化的中文语料。本数据集整理了三版主流教科书、常考化学方程式与畅销教辅等中的知识点,全部转为统一的 JSONL 格式。
2. 数据来源
普通高中教科书(苏教版、人教版、鲁教版)、高中常考化学方程式、高中参考教辅资料(一本涂书、教材帮等)均转成jsonl格式
该jsonl文件数据,部分行或许有格式错误,需要自行编写py脚本校对,以便用于大模型微调。
3. 数据格式(JSONL)
每行一条记录,可直接用于 Hugging Face datasets 库:
{"instruction": "已知0.5 mol的水(H₂O)的质量是9 g,且含有3.01×10²³个水分子。请计算1 mol水的质量和阿伏伽德罗常数。", "output":… See the full description on the dataset page: https://huggingface.co/datasets/liushuaiqian/Chinese-High-School-Chemistry-Correction-Dataset.chemistry-sft-ultra
Chemistry SFT Ultra
Modern chemistry fine-tuning data built from multiple curated upstream datasets, merged into a single English corpus with reproducible processing and analysis.
Dataset Summary
This repository merges several instruction/QA-style sources into a single, cleaned, deduplicated, English-only training corpus in chat format.
The final corpus contains 1,370,322 rows. Each row is:
messages: a list of {role, content} message dicts (chat/SFT format)
metadata: a… See the full description on the dataset page: https://huggingface.co/datasets/summykai/chemistry-sft-ultra.CoT-chemistry-SFT
CoT-chemistry-SFT
Full chemistry chain-of-thought (CoT) dataset for supervised fine-tuning (SFT), generated by o4-mini.
This is the complete 1,606-example dataset. A 100-example public preview is available at Arminzd/CoT-O4_mini.
Dataset Details
Examples: 1,606
Generated by: o4-mini
Purpose: SFT training for chemistry tool-calling agents (tool-n1 project)
Fields
Field
Description
uid=3154455(arminzd) gid=3154455(arminzd)… See the full description on the dataset page: https://huggingface.co/datasets/Arminzd/CoT-chemistry-SFT.chemistry-knowledge
ChemBricks Knowledge
Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule?
These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer.
Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.qwen3.6-35b-a3b-chemistry-benchmarks
Qwen3.6-35B-A3B Chemistry Benchmark Results
Raw outputs and scores from running Qwen3.6-35B-A3B (Q8_0 quant) through five published chemistry and biosecurity benchmarks, entirely on local hardware (two secondhand Tesla M40 24GB GPUs, no cloud compute). This is the raw data behind our blog post on locally reproducible AI capability evaluation, including the full per-item outputs, the parsing failures, and the negative results, not just the headline numbers.
Results at… See the full description on the dataset page: https://huggingface.co/datasets/CopyleftCultivars/qwen3.6-35b-a3b-chemistry-benchmarks.yher-chemistry-question-bank
YHer Chemistry Question Bank
The data layer of an evidence-bound diagnostic learning system for Shanghai high-school chemistry (Chris-TLC/YHer-skill).
Every record in this dataset is derived from publicly released Shanghai gaokao and mock examination papers through deterministic mechanical structuring: text extraction, layout repair, and answer alignment. No content is model-generated.
What's inside
The dataset ships in two configs:
Config
Records
Content… See the full description on the dataset page: https://huggingface.co/datasets/Chris-TLC/yher-chemistry-question-bank.ChemPref-DPO-for-Chemistry-data-en
Citation
@misc{zhang2024chemllm,
title={ChemLLM: A Chemical Large Language Model},
author={Di Zhang and Wei Liu and Qian Tan and Jingdan Chen and Hang Yan and Yuliang Yan and Jiatong Li and Weiran Huang and Xiangyu Yue and Dongzhan Zhou and Shufei Zhang and Mao Su and Hansen Zhong and Yuqiang Li and Wanli Ouyang},
year={2024},
eprint={2402.06852},
archivePrefix={arXiv},
primaryClass={cs.AI}
}
open-agh-chemistry-pl
Open AGH Polish chemistry textbooks
Four Polish editions: general chemistry, inorganic chemistry, polymer chemistry,
and corrosion/corrosion protection. This is an independent text-only adaptation.
Snapshot: 2026-09-05
Module occurrences: 374; unique modules: 374
Retained records: 370; measured tokens: 564028 (cl100k_base proxy)
Characters: 1446015
License: CC BY-SA 4.0, independently documented in each official EPUB rights page.
Immutable payload:… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/open-agh-chemistry-pl.task700_mmmlu_answer_generation_high_school_chemistry
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task700_mmmlu_answer_generation_high_school_chemistry
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task700_mmmlu_answer_generation_high_school_chemistry.task687_mmmlu_answer_generation_college_chemistry
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task687_mmmlu_answer_generation_college_chemistry
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task687_mmmlu_answer_generation_college_chemistry.quantum-simulation-chemistry-materials
Neura Parse — Quantum Simulation of Chemistry & Materials: Encodings, VQE/QPE & Dynamics
An application-deep, code-backed vertical on simulating quantum matter: electronic-structure problems, fermion-to-qubit encodings, Hamiltonian factorizations, ground/excited-state and real-time-dynamics algorithms, and analog simulation, with end-to-end resource estimates and honest classical-competitor accounting. Built with Qiskit Nature, OpenFermion, PennyLane-QChem, and PySCF — far… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-simulation-chemistry-materials.ChemPref-DPO-for-Chemistry-data-cn
Citation
@misc{zhang2024chemllm,
title={ChemLLM: A Chemical Large Language Model},
author={Di Zhang and Wei Liu and Qian Tan and Jingdan Chen and Hang Yan and Yuliang Yan and Jiatong Li and Weiran Huang and Xiangyu Yue and Dongzhan Zhou and Shufei Zhang and Mao Su and Hansen Zhong and Yuqiang Li and Wanli Ouyang},
year={2024},
eprint={2402.06852},
archivePrefix={arXiv},
primaryClass={cs.AI}
}
bondshift-organic-chemistry
BondShift: Organic Chemistry Mechanism Dataset
10,000 ground-truth-separated records for mechanism diagnosis, misconception repair, and chemistry tutoring.
A narrow, auditable V1 dataset built from independently constructed scenario blueprints and deterministic answer keys.
TL;DR
BondShift addresses the right answer, wrong mechanism problem. It trains models to examine electron flow, formal
charge, intermediates, pathway choice, and stereochemical… See the full description on the dataset page: https://huggingface.co/datasets/prathmeshadsod/bondshift-organic-chemistry.ChemistryConcepts-Instruct-v1
ChemistryConcepts-Instruct-v1
ChemistryConcepts-Instruct-v1 is a synthetic chemistry instruction dataset designed for supervised fine-tuning of language models on fundamental and advanced chemistry concepts. It covers diverse domains including fundamentals of matter and measurement, chemical interactions and reactions, energy and chemical change, and specialized areas of chemistry through clear explanations, chemical intuition, worked examples, and educational discussions. The… See the full description on the dataset page: https://huggingface.co/datasets/kd13/ChemistryConcepts-Instruct-v1.chemistry-ptbr
Tradução do Camel Chemisty dataset para Portuguese (PT-BR) usando NLLB 3.3b.
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Chemistry dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 chemistry topics… See the full description on the dataset page: https://huggingface.co/datasets/botbotrobotics/chemistry-ptbr.NCERT_Chemistry_11thNCERT_Chemistry_12thchemistry-organometallic-qa
chemistry-organometallic-qa
Dataset de 167 pares pregunta-respuesta en formato conversacional para fine-tuning de LLMs especializados en química organometálica. Construido sobre el paper PMC10967698 como parte del proyecto chem-rag-assistant.
Repositorio: chem-rag-assistantNotebook de construcción: 05_dataset.ipynb
Estructura del dataset
Formato messages compatible con el chat template de Qwen2.5-Instruct y modelos instruction-tuned similares:
{
"messages": [… See the full description on the dataset page: https://huggingface.co/datasets/Jesusrodriguezf90/chemistry-organometallic-qa.ChemPref-DPO-for-Chemistry-data-en
Citation
@misc{zhang2024chemllm,
title={ChemLLM: A Chemical Large Language Model},
author={Di Zhang and Wei Liu and Qian Tan and Jingdan Chen and Hang Yan and Yuliang Yan and Jiatong Li and Weiran Huang and Xiangyu Yue and Dongzhan Zhou and Shufei Zhang and Mao Su and Hansen Zhong and Yuqiang Li and Wanli Ouyang},
year={2024},
eprint={2402.06852},
archivePrefix={arXiv},
primaryClass={cs.AI}
}
podtech-chemistry
PodTech Chemistry — Chemistry Reasoning Question-Answer Pairs
A curated set of 101 challenging chemistry reasoning questions, each
with step-by-step reasoning and a single, one-step-verifiable answer. Questions
span organic, inorganic, and physical chemistry.
Dataset structure
Each record has the following fields:
field
type
description
id
int
Sequential identifier (1-based).
category
string
One of organic chemistry, inorganic chemistry, physical… See the full description on the dataset page: https://huggingface.co/datasets/Podtech/podtech-chemistry.organic-chemistry-synthesis-planning-corpus
Organic Chemistry Synthesis Planning Corpus
Status: actively ingesting. A comprehensive reaction backbone is already uploaded
(millions of reactions; see Ingested data below). Curation and
additional sources are ongoing. See Roadmap.
Quickstart (for students / first-time users)
You need a free Hugging Face account, and to accept this dataset's terms on its page (it's gated).
pip install datasets transformers
huggingface-cli login
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Solshine/organic-chemistry-synthesis-planning-corpus.whybook-chemistry-dataset
WhyBook Chemistry Dataset
This dataset contains synthetic NCERT-aligned chemistry tutoring records used for the WhyBook project.
Dataset Structure
Each record is a JSON object with the following keys:
concept
chapter
class
what
why
real_world
The writing style is designed for Indian school learners and focuses on:
simple English
concrete examples
chapter-grounded tutoring explanations
Format
Primary file:
data.jsonl
Each line contains one JSON object.… See the full description on the dataset page: https://huggingface.co/datasets/Stinger2311/whybook-chemistry-dataset.chemistry_arab_CoT_10K
Dataset Summary
The Arabic Chemistry Problem Dataset is a curated collection of Arabic-language chemistry exam questions.Each entry includes a question written in Modern Standard Arabic, its correct answer, a detailed explanation, and the underlying chemical concept.This dataset aims to support multilingual STEM reasoning, question answering, and educational AI applications across Arabic-speaking regions.
Each record follows a structured JSON format with the following fields:… See the full description on the dataset page: https://huggingface.co/datasets/Mobiusi/chemistry_arab_CoT_10K.camel-ai_chemistry-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai_chemistry-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
camel-ai/chemistry with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/camel-ai_chemistry-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.chemistry
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Chemistry dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 chemistry topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We… See the full description on the dataset page: https://huggingface.co/datasets/sdsdwww/chemistry.
