datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
chemistry
CAMEL: Communicative Agents for “Mind” Exploration of Large Scale Language Model Society
Github: https://github.com/lightaime/camel
Website: https://www.camel-ai.org/
Arxiv Paper: https://arxiv.org/abs/2303.17760
Dataset Summary
Chemistry dataset is composed of 20K problem-solution pairs obtained using gpt-4. The dataset problem-solutions pairs generating from 25 chemistry topics, 25 subtopics for each topic and 32 problems for each "topic,subtopic" pairs.
We… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/chemistry.agieval-gaokao-chemistry
Dataset Card for "agieval-gaokao-chemistry"
Dataset taken from https://github.com/microsoft/AGIEval and processed as in that repo, following dmayhem93/agieval-* datasets on the HF hub.
This dataset contains the contents of the Gaokao Chemistry subtask of AGIEval, as accessed in https://github.com/ruixiangcui/AGIEval/commit/5c77d073fda993f1652eaae3cf5d04cc5fd21d40 .
Citation:
@misc{zhong2023agieval,
title={AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models}… See the full description on the dataset page: https://huggingface.co/datasets/hails/agieval-gaokao-chemistry.agieval-gaokao-chemistry
Dataset Card for "agieval-gaokao-chemistry"
Dataset taken from https://github.com/microsoft/AGIEval and processed as in that repo.
MIT License
Copyright (c) Microsoft Corporation.
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell… See the full description on the dataset page: https://huggingface.co/datasets/dmayhem93/agieval-gaokao-chemistry.camel_ai_chemistry_instruction_datasetChemistryQAChemistryQA is a complex QA task which cannot be solved by end-to-end neural networks. To answer chemical questions, machines need to understand questions, apply chemistry and math knowledge, and do calculation and reasoning. ChemistryQA contains about 4500 questions covering around 200 chemistry topics, which are collected from https://socratic.org/chemistry.
All credits go to chemistry-qa project by Microsoft (https://github.com/microsoft/chemistry-qa)
Trademarks
This project may contain… See the full description on the dataset page: https://huggingface.co/datasets/avaliev/ChemistryQA.Jee-Chemistry-dataset-with-COTlibgen_chemistryscience_chemistrystarcoder-chemistryChinese-High-School-Chemistry-Correction-Dataset
Chinese-High-School-Chemistry-Correction-Dataset
一个面向「高中化学垂直大模型微调」的中文问答与文本生成数据集
1. 数据集缘起
为了训练一个高中化学领域的垂直大模型,我们需要大量高质量、结构化的中文语料。本数据集整理了三版主流教科书、常考化学方程式与畅销教辅等中的知识点,全部转为统一的 JSONL 格式。
2. 数据来源
普通高中教科书(苏教版、人教版、鲁教版)、高中常考化学方程式、高中参考教辅资料(一本涂书、教材帮等)均转成jsonl格式
该jsonl文件数据,部分行或许有格式错误,需要自行编写py脚本校对,以便用于大模型微调。
3. 数据格式(JSONL)
每行一条记录,可直接用于 Hugging Face datasets 库:
{"instruction": "已知0.5 mol的水(H₂O)的质量是9 g,且含有3.01×10²³个水分子。请计算1 mol水的质量和阿伏伽德罗常数。", "output":… See the full description on the dataset page: https://huggingface.co/datasets/liushuaiqian/Chinese-High-School-Chemistry-Correction-Dataset.chemistry-sft-ultra
Chemistry SFT Ultra
Modern chemistry fine-tuning data built from multiple curated upstream datasets, merged into a single English corpus with reproducible processing and analysis.
Dataset Summary
This repository merges several instruction/QA-style sources into a single, cleaned, deduplicated, English-only training corpus in chat format.
The final corpus contains 1,370,322 rows. Each row is:
messages: a list of {role, content} message dicts (chat/SFT format)
metadata: a… See the full description on the dataset page: https://huggingface.co/datasets/summykai/chemistry-sft-ultra.chemistry-reasoning-datasetPDF_and_SCP_unfiltered_organic_chemistry_questionsHandwritten-Chemistry-Notes-Dataset
English Handwritten Chemistry Notes Dataset
This dataset contains high-resolution images of handwritten chemistry notes written in English. The collection includes equations, reaction mechanisms, periodic table references, structural diagrams, and descriptive explanations. It supports AI research in handwriting recognition, chemical structure understanding, and document analysis for STEM and educational applications.
Contact
For queries or collaborations related to this… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Chemistry-Notes-Dataset.Chemistrychemistry_qaCoT-chemistry-SFT
CoT-chemistry-SFT
Full chemistry chain-of-thought (CoT) dataset for supervised fine-tuning (SFT), generated by o4-mini.
This is the complete 1,606-example dataset. A 100-example public preview is available at Arminzd/CoT-O4_mini.
Dataset Details
Examples: 1,606
Generated by: o4-mini
Purpose: SFT training for chemistry tool-calling agents (tool-n1 project)
Fields
Field
Description
uid=3154455(arminzd) gid=3154455(arminzd)… See the full description on the dataset page: https://huggingface.co/datasets/Arminzd/CoT-chemistry-SFT.euro_pmc_chemistry_paperschemistry_textbookUniversity-level_Mathematics_Physics_Chemistry_Computer_Science_Reasoning_Corpus
Title
University-level Mathematics, Physics, Chemistry, Computer Science Reasoning Corpus
Size
200,000+ text+ multimodal university-level problems, each with step-by-step solutions and final answers
Format
Natural language explanations with multimodal samples include images (graphs, diagrams, etc.)
Subject
Mathematics, Physics, Chemistry, Computer Science
Labeling Details
Question ID/Question Stem (Full text/content) /Subject/Question Type… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/University-level_Mathematics_Physics_Chemistry_Computer_Science_Reasoning_Corpus.social_chemistry_101
Social Chemistry 101 (clean mirror)
Why. The original social_chem_101 dataset was removed from the Hub when Hugging Face
deprecated loading scripts, leaving only partial or translated mirrors. This is a faithful,
full mirror of the v1.0 release so it loads with a plain load_dataset, mainly so downstream
work (e.g. joining moral-foundation labels onto other datasets) has a stable source.
What. 356k rule-of-thumb (RoT) annotation rows over ~104k situations from Reddit and… See the full description on the dataset page: https://huggingface.co/datasets/wassname/social_chemistry_101.OpenSciReasoning-Chemistry-20K
OpenSciReasoning-Chemistry-20K
Three-domain release derived from nvidia/OpenScienceReasoning-2 for
domain-specific reasoner training and cross-domain transfer experiments.
Each row preserves the stable source_row_id and has exactly one mutually
exclusive domain value: CHEMISTRY. Domain acceptance was checked from the
question and choices with two independent question-only verifiers; answer and
source-ID gates were also replayed.
The audit records list any remaining source-output… See the full description on the dataset page: https://huggingface.co/datasets/TerryJCZhang/OpenSciReasoning-Chemistry-20K.scbe-chemistry-sft
Status: experimental. Research artifact, not a production candidate. Canonical dataset: scbe-aethermoore-training-data.
SCBE chemistry SFT
Chemistry adapter training data, split by purpose rather than one flat pile:
chemistry_adapter_invariants_v1_{train,eval}.sft.jsonl - conservation and
invariant rows
chemistry_adapter_verification_v1_{train,eval}.sft.jsonl - verification rows
chemistry_gate_repair_v1_{train,eval}.sft.jsonl - gate-repair rows
aligned_foundations_v2_{train… See the full description on the dataset page: https://huggingface.co/datasets/issdandavis/scbe-chemistry-sft.chemistry-knowledge
ChemBricks Knowledge
Does caffeine prefer water or an oil-like liquid?Why can adding one small group change a molecule's behavior?Can we design a molecule that interacts more favorably with water while meeting other constraints?How much energy does it take to remove an electron from a molecule?
These are the kinds of questions behind this dataset. Each investigation connects a question to recorded calculations, an answer, and the evidence needed to examine that answer.
Created… See the full description on the dataset page: https://huggingface.co/datasets/chembricks/chemistry-knowledge.C-MHChem-Benchmark-Chinese-Middle-high-school-Chemistry-Test
Introduction
C-MHChem-Benchmark-Chinese-Middle-high-school-Chemistry-Test is a High-quality single-choice full-human-writen Benchmark of 600 entries collected from Chinese Chemistry test of middle and high schools past 25 years.
C-MHChem 是一个包含了600个高质量的全人工编写的单选题测评基准,收集自过去25年间中国各地初高中中高考测试题目。
Citation
@misc{zhang2024chemllm,
title={ChemLLM: A Chemical Large Language Model},
author={Di Zhang and Wei Liu and Qian Tan and Jingdan Chen and Hang Yan and Yuliang Yan… See the full description on the dataset page: https://huggingface.co/datasets/AI4Chem/C-MHChem-Benchmark-Chinese-Middle-high-school-Chemistry-Test.HLE_SFT_Chemistry
HLE_SFT_Chemistry
データセットの説明
このデータセットは、以下の分割(split)ごとに整理された処理済みデータを含みます。
chempile: 1 JSON files, 1 Parquet files
データセット構成
各 split は JSON 形式と Parquet 形式の両方で利用可能です:
JSONファイル: 各 split 用サブフォルダ内の元データ(chempile/)
Parquetファイル: split名をプレフィックスとした最適化データ(data/chempile_*.parquet)
各 JSON ファイルには、同名の split プレフィックス付き Parquet ファイルが対応しており、大規模データセットの効率的な処理が可能です。
使い方
from datasets import load_dataset
# 特定の split を読み込む
chempile_data =… See the full description on the dataset page: https://huggingface.co/datasets/neko-llm/HLE_SFT_Chemistry.chemistry_stackexchange
Dataset Details
Dataset Description
Questions and answers mined from chemistry.stackexchange.com.
Curated by:
License: CC BY-SA
Dataset Sources
original data source
information about the license
Citation
BibTeX:
No citations provided
SciBench_Chemistrytextbookreasoning_chemistrychemistry
