CoolFace
Datasetpublic

jablonkagroup/chempile-lift

ChemPile-LIFT A comprehensive dataset for chemistry property prediction using large language models πŸ“‹ Dataset Summary ChemPile-LIFT is a dataset designed for chemistry property prediction tasks, specifically focusing on the prediction of chemical properties using large language models (LLMs). It is part of the ChemPile project, which aims to create a comprehensive collection of chemistry-related data for training LLMs. The dataset includes a variety of… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/chempile-lift.

sourceHugging Facecc-by-nc-sa-4.0updated 9mo agoView on Hugging Face
7likes8.8kdownloads
Dataset Card

ChemPile-LIFT

<div align="center">

[image]

![Dataset](https://huggingface.co/datasets/jablonkagroup/chempile-lift) ![License: CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/) ![Paper](https://arxiv.org/abs/2505.12534) ![Website](https://chempile.lamalab.org/)

A comprehensive dataset for chemistry property prediction using large language models

</div>

πŸ“‹ Dataset Summary

ChemPile-LIFT is a dataset designed for chemistry property prediction tasks, specifically focusing on the prediction of chemical properties using large language models (LLMs). It is part of the ChemPile project, which aims to create a comprehensive collection of chemistry-related data for training LLMs. The dataset includes a variety of chemical properties embedded in text using SMILES (Simplified Molecular Input Line Entry System) strings for represent the molecules. The dataset is structured to facilitate the training of LLMs in predicting chemical properties based on textual descriptions.

The origin of the dataset property data is from well-known chemistry datasets such as the QM9 dataset, which contains quantum mechanical properties of small organic molecules, and the RDKit dataset, which includes a wide range of chemical properties derived from molecular structures. Each of the subsets or Hugging Face configurations corresponds to a different source of chemical property data, allowing for diverse training scenarios.

πŸ“Š Dataset Statistics

The resulting dataset contains 263.8M examples across all configurations, with a total size of 35.5B tokens.

πŸ—‚οΈ Dataset Configurations

The dataset includes diverse configurations covering various chemical property domains:

🧬 Drug Discovery and ADMET Properties

  • β€”BACE, BBBP, bioavailabilitymaetal, bloodbrainbarriermartinsetal, caco2wang, clearanceastrazeneca, cyp2c9substratecarbonmangels, cyp2d6substratecarbonmangels, cyp3a4substratecarbonmangels, cypp450*inhibitionveithetal series, druginducedliverinjury, halflifeobach, humanintestinalabsorption, lipophilicity, pampancats, volumeofdistributionatsteadystatelombardoetal, freesolv

⚠️ Toxicology and Safety Assessment

  • β€”amesmutagenicity, carcinogens, ld50catmos, ld50zhu, Tox21 initiative datasets (nrahrtox21, nrarlbdtox21, nrartox21, nraromatasetox21, nrertox21, nrppargammatox21, sraretox21, sratad5tox21, srhsetox21, srmmptox21, srp53tox21), hERG channel datasets (hergblockers, hergcentralat10uM, hergcentralat1uM, hergcentralinhib, hergkarimet_al), clintox, SIDER

🎯 Bioactivity and Target Interaction

  • β€”MUV series (MUV466, MUV548, MUV600, MUV644, MUV652, MUV689, MUV692, MUV712, MUV713, MUV733, MUV737, MUV810, MUV832, MUV846, MUV852, MUV858, MUV859), Butkiewicz collection (cav3t-typecalciumchannelsbutkiewicz, cholinetransporterbutkiewicz, kcnq2potassiumchannelbutkiewicz, m1muscarinicreceptoragonistsbutkiewicz, orexin1receptorbutkiewicz, potassiumionchannelkir21butkiewicz), hiv, sarscov23clprodiamond, sarscov2vitro_touret

πŸ”¬ Materials Science and Physical Properties

  • β€”formationenergies, flashpoint, meltingpoints, qm8, qm9, metal-organic frameworks (coremofnotopo, mofdscribe, qmofgcmc, qmofquantum), perovskitedb, polymer datasets (blockpolymersmorphology, biceranodataset), crystallographic data (oqmd, nomadstructure), rdkit_features

βš—οΈ Chemical Reactions and Synthesis

  • β€”Open Reaction Database collections (ordpredictions, ordrxnsmilesprocedure, ordrxnsmilesyieldpred, ordstepsyield), RHEA database (rheadbmasked, rheadbpredictions), patent chemistry (uspto, usptoyield), buchwaldhartwig

πŸ“– Biomedical Text Mining and Chemical Entity Recognition

  • β€”Named entity recognition (bioner series, bc5chem, bc5disease, chemdner, nlmchem, ncbidisease), chebi20, protein information (uniprotbindingsingle, uniprotbindingsitesmultiple, uniprotorganisms, uniprotreactions, uniprotsentences), msds, drugchatliangzhanget_al, mona

🧬 Peptide and Protein Properties

  • β€”peptideshemolytic, peptidesnonfouling, peptides_soluble

πŸ€– Machine Learning Benchmarks and Evaluation

  • β€”RedDB, oddoneout, materials property descriptions (mpdescriptions, mpself_supervised)

πŸ“œ License

All content is made open-source under the CC BY-NC-SA 4.0 license, allowing for:

  • β€”βœ… Non-commercial use and sharing with attribution
  • β€”βœ… Modification and derivatives
  • β€”βš οΈ Attribution required
  • β€”βš οΈ Non-commercial use only

πŸ“– Data Fields

The dataset contains the following fields for all the configurations allowing for a consistent structure across different chemical property datasets:

  • β€”`text`: The textual representation of the chemical property or question. It includes the input question or prompt related to the chemical property, often formatted as a natural language query, and the correct answer.

πŸ”¬ Dataset Groups Detailed Description

πŸ’Š Drug Discovery and ADMET Properties

The Drug Discovery and ADMET (Absorption, Distribution, Metabolism, Excretion, Toxicity) group encompasses datasets focused on pharmaceutical compound development and safety assessment. This collection includes bioavailability prediction (bioavailabilitymaetal), blood-brain barrier permeability (BBBP, bloodbrainbarriermartinsetal), intestinal permeability (caco2wang, pampancats), hepatic clearance (clearanceastrazeneca), drug-induced liver injury assessment (druginducedliverinjury), and various CYP450 enzyme interactions (cyp2c9substratecarbonmangels, cyp2d6substratecarbonmangels, cyp3a4substratecarbonmangels, cypp450*inhibitionveithetal series). The group also covers crucial pharmacokinetic properties such as half-life prediction (halflifeobach), human intestinal absorption (humanintestinalabsorption), lipophilicity (lipophilicity), volume of distribution (volumeofdistributionatsteadystatelombardoetal), and solubility (freesolv). These datasets are essential for early-stage drug development, helping predict whether compounds will have favorable drug-like properties before expensive clinical trials.

⚠️ Toxicology and Safety Assessment

The Toxicology and Safety Assessment group focuses on predicting harmful effects of chemical compounds across various biological systems. This includes mutagenicity prediction (amesmutagenicity), carcinogenicity assessment (carcinogens), acute toxicity (ld50catmos, ld50zhu), and comprehensive toxicity screening through the Tox21 initiative datasets (nrahrtox21, nrarlbdtox21, nrartox21, nraromatasetox21, nrertox21, nrppargammatox21, sraretox21, sratad5tox21, srhsetox21, srmmptox21, srp53tox21). The group also includes specialized toxicity assessments such as hERG channel blocking potential (hergblockers, hergcentralat10uM, hergcentralat1uM, hergcentralinhib, hergkarimet_al), clinical toxicity prediction (clintox), and adverse drug reactions (SIDER). These datasets are crucial for environmental safety assessment and pharmaceutical safety profiling.

🎯 Bioactivity and Target Interaction

The Bioactivity and Target Interaction group contains datasets focused on molecular interactions with specific biological targets. This includes enzyme inhibition studies (BACE for Ξ²-secretase), receptor binding and modulation data from the Butkiewicz collection (cav3t-typecalciumchannelsbutkiewicz, cholinetransporterbutkiewicz, kcnq2potassiumchannelbutkiewicz, m1muscarinicreceptoragonistsbutkiewicz, orexin1receptorbutkiewicz, potassiumionchannelkir21butkiewicz), and comprehensive screening datasets like MUV (Maximum Unbiased Validation) series covering various protein targets. The group also includes antiviral activity data (hiv, sarscov23clprodiamond, sarscov2vitrotouret) and specialized bioactivity assessments. These datasets enable the development of target-specific therapeutics and understanding of molecular mechanisms of action.

πŸ”¬ Materials Science and Physical Properties

The Materials Science and Physical Properties group encompasses datasets for predicting fundamental chemical and physical characteristics of compounds and materials. This includes thermodynamic properties (formationenergies, flashpoint, meltingpoints), quantum mechanical properties (qm8, qm9), and specialized materials databases such as metal-organic frameworks (coremofnotopo, mofdscribe, qmofgcmc, qmofquantum), perovskite materials (perovskitedb), polymer morphology (blockpolymersmorphology, biceranodataset), and crystallographic data (oqmd, nomadstructure). The group also covers molecular descriptors and features (rdkit_features) that are fundamental for computational chemistry and materials design applications.

βš—οΈ Chemical Reactions and Synthesis

The Chemical Reactions and Synthesis group focuses on datasets related to chemical reaction prediction, optimization, and mechanistic understanding. This includes the Open Reaction Database collections (ordpredictions, ordrxnsmilesprocedure, ordrxnsmilesyieldpred, ordstepsyield), reaction classification and prediction from the RHEA database (rheadbmasked, rheadbpredictions), patent chemistry data (uspto, usptoyield), and specialized reaction datasets like Buchwald-Hartwig coupling reactions (buchwaldhartwig). These datasets are essential for developing AI-driven synthetic chemistry tools and understanding reaction mechanisms.

πŸ“– Biomedical Text Mining and Chemical Entity Recognition

The Biomedical Text Mining and Chemical Entity Recognition group contains datasets designed for natural language processing tasks in chemistry and biology domains. This includes named entity recognition datasets (bioner series, bc5chem, bc5disease, chemdner, nlmchem, ncbidisease), chemical database integration (chebi20), protein and biological pathway information (uniprotbindingsingle, uniprotbindingsitesmultiple, uniprotorganisms, uniprotreactions, uniprotsentences), material safety data sheets processing (msds), conversational chemistry datasets (drugchatliangzhanget_al), and specialized chemical databases (mona for mass spectrometry data). These datasets enable the development of AI systems that can extract, understand, and reason about chemical information from scientific literature and databases.

🧬 Peptide and Protein Properties

The Peptide and Protein Properties group focuses on datasets specifically designed for understanding and predicting properties of peptides and proteins. This includes peptide-specific properties such as hemolytic activity (peptideshemolytic), antifouling characteristics (peptidesnonfouling), and solubility prediction (peptides_soluble). These datasets are particularly valuable for developing peptide-based therapeutics and understanding protein-protein interactions in biological systems.

πŸ€– Machine Learning Benchmarks and Evaluation

The Machine Learning Benchmarks and Evaluation group contains specialized datasets designed for testing and evaluating machine learning models in chemistry. This includes the RedDB dataset for chemical space exploration and benchmarking, the oddoneout dataset for anomaly detection and pattern recognition tasks, and materials property description datasets (mpdescriptions, mpself_supervised) that combine textual descriptions with materials properties for multimodal learning approaches. These datasets are essential for developing robust AI models that can handle diverse chemical and materials science challenges and for establishing standardized evaluation protocols in computational chemistry.

πŸš€ Usage

python
from datasets import load_dataset, get_dataset_config_names

# Print available configs for the dataset
configs = get_dataset_config_names("jablonkagroup/chempile-lift")
print(f"Available configs: {configs}")
# Available configs: ['BACE', 'BBBP', 'USPTO...

dataset = load_dataset("jablonkagroup/chempile-lift", name=configs[0])
# Loading config: BACE-completion_0

print(dataset)
# DatasetDict({
#     train: Dataset({
#         features: ['text', 'input', 'output', 'answer_choices', 'correct_output_index'],
#         num_rows: 1142
#     })
#     test: Dataset({
#         features: ['text', 'input', 'output', 'answer_choices', 'correct_output_index'],
#         num_rows: 118
#     })
#     val: Dataset({
#         features: ['text', 'input', 'output', 'answer_choices', 'correct_output_index'],
#         num_rows: 253
#     })
# })

split_name = list(dataset.keys())[0]
sample = dataset[split_name][0]
print(sample)
# {
#     'text': 'The compound with the SMILES of O1CC[C@@H](NC(=O)[C@@H]...',
# }

🎯 Use Cases

  • β€”πŸ§ͺ Chemical Property Prediction: Training models for predicting molecular properties and characteristics
  • β€”πŸ’Š Drug Discovery: Building systems for pharmaceutical compound screening and optimization
  • β€”βš οΈ Safety Assessment: Developing models for toxicity and environmental impact prediction
  • β€”πŸ”¬ Materials Design: Creating AI tools for materials science and property prediction
  • β€”πŸ“– Scientific Text Understanding: Training models to understand and reason about chemical information

⚠️ Limitations & Considerations

  • β€”Scope: Focused on chemistry and materials science; domain-specific terminology and concepts
  • β€”Quality: Variable quality across sources; expert curation applied but some noise may remain
  • β€”Bias: Reflects biases present in chemical databases and scientific literature
  • β€”License: Non-commercial use only under CC BY-NC-SA 4.0
  • β€”Language: Primarily English content
  • β€”Completeness: Some datasets may have missing values or incomplete property annotations

πŸ› οΈ Data Processing Pipeline

  1. 1.Collection: Automated extraction from well-known chemistry datasets (QM9, RDKit, Tox21, etc.)
  2. 2.Standardization: Consistent formatting and SMILES representation across all configurations
  3. 3.Text Generation: Conversion of structured data into natural language question-answer pairs
  4. 4.Quality Control: Expert curation and validation of chemical property representations
  5. 5.Deduplication: Removal of duplicate entries and data cleaning
  6. 6.Validation: Train/validation/test splits and quality checks

πŸ—οΈ ChemPile Collection

This dataset is part of the ChemPile collection, a comprehensive open dataset containing over 75 billion tokens of curated chemical data for training and evaluating general-purpose models in the chemical sciences.

Collection Overview

  • β€”πŸ“Š Scale: 75+ billion tokens across multiple modalities
  • β€”πŸ§¬ Modalities: Structured representations (SMILES, SELFIES, IUPAC, InChI), scientific text, executable code, and molecular images
  • β€”πŸŽ― Design: Integrates foundational educational knowledge with specialized scientific literature
  • β€”πŸ”¬ Curation: Extensive expert curation and validation
  • β€”πŸ“ˆ Benchmarking: Standardized train/validation/test splits for robust evaluation
  • β€”πŸŒ Availability: Openly released via Hugging Face

πŸ“„ Citation

If you use this dataset in your research, please cite:

bibtex
@article{mirza2025chempile0,
  title   = {ChemPile: A 250GB Diverse and Curated Dataset for Chemical Foundation Models},
  author  = {Adrian Mirza and Nawaf Alampara and MartiΓ±o RΓ­os-GarcΓ­a and others},
  year    = {2025},
  journal = {arXiv preprint arXiv:2505.12534}
}

πŸ‘₯ Contact & Support


<div align="center">

[image]

<i>Part of the ChemPile project - Advancing AI for Chemical Sciences</i>

</div>