datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cti-bench
Dataset Card for CTIBench
A set of benchmark tasks designed to evaluate large language models (LLMs) on cyber threat intelligence (CTI) tasks.
Dataset Details
Dataset Description
CTIBench is a comprehensive suite of benchmark tasks and datasets designed to evaluate LLMs in the field of CTI.
Components:
CTI-MCQ: A knowledge evaluation dataset with multiple-choice questions to assess the LLMs' understanding of CTI standards, threats, detection strategies… See the full description on the dataset page: https://huggingface.co/datasets/AI4Sec/cti-bench.Sombench-Ice-Prospectivity-Regression
SomBench Benchmark: Polar Ice Prospectivity Regression
Science theme: Polar volatiles
Task: Regression
Dataset Summary
A polar, multi-layer benchmark for predicting near-surface water-ice
prospectivity within ~10° latitude of each pole at 240 m/pixel. Following
the ice-prospectivity workflow of Coyan et al. (2025), the dataset includes a
group of physically motivated evidential layers (thermophysical,
illumination, and terrain) alongside a continuous prospectivity… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-Ice-Prospectivity-Regression.MoleculeQA
Dataset Card for MoleculeQA
Dataset Details
Dataset Description
MoleculeQA: A Dataset to Evaluate Factual Accuracy in Molecular Comprehension (EMNLP 2024)
Curated by: IDEA-XL
Language(s) (NLP): en
License: mit
Dataset Sources
Repository: https://github.com/IDEA-XL/MoleculeQA
Paper [optional]: https://arxiv.org/abs/2403.08192
Dataset Structure
- JSON
- All
- train.json # 49,993
- valid.json # 5,795
- test.json # 5… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-AI4S/MoleculeQA.Sombench-pretraining-data
SomBench Pre-training Corpus: Multimodal Lunar Tiles
Dataset Summary
This includes a small sample from SomBench: a corpus of co-registered, multimodal lunar image tiles built for
large-scale self-supervised (foundation-model) pre-training. It contains a subset of modalities from the
low-resolution (WAC-anchored) and high-resolution (NAC-anchored) tracks specifically used in pretraining.
Tiles are anchored to individual LROC Experiment Data Record (EDR) image… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-pretraining-data.Surya-bench-solarwind
Solar Wind Forecasting Dataset
Dataset Summary
This dataset provides hourly solar wind plasma and interplanetary magnetic field (IMF) parameters at L1, derived from NASA’s OMNI dataset. The primary forecasting target is the solar wind speed (V), while additional parameters are included for completeness:
Solar wind speed (V)
IMF Bx (GSE)
IMF By (GSM)
IMF Bz (GSM)
Proton number density (N)
The dataset is structured for machine learning experiments, particularly… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Surya-bench-solarwind.surya-bench-ar-segmentation
A Dataset of Binary Maps of Active Regions with Polarity Inversion Lines
Dataset Summary
This dataset provides hourly binary segmentation maps (4096×4096 resolution) derived from Solar Dynamics Observatory (SDO) / Helioseismic and Magnetic Imager (HMI) line-of-sight magnetograms. The maps highlight regions containing Active Regions (ARs) and Polarity Inversion Lines (PILs). The dataset spans observations from May 13, 2010 to December 31, 2024 and is intended for image… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/surya-bench-ar-segmentation.Sombench-WAC-Crater-Detection
SomBench Benchmark: Robbins Crater Detection, WAC
Science theme: Impact processes
Task: Object detection
Dataset Summary
An impact-crater object-detection benchmark built from the
Robbins (2019) global lunar crater
catalog, a manually compiled, near-complete census of
lunar impact craters (≥ ~1–2 km). Catalog crater centers and diameters are
converted to bounding boxes and packaged over LROC WAC visible tiles drawn
from the pre-training corpus test split, in COCO… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-WAC-Crater-Detection.Sombench-IMP-Segmentation
SomBench Benchmark: Irregular Mare Patch (IMP) Segmentation
Science theme: Volcanic history
Task: Binary semantic segmentation
Dataset Summary
A binary semantic-segmentation benchmark for irregular mare patches
(IMPs): rare, morphologically subtle features interpreted as unusually young
volcanic landforms. Each sample is an LROC NAC image tile paired with a
binary IMP mask (IMP vs. background). The set is derived from published IMP
polygon annotations, framed as a… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-IMP-Segmentation.ChemCoTDataset|- mol_edit/
|---- add.json
|---- delete.json
|---- sub.json
|- mol_opt/
|---- drd.json
|---- gsk.json
|---- jnk.json
|---- qed.json
|---- solubility.json
|---- logp.json
|- mol_und/
|---- fg_count.json
|---- Murcko_scaffold.json
|---- ring_count.json
|---- ring_system_scaffold.json
|---- equivalance.json
|- rxn
|---- fs_by_product.json
|---- fs_major_product.json
|---- retro.json
|---- rcr.json
|---- mech_sel.json
|---- nepp.json
📰 News
[2026.1.19] 🤝 Add… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-AI4S/ChemCoTDataset.ChemCoTBench
Tasks
mol_und
fg-level
fg_count.json
100 samples across 38 different functional groups detection
ring_count.json
20 samples, 9 types of ring unit
scaffold-level
Murcko_scaffold.json
40 samples, using MurckoScaffold extraction
ring_system_scaffold.json
60 samples, extract ring system as scaffolds
SMILES-level
equivalence.json
50 samples, each smiles -> mutate -> permutate, mutated smiles differs from original smiles
50 samples, each smiles -> permutate… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-AI4S/ChemCoTBench.euv-spectra
Solar EUV Spectra Modeling Dataset
Dataset Summary
This dataset provides time-aligned Extreme Ultraviolet (EUV) irradiance spectra from NASA’s SDO/EVE (Extreme Ultraviolet Variability Experiment) instrument. This dataset enables machine learning models to learn from and predict EUV spectral behavior driven by solar dynamics. It addresses the need for high-resolution, calibrated spectral data paired with physics-based contextual input. It is designed for image-to-spectra… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/euv-spectra.surya-bench-flare-forecasting
Full-disk Solar Flare Forecasting Dataset
Dataset Summary
This dataset provides labels for solar flare forecasting derived from NOAA GOES flare events from May 2010 to December 2024. Labels are constructed using a 24h rolling prediction window sampled at an hourly cadence. Each window is annotated with both max GOES class (based on peak X-ray flux) and cumulative flare index.
Two derived binary labels are included for forecasting tasks:
label_max: 1 if the maximum… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/surya-bench-flare-forecasting.RCR_RP_57K_SMILES-MMChatReaction Condition Prediction Dataset (Reagent Prediction)
molecule representation format: 1D SMILES
will further encode into 2D graph features
For detail, refer to PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes: https://arxiv.org/pdf/2406.13193
ar_emergence
Active Region Emergence Dataset
Dataset Summary
The Active Region Emergence Dataset is designed to support research on the early detection of solar Active Regions (ARs) and the development of predictive models for space weather. By characterizing the evolution of ARs before, during, and after their emergence, the dataset enables studies of pre-emergence signatures and early warning methods.
This dataset is derived from NASA’s Solar Dynamics Observatory (SDO) using… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/ar_emergence.surya-bench-coronal-extrapolation
Coronal Field Extrapolation Dataset
Dataset Summary
This dataset contains spherical harmonic coefficients of the coronal magnetic potential generated by emulating the physics-based ADAPT-WSA PFSS (Potential Field Source Surface) code, driven by SDO/HMI solar magnetogram observations. The target spherical harmonic coefficients represent the magnetic potential between the photosphere and the source surface (set to 2.51 Rs).Each file also contains additional variables from… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/surya-bench-coronal-extrapolation.ReactBench
ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams
Dataset Summary
ReactBench is a comprehensive benchmark designed to evaluate the topological reasoning capabilities of Multimodal Large Language Models (MLLMs) specifically on chemical reaction diagrams. The dataset challenges models to interpret complex visual layouts, understand molecular transformations, and trace reaction pathways across different diagram structures.
Data… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-AI4S/ReactBench.SMol_RS_Filtered_825K_SMILES-MMChatRetrosynthesis Prediction Dataset (derived from SMolInstruct)
molecule representation format: 1D SMILES
will further encode into 2D graph features
We filtered out overlapping samples from the original train-split (test-set: MolInstruct-Retrosynthesis Prediction)
We only include single-step retrosynthesis prediction.
For Detail, refer to PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes: https://arxiv.org/pdf/2406.13193
SMol_S2F_270K-MMChatMolInst_RS_125K_SMILES-MMChatUSPTO_1k_TPL-SFT
USPTO 1K TPL Reaction Classification Dataset
The USPTO 1K TPL Reaction Classification Dataset is a collection of chemical reactions labeled with their corresponding reaction classes. The dataset is derived from the USPTO (United States Patent and Trademark Office) database and consists of 1,000 different reaction templates (classes).
Dataset Structure
The dataset is organized into the following structure:
.
├── README.md
├── dataset_infos.json
├── instructions.txt
├──… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-AI4S/USPTO_1k_TPL-SFT.SMol_FS_Filtered_875K_SMILES-MMChatForward Reaction Prediction Dataset (derived from SMolInstruct)
molecule representation format: 1D SMILES
will further encode into 2D graph features
We filtered out overlapping samples from original train-split (test-set: MolInstruct-Forward Reaction Prediction)
For Detail, refer to PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes: https://arxiv.org/pdf/2406.13193
RCR_SP_70K_SMILES-MMChatReaction Condition Prediction Dataset (Solvent Prediction)
molecule representation format: 1D SMILES
will further encode into 2D graph features
For Detail, refer to PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes: https://arxiv.org/pdf/2406.13193
MolInst_RS_125K_Scaffold_SMILES-MMChatRetrosynthesis Prediction Dataset (derived from MolInstruct)
molecule representation format: 1D SMILES
will further encode into 2D graph features
We use scaffold splitting to reconstruct the train-split. We use SMolInstruct RS train split as the sample pool.
We only include single-step retrosynthesis prediction.
For Detail, refer to PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes: https://arxiv.org/pdf/2406.13193
RCR_CP_10K_SMILES-MMChatReaction Condition Prediction Dataset (Catalyst Prediction)
molecule representation format: 1D SMILES
will further encode into 2D graph features
Detail refer to PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes: https://arxiv.org/pdf/2406.13193
MolInst_FS_125K_SMILES-MMChatSMol_I2S_270K-MMChatSMol_I2F_270K-MMChatSMol_S2I_270K-MMChatBH-SM_YR_10K-MMChatHTE_RAS_4K-MMChat
