datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RealPDEBench
RealPDEBench
RealPDEBench is a benchmark of paired real-world measurements and matched numerical simulations for complex physical systems. It is designed for spatiotemporal forecasting and sim-to-real transfer evaluation on real data.
This Hub repository (AI4Science-WestlakeU/RealPDEBench) is the release repo for RealPDEBench.
Website & documentation: realpdebench.github.io
Raw HDF5 distribution: realpdebench.westlake.edu.cn
Benchmark codebase:… See the full description on the dataset page: https://huggingface.co/datasets/AI4Science-WestlakeU/RealPDEBench.ai4sci-virtual-cell-zeroshot-data
ai4sci virtual-cell-zeroshot — release v1
Prepared data for the virtual-cell-zeroshot
task: predict single-cell CRISPRi knockdown responses in cellular contexts a model has never seen
perturbed (the Arc Virtual Cell Challenge 2026 zero-shot setting), rebuilt from public data.
dev/ what the agent sees (mount read-only at /workspace/data)
train/{k562,jurkat,hct116,hek293t}/ cells.h5ad, pseudobulk.h5ad, se_embeddings.npy, dev_split/
test/{ctx_near,ctx_mid… See the full description on the dataset page: https://huggingface.co/datasets/sunweiwei/ai4sci-virtual-cell-zeroshot-data.SoccerWiki
Dataset Card for SoccerWiki
This repository contains the database for paper "Multi-Agent System for Comprehensive Soccer Understanding" in ACM Multimidia 2025.
SoccerWiki is a large-scale multimodal soccer knowledge base. The dataset integrates rich domain knowledge about soccer players, teams, referees, and venues, which is used to facilitate knowledge-driven reasoning and decision-making in various soccer-related tasks.
This dataset was built using data from Wikipedia and… See the full description on the dataset page: https://huggingface.co/datasets/SJTU-AI4Sports/SoccerWiki.cti-bench
Dataset Card for CTIBench
A set of benchmark tasks designed to evaluate large language models (LLMs) on cyber threat intelligence (CTI) tasks.
Dataset Details
Dataset Description
CTIBench is a comprehensive suite of benchmark tasks and datasets designed to evaluate LLMs in the field of CTI.
Components:
CTI-MCQ: A knowledge evaluation dataset with multiple-choice questions to assess the LLMs' understanding of CTI standards, threats, detection strategies… See the full description on the dataset page: https://huggingface.co/datasets/AI4Sec/cti-bench.core-sdo
ML-Ready Multi-Modal Image Dataset from SDO
Overview
This dataset provides machine learning (ML)-ready solar data curated from NASA’s Solar Dynamics Observatory (SDO), covering observations from May 13, 2010, to Dec 31, 2024. It includes Level-1.5 processed data from: Atmospheric Imaging Assembly (AIA)
and Helioseismic and Magnetic Imager (HMI).
The dataset is designed to facilitate large-scale learning applications in heliophysics, such as space weather forecasting… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/core-sdo.Surya-1.0_validation_data
Validation data for Surya 1.0
This dataset comprises imagery from NASA's Solar Dynamics Observatory (SDO). The data can and should be used to validate a local installation of the Surya Foundation Model for Heliophysics. The data is compressed; you should use the hdf5plugin to read it directly.
Sombench-Ice-Prospectivity-Regression
SomBench Benchmark: Polar Ice Prospectivity Regression
Science theme: Polar volatiles
Task: Regression
Dataset Summary
A polar, multi-layer benchmark for predicting near-surface water-ice
prospectivity within ~10° latitude of each pole at 240 m/pixel. Following
the ice-prospectivity workflow of Coyan et al. (2025), the dataset includes a
group of physically motivated evidential layers (thermophysical,
illumination, and terrain) alongside a continuous prospectivity… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-Ice-Prospectivity-Regression.RealPDE-Competition-Data
RealPDE Competition Data (NeurIPS 2026)
Training data and baseline checkpoints for the NeurIPS 2026 RealPDE
Competition. This is a mirror of the
competition's Google Drive release, hosted here because the Drive link runs into
a per-file anonymous download quota when many people fetch it at once.
Both tracks share this release:
Track 1, Sim2Real — codabench.org/competitions/17363
Track 2, LTTTA — codabench.org/competitions/17385
Contents
train_sim.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/AI4Science-WestlakeU/RealPDE-Competition-Data.MoleculeQA
Dataset Card for MoleculeQA
Dataset Details
Dataset Description
MoleculeQA: A Dataset to Evaluate Factual Accuracy in Molecular Comprehension (EMNLP 2024)
Curated by: IDEA-XL
Language(s) (NLP): en
License: mit
Dataset Sources
Repository: https://github.com/IDEA-XL/MoleculeQA
Paper [optional]: https://arxiv.org/abs/2403.08192
Dataset Structure
- JSON
- All
- train.json # 49,993
- valid.json # 5,795
- test.json # 5… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-AI4S/MoleculeQA.Sombench-pretraining-data
SomBench Pre-training Corpus: Multimodal Lunar Tiles
Dataset Summary
This includes a small sample from SomBench: a corpus of co-registered, multimodal lunar image tiles built for
large-scale self-supervised (foundation-model) pre-training. It contains a subset of modalities from the
low-resolution (WAC-anchored) and high-resolution (NAC-anchored) tracks specifically used in pretraining.
Tiles are anchored to individual LROC Experiment Data Record (EDR) image… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-pretraining-data.Surya-bench-solarwind
Solar Wind Forecasting Dataset
Dataset Summary
This dataset provides hourly solar wind plasma and interplanetary magnetic field (IMF) parameters at L1, derived from NASA’s OMNI dataset. The primary forecasting target is the solar wind speed (V), while additional parameters are included for completeness:
Solar wind speed (V)
IMF Bx (GSE)
IMF By (GSM)
IMF Bz (GSM)
Proton number density (N)
The dataset is structured for machine learning experiments, particularly… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Surya-bench-solarwind.Sombench-NAC-Crater-Detection
SOMBench Benchmark: Hand-Labeled Crater Detection in NAC data (nac_craters_dataset)
Science theme: Impact cratering
Task: Single-class object detection (COCO bounding boxes)
Dataset Summary
A crater-detection benchmark built from expert hand-labeled crater outlines
on LROC NAC orthophotos across six lunar NAC PHO sites (Apollo 15 SIVB impact,
Apollo 17, Highland Photom, King Ejecta, March 17 Impact, Reiner Gamma). Subject matter expert (SME) annotators… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-NAC-Crater-Detection.surya-bench-ar-segmentation
A Dataset of Binary Maps of Active Regions with Polarity Inversion Lines
Dataset Summary
This dataset provides hourly binary segmentation maps (4096×4096 resolution) derived from Solar Dynamics Observatory (SDO) / Helioseismic and Magnetic Imager (HMI) line-of-sight magnetograms. The maps highlight regions containing Active Regions (ARs) and Polarity Inversion Lines (PILs). The dataset spans observations from May 13, 2010 to December 31, 2024 and is intended for image… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/surya-bench-ar-segmentation.Sombench-WAC-Crater-Detection
SomBench Benchmark: Robbins Crater Detection, WAC
Science theme: Impact processes
Task: Object detection
Dataset Summary
An impact-crater object-detection benchmark built from the
Robbins (2019) global lunar crater
catalog, a manually compiled, near-complete census of
lunar impact craters (≥ ~1–2 km). Catalog crater centers and diameters are
converted to bounding boxes and packaged over LROC WAC visible tiles drawn
from the pre-training corpus test split, in COCO… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-WAC-Crater-Detection.SMolInst-Reactionsai4sci-hydrogym-research-v2
HydroGym Research v2: actual CFD resources
This repository hosts data, not the Harbor task or training/evaluation code.
Harbor instructions, starter/reference implementations, model-delivery contract
and evaluator belong in T0-RSI/ai4sci-tasks.
Publication status
The actual resource upload is complete: all 558 selected files and their exact
453,756,315,982-byte total are present. Use the immutable Hub revision pinned
in the GitHub task's… See the full description on the dataset page: https://huggingface.co/datasets/sunweiwei/ai4sci-hydrogym-research-v2.Sombench-IMP-Segmentation
SomBench Benchmark: Irregular Mare Patch (IMP) Segmentation
Science theme: Volcanic history
Task: Binary semantic segmentation
Dataset Summary
A binary semantic-segmentation benchmark for irregular mare patches
(IMPs): rare, morphologically subtle features interpreted as unusually young
volcanic landforms. Each sample is an LROC NAC image tile paired with a
binary IMP mask (IMP vs. background). The set is derived from published IMP
polygon annotations, framed as a… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/Sombench-IMP-Segmentation.ChemO
🧪 ChemO Dataset
📄 Paper: ChemLabs on ChemO: A Multi-Agent System for Multimodal Reasoning on IChO 2025
ChemO Version 1.1
Now with CDXML Files! 🎉
The ChemO dataset has been officially released after meticulous proofreading and preparation. This benchmark is built from the International Chemistry Olympiad (IChO) 2025 and represents a new frontier in automated chemical problem-solving.
🌟 Key Features
🏆 Olympic-Level Benchmark - Challenging problems… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-AI4S/ChemO.ChemCoTDataset|- mol_edit/
|---- add.json
|---- delete.json
|---- sub.json
|- mol_opt/
|---- drd.json
|---- gsk.json
|---- jnk.json
|---- qed.json
|---- solubility.json
|---- logp.json
|- mol_und/
|---- fg_count.json
|---- Murcko_scaffold.json
|---- ring_count.json
|---- ring_system_scaffold.json
|---- equivalance.json
|- rxn
|---- fs_by_product.json
|---- fs_major_product.json
|---- retro.json
|---- rcr.json
|---- mech_sel.json
|---- nepp.json
📰 News
[2026.1.19] 🤝 Add… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-AI4S/ChemCoTDataset.ChemCoTBench
Tasks
mol_und
fg-level
fg_count.json
100 samples across 38 different functional groups detection
ring_count.json
20 samples, 9 types of ring unit
scaffold-level
Murcko_scaffold.json
40 samples, using MurckoScaffold extraction
ring_system_scaffold.json
60 samples, extract ring system as scaffolds
SMILES-level
equivalence.json
50 samples, each smiles -> mutate -> permutate, mutated smiles differs from original smiles
50 samples, each smiles -> permutate… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-AI4S/ChemCoTBench.euv-spectra
Solar EUV Spectra Modeling Dataset
Dataset Summary
This dataset provides time-aligned Extreme Ultraviolet (EUV) irradiance spectra from NASA’s SDO/EVE (Extreme Ultraviolet Variability Experiment) instrument. This dataset enables machine learning models to learn from and predict EUV spectral behavior driven by solar dynamics. It addresses the need for high-resolution, calibrated spectral data paired with physics-based contextual input. It is designed for image-to-spectra… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/euv-spectra.surya-bench-flare-forecasting
Full-disk Solar Flare Forecasting Dataset
Dataset Summary
This dataset provides labels for solar flare forecasting derived from NOAA GOES flare events from May 2010 to December 2024. Labels are constructed using a 24h rolling prediction window sampled at an hourly cadence. Each window is annotated with both max GOES class (based on peak X-ray flux) and cumulative flare index.
Two derived binary labels are included for forecasting tasks:
label_max: 1 if the maximum… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/surya-bench-flare-forecasting.RCR_RP_57K_SMILES-MMChatReaction Condition Prediction Dataset (Reagent Prediction)
molecule representation format: 1D SMILES
will further encode into 2D graph features
For detail, refer to PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes: https://arxiv.org/pdf/2406.13193
ar_emergence
Active Region Emergence Dataset
Dataset Summary
The Active Region Emergence Dataset is designed to support research on the early detection of solar Active Regions (ARs) and the development of predictive models for space weather. By characterizing the evolution of ARs before, during, and after their emergence, the dataset enables studies of pre-emergence signatures and early warning methods.
This dataset is derived from NASA’s Solar Dynamics Observatory (SDO) using… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/ar_emergence.PubChemSFT
all_clean.json
remove overlapped parts with ChEBI-20 test
remove no description SMILES
Format:{
SMILES <str>:
[
["Please describe the molecule", DESCRIPTION],
...,
]
}
Stats
max tokens length: 6113
min tokens length: 20
mean tokens length: 191
median tokens length: 149
Total 326,689 single turn dialogue. Total 293,302 SMILES examples.
Size: Train: 264,391 Valid: 33,072 Test: 32,987
conversation template
'conversation':{
[
"from": "human",
"value": <QUERY>, #… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-AI4S/PubChemSFT.surya-bench-coronal-extrapolation
Coronal Field Extrapolation Dataset
Dataset Summary
This dataset contains spherical harmonic coefficients of the coronal magnetic potential generated by emulating the physics-based ADAPT-WSA PFSS (Potential Field Source Surface) code, driven by SDO/HMI solar magnetogram observations. The target spherical harmonic coefficients represent the magnetic potential between the photosphere and the source surface (set to 2.51 Rs).Each file also contains additional variables from… See the full description on the dataset page: https://huggingface.co/datasets/nasa-ibm-ai4science/surya-bench-coronal-extrapolation.ReactBench
ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams
Dataset Summary
ReactBench is a comprehensive benchmark designed to evaluate the topological reasoning capabilities of Multimodal Large Language Models (MLLMs) specifically on chemical reaction diagrams. The dataset challenges models to interpret complex visual layouts, understand molecular transformations, and trace reaction pathways across different diagram structures.
Data… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-AI4S/ReactBench.ai4sci-probabilistic-global-weather
AI4Science probabilistic global weather: prepared WN2 resources
Frozen, ready-to-use data for the native WeatherNext 2 track in
T0-RSI/ai4sci-tasks.
This release contains approximately 205.2 GiB of prepared NetCDF/Zarr resources.
It avoids rebuilding training examples, normalizers and evaluation climatologies
from large, rolling upstream archives. It contains no pretrained model weights,
candidate predictions or experiment logs.
Contents and intended use
32 fixed… See the full description on the dataset page: https://huggingface.co/datasets/sanxing/ai4sci-probabilistic-global-weather.SMol_RS_Filtered_825K_SMILES-MMChatRetrosynthesis Prediction Dataset (derived from SMolInstruct)
molecule representation format: 1D SMILES
will further encode into 2D graph features
We filtered out overlapping samples from the original train-split (test-set: MolInstruct-Retrosynthesis Prediction)
We only include single-step retrosynthesis prediction.
For Detail, refer to PRESTO: Progressive Pretraining Enhances Synthetic Chemistry Outcomes: https://arxiv.org/pdf/2406.13193
SMol_S2F_270K-MMChat
