datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opendatalab-experimental-nmr-peaks
OpenDataLab Experimental NMR Peaks Dataset
Dataset Description
This dataset contains experimental NMR (Nuclear Magnetic Resonance) peak sequences extracted from the OpenDataLab experimental spectra database. The dataset includes both H-NMR and C-NMR peak sequences for chemical compounds, along with their SMILES representations and molecular formulas.
Dataset Summary
Total Samples: 533,595 compounds
Batches: 333 batch files
Data Source: Experimental spectra… See the full description on the dataset page: https://huggingface.co/datasets/snehasis19/opendatalab-experimental-nmr-peaks.NMR-analysisnmr-canonical-cleaned
Canonical NMR Dataset Collection — Data Card
Dataset release: v4
Canonical schema: v2
Spectral modalities: 1H and 13C resonance-level peak lists
Collection overview
This release brings several of the largest openly available processed NMR
corpora used by current deep-learning methods into one model-independent
schema. It combines simulated and literature-derived spectra while preserving
the provenance and annotation coverage of every source.
The collection has… See the full description on the dataset page: https://huggingface.co/datasets/niccogreek/nmr-canonical-cleaned.nmrexp-cnmr-400knmr-baseline-datasetnmrexp-cnmr-peaklist-1.5Mnmrshiftdb2NMRBank
1. NMRBank data (225809)
We batch processed 380,220 NMR segments using NMRExtractor. After removing entries with empty 13C NMR chemical shifts, we obtained about 260,000 entries. Further filtering out entries with empty IUPAC names and NMR chemical shifts resulted in 225,809 entries.
NMRBank_data_225809.zip
2. NMRBank data with SMILES (156621)
To normalize these data, we converted IUPAC names to SMILES using ChemDraw and OPSIN, successfully converting 156,621 entries.… See the full description on the dataset page: https://huggingface.co/datasets/sweetssweets/NMRBank.nmrexp-hnmr-raw-500knmr-belief-cascade
BeliefCascade Branch Grid
Each row is one complete sequential belief-revision episode. The benchmark
uses a 432-condition grid: nodes per level {2, 3, 4, 5}, level counts
{3, 4, 5}, out-/in-degree complexity bands {20, 50, 80}, and revision
types {monotonic, nmr_retraction, nmr_newinfo, nmr_mixed}. There are 10
train and 50 test episodes for every condition (4,320 train / 21,600 test).
Columns
text: atoms, static dependencies, and inference policy.
belief:… See the full description on the dataset page: https://huggingface.co/datasets/leo-bjpark/nmr-belief-cascade.NMRGym
NMRGym
Benchmark on NMR Spectrum.
!!! Important: !!!
If you need access to the dataset, please email zhengf723@connect.hkust-gz.edu.cn
Data Sources (Before Cleaning)
Source
Records
Unique SMILES
Total Spectrums
¹H NMR
¹³C NMR
CH-NP
12,165
12,165
24,326
12,161
12,165
HMDB
1,791
896
3,278
1,566
1,712
NMRBank
148,914
142,964
297,043
148,437
148,606
NMRShiftDB 2024
41,019
39,631
50,416
18,570
31,846
NP-MRD
489,569
243,598
950,242
462,233
488,009… See the full description on the dataset page: https://huggingface.co/datasets/meaw0415/NMRGym.nmr-bench
NMR-Bench canonical183
NMR-Bench is a naturalistic long-context multi-hop reasoning benchmark over full real documents. This anonymous release is the canonical183 construction snapshot prepared for NeurIPS 2026 Evaluations & Datasets review. It contains 183 rubric-scored questions over 100 source documents, with evidence clues distributed across long contexts and tasks grouped into seven reasoning paradigms.
This is a v0.2.1 canonical dataset snapshot. It intentionally excludes… See the full description on the dataset page: https://huggingface.co/datasets/nmrbench/nmr-bench.nmrexp-cnmr-200knmr-dataset-synthetic-retraction
NMR Dataset: Synthetic Retraction
This dataset contains 360 complete belief-revision episodes: 90 each for
monotonic, nmr_new_evidence, nmr_retraction, and nmr_mixed.
Each JSONL row is one episode, with a fixed dependency graph, an initial
belief base, three revisions, and complete gold T/F/U states at all four
checkpoints. It is intentionally not expanded into one row per proposition.
An evaluator can construct a query for every proposition at any checkpoint
from the episode… See the full description on the dataset page: https://huggingface.co/datasets/leo-bjpark/nmr-dataset-synthetic-retraction.nmrexp-cnmrnmrexpsecs-chemotion-experimental-1H_NMR
1H NMR data from the Chemotion repository
The data has two columns: SMILES and 1H NMR. The SMILES are canonicalized using RDKit. 1H NMR is represented as a dictionary containing x and y coordinates.
x ranges from -2 to 10 (chemical shift).
y ranges from 0 to 1 (intensity).
The dataset can simply be loaded by:
from datasets import load_dataset
chemotion_data = load_dataset("jablonkagroup/secs-chemotion-experimental-1H_NMR")
31p-nmr-counterNMRTrans-Data
NMRTrans-Data
This dataset repository contains the cached NMRTrans splits:
train.pkl.lz4
val.pkl.lz4
test.pkl.lz4
These files are consumed by the NMRTrans MergedDataset loader.
secs-cnmr-nmrshiftdbnmr-retraction-hub
NMR Retraction Hub
Synthetic forward-DAG belief-revision episodes for evaluating policy-bounded
hub retraction. Each row is one complete policy variant of a shared base case.
The train.jsonl and test.jsonl files contain the evaluation splits;
variants.jsonl is the complete prototype collection and
policy_comparison.json is a compact review table.
Rules have one premise. Internal hub nodes have two incoming and two outgoing
rules, while internal simple nodes have one incoming and… See the full description on the dataset page: https://huggingface.co/datasets/leo-bjpark/nmr-retraction-hub.NMR-CRAFT-results
NMR-CRAFT Results
This dataset contains full workflow outputs from NMR-CRAFT on the NMRGym and
NMRSpec test sets. Each result file is JSONL with one record per test sample;
the corresponding metrics file contains aggregate evaluation results.
Files
nmrgym/
results.jsonl
metrics.json
nmrspec/
results.jsonl
metrics.json
NMRGym
The NMRGym test set contains 27,054 samples. Retrieval uses the
exclude_scaffold setting.
Metric
Top-1
Top-5… See the full description on the dataset page: https://huggingface.co/datasets/Anonymouszzzz/NMR-CRAFT-results.nmrexp-cnmr-peaklist-1.2Mnmr-description-flipnmr-flipnmrexp-cnmr-peaklist-400kbmrb-hsqc-nmr-1H13CThis is a set of Heteronuclear single quantum coherence spectroscopy (HSQC) nuclear magnetic resonance (NMR) spectra for various biological molecules. Specifically, these are 1H-13C spectra. The "train" fold originates from the Biological Magnetic Resonance Data Bank (https://bmrb.io/), while the "test" fold corresponds to measurements of fish liver. HSQC NMR spectra are 2D spectra with an intensity recorded for each point in 2D space. These results were collected using a Bruker instrument and the raw data can be read with python using the nmrglue (https://pypi.org/project/nmrglue/) package. This dataset is meant to accompany FINCHnmr (https://github.com/mahynski/FINCHnmr) as an example library.chemotion-experimental-1H_NMRafrica-unsdg-neonatal-mortality-rate-deaths-per-1-000-live-births-sh-dyn-nmrt
Africa Unsdg Neonatal Mortality Rate Deaths Per 1-000 Live Births Sh Dyn Nmrt | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-unsdg-neonatal-mortality-rate-deaths-per-1-000-live-births-sh-dyn-nmrt.human_dopamine_D1_receptor_antagonists_SMILES_13C_NMR_spectrosopy_conciseThe human_dopamine_D1_receptor_antagonists_SMILES_13C_NMR_spectrosopy_concised dataset is a part of the study "Leveraging 13C NMR spectroscopic data derived from SMILES to predict the functionality of small biomolecules by machine learning: a case study on human Dopamine D1 receptor antagonists "
https://doi.org/10.48550/arXiv.2501.14044
Dataset content: The Dataset has 59,567 rows of samples and 224 columns of features. Each feature corresponds to a range of the 13C NMR spectroscopy chemical… See the full description on the dataset page: https://huggingface.co/datasets/ivanovaml/human_dopamine_D1_receptor_antagonists_SMILES_13C_NMR_spectrosopy_concise.
