nvidia/Nemotron-RL-litmus-bench-v0.1
Dataset Description: Litmus-Bench v0.1 is an open dataset for training and evaluating chemical reasoning in language models. It includes 5,232 training questions and 482 test questions, each in short-answer format and was created from the ChEMBL dataset with RDKit descriptors requiring short answers. The dataset is for RL training. This dataset is released as part of NVIDIA NeMo-Gym, an open-source library within the NVIDIA NeMo framework, designed for large-scale, verifiable… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-litmus-bench-v0.1.
Dataset Description:
Litmus-Bench v0.1 is an open dataset for training and evaluating chemical reasoning in language models. It includes 5,232 training questions and 482 test questions, each in short-answer format and was created from the ChEMBL dataset with RDKit descriptors requiring short answers. The dataset is for RL training.
This dataset is released as part of NVIDIA NeMo-Gym), an open-source library within the NVIDIA NeMo framework, designed for large-scale, verifiable reinforcement-learning (RL) data used in post-training generativeAI models. It was utilized in the development of the NVIDIA Nemotron family of models.
NeMo Gym is a framework for building and serves as a unified hub for all RL data, environments, and reward signals tailored for post-training generativeAI models. Its purpose is to provide the simulation "worlds" in which AI agents can learn. Specifically, NeMo-Gym helps generate large-scale, high-quality data for a new training paradigm: Reinforcement Learning from Verifiable Reward (RLVR).
This dataset is part of the Hugging Face Org.
This dataset is ready for commercial or non-commercial uses.
Dataset Owner(s):
NVIDIA Corporation
Dataset Creation Date:
03/15/2026
Version:
v0.1
Previous Version(s): None
License/Terms of Use:
This dataset is licensed under CC BY 4.0.
Intended Usage:
To be used with NeMo-Gym for post-training LLMs.
Dataset Characterization
Data Collection Method
- \[Human\]
Labeling Method
- \[Human\]
Dataset Format
Text Only, Compatible with NeMo-Gym
Dataset Quantification
Record Count: 5232 (train), 482 (test) Feature Count: 123 different properties across dataset, each record contains a SMILES and a single property value as the correct answer Total Data Storage: 5.58 MB
Reference(s):
Ethical Considerations:
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. Developers should work with their internal developer teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns here.
