datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-SpecializedDomains-Finance-v1
Dataset Description
Nemotron-SpecializedDomains-Finance is a large-scale synthetic financial question-answering dataset designed to improve LLM performance on specialized financial reasoning and document comprehension tasks. The dataset comprises 326K+ high-quality Q&A pairs generated from SEC filings of S&P 500 companies spanning 2019-2024.
This dataset is ready for commercial use.
Overview
The dataset leverages template-based Synthetic Data Generation (SDG) to… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SpecializedDomains-Finance-v1.Nemotron-RL-ARC-AGI-v1
Dataset Description:
Nemotron-RL-ARC-AGI-v1 is a reinforcement-learning (RL) gym environment dataset of single-step ARC-AGI puzzle prompts intended for RL post-training of large language models. Each row corresponds to one ARC puzzle (a set of (input grid, output grid) demonstration pairs plus a single test input grid) rendered as a text prompt; reward is binary (1.0 / 0.0) determined by exact-match comparison against the ground-truth output grid. No LLM judge is used, no… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-ARC-AGI-v1.Nemotron-RL-litmus-bench-v0.1
Dataset Description:
Litmus-Bench v0.1 is an open dataset for training and evaluating chemical reasoning in language models. It includes 5,232 training questions and 482 test questions, each in short-answer format and was created from the ChEMBL dataset with RDKit descriptors requiring short answers. The dataset is for RL training.
This dataset is released as part of NVIDIA NeMo-Gym, an open-source library within the NVIDIA NeMo framework, designed for large-scale, verifiable… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-litmus-bench-v0.1.Nemotron-RL-QA-Abstention-v1Nemotron-RL-QA-Abstention-v1
License: cc-by-4.0
Language: en
Task Categories: reinforcement-learning, question-answering, text-generation
Tags: abstention, question-answering, hotpotqa, software-engineering, health, law, rl, rlvr
Configs: default train split at data/train.jsonl
Domain: multi-domain question answering, abstention
Modality: text
Capability Breakdown: Abstention-aware factoid question answering [100%]
Source: Hybrid: Automated, Manually Collected, Synthetic
Size Bin: <10K… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-QA-Abstention-v1.ruler-100-nemotron
RULER-100 — Nemotron-Nano-v3 tokenized
RULER long-context evaluation data, regenerated with the
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 (instruct) tokenizer so the labeled context
lengths are exact for that model — instead of drifting, as they do when RULER data tokenized for a
different model (e.g. Qwen3) is fed to Nemotron.
What's here
7 context lengths: 4096, 8192, 16384, 32768, 65536, 131072, 262144 (the model's max).
13 RULER tasks: niah_single_1/2/3… See the full description on the dataset page: https://huggingface.co/datasets/jet-ai/ruler-100-nemotron.persona_nemotron
Persona Nemotron PT Datasets
This is a collection of Portuguese synthetic datasets, consisting of 3 datasets, one with general questions from varied topics, one with math questions, and one with instruction-following requests.
The prompts were generated using an approach similar to PersonaHub, with a translated version of Nemotron Personas. Both prompts and answers were generated using Gemma 3-27B.
This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/persona_nemotron.nvidia-Nemotron-Science-Math
NVIDIA Nemotron Science and Math Reasoning
This is an unofficial, curated collection derived from NVIDIA's open-source Nemotron datasets. It is designed specifically to train language models in complex scientific and mathematical reasoning by providing structured chain-of-thought (CoT) examples.
To ensure efficiency, the shortest available CoT sequence was chosen for each question, filtering out redundant variations while preserving the core logical progression toward the final… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/nvidia-Nemotron-Science-Math.Nemotron-RL-litmus-bench-v0.1
Dataset Description:
Litmus-Bench v0.1 is an open dataset for training and evaluating chemical reasoning in language models. It includes 5,232 training questions and 482 test questions, each in short-answer format and was created from the ChEMBL dataset with RDKit descriptors requiring short answers. The dataset is for RL training.
This dataset is released as part of NVIDIA NeMo-Gym, an open-source library within the NVIDIA NeMo framework, designed for large-scale, verifiable… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-RL-litmus-bench-v0.1.STXBP1-RAG-Nemotron
🧬⚡ STXBP1-ARIA RAG Database v10.1 - NVIDIA Nemotron Embeddings
The most advanced RAG database for STXBP1 therapeutic research.
A pre-built ChromaDB vector database containing:
571,816 indexed text chunks from ~17,000 curated PubMed Central (PMC) biomedical papers + 165 base editing analysis entries,
(https://huggingface.co/datasets/SkyWhal3/stxbp1-base-editing-sweep),
embedded with NVIDIA's state-of-the-art Llama-Nemotron-Embed-1B-v2 model featuring 2048-dimensional embeddings.
⚡… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/STXBP1-RAG-Nemotron.SNAP25-RAG-Nemotron
🧬⚡ SNAP25-ARIA RAG Database v2.0 - NVIDIA Nemotron Embeddings
The first comprehensive RAG database for SNAP25 therapeutic research.
A pre-built ChromaDB vector database containing:
76,592 indexed text chunks from 2,043 curated PubMed Central (PMC) biomedical papers + 683 OpenFold3 structural analysis reports + 53 expert-curated knowledge entries,
embedded with NVIDIA's state-of-the-art Llama-Nemotron-Embed-1B-v2 model featuring 2048-dimensional embeddings.
🤝 Built for the SNAP25… See the full description on the dataset page: https://huggingface.co/datasets/SkyWhal3/SNAP25-RAG-Nemotron.amalia-Nemotron-SpecializedDomains-Finance-v1
AMALIA Nemotron-SpecializedDomains-Finance-v1
Version of the nvidia/Nemotron-SpecializedDomains-Finance-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage.
This dataset went through a processing pipeline to:
Remove entries that reference other LLMs or research labs;
Remove the reasoning_content field;
Original Dataset: https://huggingface.co/datasets/nvidia/Nemotron-SpecializedDomains-Finance-v1
This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SpecializedDomains-Finance-v1.Nemotron-RL-litmus-bench-v0.1
Dataset Description:
Litmus-Bench v0.1 is an open dataset for training and evaluating chemical reasoning in language models. It includes 5,232 training questions and 482 test questions, each in short-answer format and was created from the ChEMBL dataset with RDKit descriptors requiring short answers. The dataset is for RL training.
This dataset is released as part of NVIDIA NeMo-Gym, an open-source library within the NVIDIA NeMo framework, designed for large-scale, verifiable… See the full description on the dataset page: https://huggingface.co/datasets/Iambackup/Nemotron-RL-litmus-bench-v0.1.nemotron-nano-rl-mcqa-19k
Nemotron Nano RL MCQA 19K
Nemotron Nano RL MCQA 19K is a 19,670-example English multiple-choice question answering dataset prepared for reinforcement learning with verifiable rewards (RLVR). Its nano_v3_sft_profiled_stem_mcqa identifier and schema correspond to the knowledge-MCQA component of NVIDIA's Nemotron-3-Nano-RL-Training-Blend, represented here in a compact prompt / label / metadata JSONL format.
Each record contains a formatted user prompt, the correct option identifier… See the full description on the dataset page: https://huggingface.co/datasets/wflying/nemotron-nano-rl-mcqa-19k.aya-ja-nemotron-dpo-masked
aya-ja-nemotron-dpo-masked
LLMの推論能力を向上させるためのデータセット
CohereForAI/aya_datasetから日本語パートを抜粋
deepinfraのnvidia/Nemotron-4-340B-Instructで応答を再生成
2024年8月現在ではnvidia/Nemotron-4-340B-Instructは使用不可
5,651件(6,259件の内608件削除)
ルールベース・機械学習ベースのフィルタリング処理の後、目視確認を行い、個人情報が含まれ得るサンプルについては全件削除対応を実施
Format
{
"idx": インデックス,
"prompt": 日本語の指示文,
"chosen": chosenの応答文,
"rejected": rejectedの応答文,
"chosen_model": chosenとしたモデル,
"rejected_model":… See the full description on the dataset page: https://huggingface.co/datasets/weblab-GENIAC/aya-ja-nemotron-dpo-masked.
