datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineinstructions_nemotron
✨ Note: For all FineInstructions resources please visit: https://huggingface.co/fineinstructions
This dataset is ~1B+ synthetic instruction-answer pairs or ~300B tokens created using the FineInstructions pipeline.
The FineInstructions pipeline was run over the raw pre-training documents in the Nemotron-CC pre-training corpus (a subset of high-quality documents from CommonCrawl). See our paper for more details.
Each .parquet file in the data folder has a corresponding judge-*.json file that… See the full description on the dataset page: https://huggingface.co/datasets/fineinstructions/fineinstructions_nemotron.Nemotron-Research-Reasoning-Qwen-1.5B_eval_569anemotron-mc-en-ar-midtrain
nemotron-mc-en-ar-midtrain
Arabic translation of the Nemotron-Pretraining-Multiple-Choice config of Nemotron-Pretraining-Specialized-v1.2 (pinned revision 807afc1). Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 23,926,492 source rows are present, none dropped. English source and Arabic translation sit in the same row, so the dataset serves as a parallel corpus as well as an Arabic one. A sibling corpus from the same pipeline is available at… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-mc-en-ar-midtrain.AceReason-Nemotron-7B_eval_118b
mlfoundations-dev/AceReason-Nemotron-7B_eval_118b
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBenchv5_official
Average Accuracy: 43.85% ± 0.24%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
43.37%
121
279
2
44.09%
123
279
3
44.09%
123
279
nemotron-r1-en-ar-midtrain
nemotron-r1-en-ar-midtrain
Arabic translation of the Llama_Nemotron_Post_Training_Dataset_reasoning_r1 split of smoltalk2 (config Mid, pinned revision fc6cc21): reasoning traces with <think> blocks in a conversational format. Translated with RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic (greedy) on H100s. FP8 was verified lossless against its bf16 parent before the run (chrF 96.4, 0 of 510 chunks materially diverged). All 3,644,790 source rows are present, none dropped. Sibling… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-r1-en-ar-midtrain.OpenReasoning-Nemotron-7B_eval_8179
mlfoundations-dev/OpenReasoning-Nemotron-7B_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
79.0
98.8
89.0
81.7
60.1
62.5
50.6
46.8
68.7
13.3
49.6
59.7
AIME24
Average Accuracy: 79.00% ± 1.42%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
70.00%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/OpenReasoning-Nemotron-7B_eval_8179.nemotron_cc_v2_hq_packed4096_200shard
Nemotron-CC-v2 High-Quality, packed to 4096 tokens (train)
Documents from nvidia/Nemotron-CC-v2
High-Quality subset, tokenized with google/t5gemma-2-270m-270m (add_special_tokens=False, EOS
appended per document) and greedily packed into sequences of at most 4096 tokens.
A document is never split across a pack boundary; documents longer than 4096
are truncated to their own pack. Every pack ends on an EOS/document boundary.
Schema
index (int64): running pack id… See the full description on the dataset page: https://huggingface.co/datasets/lillian039/nemotron_cc_v2_hq_packed4096_200shard.OpenReasoning-Nemotron-1.5B_eval_8179
mlfoundations-dev/OpenReasoning-Nemotron-1.5B_eval_8179
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
HMMT
Accuracy
49.7
83.0
78.0
49.4
31.0
35.5
19.8
14.6
40.7
12.0
24.3
32.3
AIME24
Average Accuracy: 49.67% ± 1.20%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
50.00%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/OpenReasoning-Nemotron-1.5B_eval_8179.nemotron-sft-balanced-2b-v1
Nemotron SFT Dataset
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Statistics
Total Samples: 200,000
Total Tokens: 1,252,287,904
Average Tokens per Sample: 6261.4
Tokenizer: Qwen/Qwen3-0.6B
Random Seed: 42
Strategy: balanced
Subset Distribution
Subset
Samples
Tokens
Target
Completion
Avg Tokens/Sample
Stage-1/math
20,000
151,546,125
20,000… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-balanced-2b-v1.AceReason-Nemotron-7B_eval_c64a
mlfoundations-dev/AceReason-Nemotron-7B_eval_c64a
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBenchv5_v3
Average Accuracy: 41.67% ± 0.45%
Number of Runs: 3
Run
Accuracy
Questions Solved
Total Questions
1
41.04%
110
268
2
41.42%
111
268
3
42.54%
114
268
generations-nemotron-nano-9b-v2-simnpo-gentle-igm-10bnemotron-sft-balanced-stage1-2
Nemotron SFT Dataset
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Statistics
Total Samples: 100,000
Total Tokens: 624,101,275
Average Tokens per Sample: 6241.0
Tokenizer: Qwen/Qwen3-0.6B
Random Seed: 42
Strategy: balanced
Subset Distribution
Subset
Samples
Tokens
Target
Completion
Avg Tokens/Sample
Stage-1/math
10,000
75,402,505
10,000… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-balanced-stage1-2.generations-nemotron-nano-9b-v2-simnpo-gentle-baselineOrnith-1.5-35B-A3B-Nemotron-v2-100M
Ornith 1.5 35B A3B Nemotron v2 100M
This dataset contains 108,729 English conversations with
108,729 regenerated assistant turns and 100,014,884 generated
assistant completion tokens. 100M refers to the completion-token target, not the
number of examples.
The prompt mix is a deterministic sample from
nvidia/Nemotron-Post-Training-Dataset-v2.
It covers the source dataset's chat, code, math, and STEM subsets. Every assistant turn
was regenerated with ornith-ai/Ornith-1.5-35B-A3B;… See the full description on the dataset page: https://huggingface.co/datasets/jzinno/Ornith-1.5-35B-A3B-Nemotron-v2-100M.details_MarinaraSpaghetti__NemoReRemix-12B
Dataset Card for Evaluation run of MarinaraSpaghetti/NemoReRemix-12B
Dataset automatically created during the evaluation run of model MarinaraSpaghetti/NemoReRemix-12B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MarinaraSpaghetti__NemoReRemix-12B.pretrain-nemotron-math-mixNormalized documents plus aligned Dolma-2 tokens and target masks.
Size
Tokens
22,927,812,461 (22.9B)
Trainable tokens
22,927,812,461 (22.9B)
Documents
21,377,358
Shards
180
UTF-8 bytes
77,994,866,327
Tokenizer
allenai/dolma2-tokenizer@5292e5d6c0f4
documents.parquet - document_id, text, part_ends, part_trainable,
must_not_split. The readable payload and the mask intent.
metadata.parquet - one text-free row per document: token span, source,
stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix.generations-nemotron-nano-9b-v2-simnpo-baselineNemotron-Reason2-GPTgenerations-nemotron-nano-9b-v2-pre_valNemotron-RL-math-advanced_calculations
Dataset Description:
The Nemotron-RL-math-advanced_calculations is a dataset designed to test a model's ability to solve complex, multi-step math problems in a multi-step agentic environment. It involves counterintuitive calculations with varying levels of function composition.
This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection of training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-math-advanced_calculations.nemotron-cc-10K-sample-translated-judgedexp-pool-nemotron-math-dolma2-tokenized
Locus EXP Nemotron Math - OLMo 2 tokenized
Pretokenized experiment pool for reproducible proxy-training runs.
MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment.
shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs.
offsets.bin stores little-endian int64 document boundaries.
index.parquet stores document IDs, offsets, and compact filter fields.
metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-nemotron-math-dolma2-tokenized.mistral_nemo_base-mmlu-valnemotron-cc-v2.1-hq-dqa-qwen3-tokens
Nemotron-CC-v2.1 / High-Quality-DQA — tokenized with the Qwen3-8B tokenizer
Question/answer pairs extracted from nvidia/Nemotron-CC-v2.1
(High-Quality-DQA subset) and tokenized with the Qwen/Qwen3-8B tokenizer.
In the source data each row is a web document whose tail carries synthetic QA pairs marked
Question: / Answer:. Here that document is split into its original prose (context) and
the individual QA pairs, each tokenized separately. The Question: / Answer: marker keywords… See the full description on the dataset page: https://huggingface.co/datasets/jackyk02/nemotron-cc-v2.1-hq-dqa-qwen3-tokens.nemotron-sft-code-focused-stage1-2-ChatML
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 50,000
Total Tokens: 415,605,764
Average Tokens per Sample: 8312.1
Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-code-focused-stage1-2-ChatML.fsid-curated-nemotron-9b-target-100nemotron_cc_v2_hq_packed4096
Nemotron-CC-v2 High-Quality, packed to 4096 tokens
5% subset of nvidia/Nemotron-CC-v2
High-Quality documents, tokenized with google/t5gemma-2-270m-270m (add_special_tokens=False, EOS
appended per document) and greedily packed into sequences of at most 4096 tokens.
A document is never split across a pack boundary; documents longer than 4096
are truncated to their own pack. Every pack ends on an EOS/document boundary.
Schema
index (int64): running pack id
input_ids… See the full description on the dataset page: https://huggingface.co/datasets/lillian039/nemotron_cc_v2_hq_packed4096.details_cognitivecomputations__dolphin-2.9.3-mistral-nemo-12b
Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9.3-mistral-nemo-12b
Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9.3-mistral-nemo-12b.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_cognitivecomputations__dolphin-2.9.3-mistral-nemo-12b.nemotron-sft-advanced-stage1-2-ChatML-V1
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 50,000
Total Tokens: 306,946,919
Average Tokens per Sample: 6138.9
Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-advanced-stage1-2-ChatML-V1.nemotron-sft-benchmark-focused-stage1-2-ChatML-V1
Nemotron SFT Dataset (Chat Template Formatted)
Overview
This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets.
Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields.
Statistics
Total Samples: 50,000
Total Tokens: 428,330,639
Average Tokens per Sample: 8566.6
Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-benchmark-focused-stage1-2-ChatML-V1.
