CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fineinstructions /fineinstructions_nemotron ✨ Note: For all FineInstructions resources please visit: https://huggingface.co/fineinstructions This dataset is ~1B+ synthetic instruction-answer pairs or ~300B tokens created using the FineInstructions pipeline. The FineInstructions pipeline was run over the raw pre-training documents in the Nemotron-CC pre-training corpus (a subset of high-quality documents from CommonCrawl). See our paper for more details. Each .parquet file in the data folder has a corresponding judge-*.json file that… See the full description on the dataset page: https://huggingface.co/datasets/fineinstructions/fineinstructions_nemotron.tabular1B<n<10B28 likes2.6k downloads8mo agoHugging Face02mlfoundations-dev /Nemotron-Research-Reasoning-Qwen-1.5B_eval_569atabular1K<n<10K0 likes1.1k downloads1y agoHugging Face03SultanR /nemotron-mc-en-ar-midtrain nemotron-mc-en-ar-midtrain Arabic translation of the Nemotron-Pretraining-Multiple-Choice config of Nemotron-Pretraining-Specialized-v1.2 (pinned revision 807afc1). Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 23,926,492 source rows are present, none dropped. English source and Arabic translation sit in the same row, so the dataset serves as a parallel corpus as well as an Arabic one. A sibling corpus from the same pipeline is available at… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-mc-en-ar-midtrain.tabulartext-generation10M<n<100M0 likes472 downloads1mo agoHugging Face04mlfoundations-dev /AceReason-Nemotron-7B_eval_118b mlfoundations-dev/AceReason-Nemotron-7B_eval_118b Precomputed model outputs for evaluation. Evaluation Results LiveCodeBenchv5_official Average Accuracy: 43.85% ± 0.24% Number of Runs: 3 Run Accuracy Questions Solved Total Questions 1 43.37% 121 279 2 44.09% 123 279 3 44.09% 123 279 tabularn<1K0 likes402 downloads1y agoHugging Face05SultanR /nemotron-r1-en-ar-midtrain nemotron-r1-en-ar-midtrain Arabic translation of the Llama_Nemotron_Post_Training_Dataset_reasoning_r1 split of smoltalk2 (config Mid, pinned revision fc6cc21): reasoning traces with <think> blocks in a conversational format. Translated with RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic (greedy) on H100s. FP8 was verified lossless against its bf16 parent before the run (chrF 96.4, 0 of 510 chunks materially diverged). All 3,644,790 source rows are present, none dropped. Sibling… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-r1-en-ar-midtrain.tabulartext-generation1M<n<10M0 likes398 downloads1mo agoHugging Face06mlfoundations-dev /OpenReasoning-Nemotron-7B_eval_8179 mlfoundations-dev/OpenReasoning-Nemotron-7B_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 79.0 98.8 89.0 81.7 60.1 62.5 50.6 46.8 68.7 13.3 49.6 59.7 AIME24 Average Accuracy: 79.00% ± 1.42% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 70.00%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/OpenReasoning-Nemotron-7B_eval_8179.tabular10K<n<100K0 likes361 downloads1y agoHugging Face07lillian039 /nemotron_cc_v2_hq_packed4096_200shard Nemotron-CC-v2 High-Quality, packed to 4096 tokens (train) Documents from nvidia/Nemotron-CC-v2 High-Quality subset, tokenized with google/t5gemma-2-270m-270m (add_special_tokens=False, EOS appended per document) and greedily packed into sequences of at most 4096 tokens. A document is never split across a pack boundary; documents longer than 4096 are truncated to their own pack. Every pack ends on an EOS/document boundary. Schema index (int64): running pack id… See the full description on the dataset page: https://huggingface.co/datasets/lillian039/nemotron_cc_v2_hq_packed4096_200shard.tabulartext-generation10M<n<100M0 likes318 downloads2mo agoHugging Face08mlfoundations-dev /OpenReasoning-Nemotron-1.5B_eval_8179 mlfoundations-dev/OpenReasoning-Nemotron-1.5B_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 49.7 83.0 78.0 49.4 31.0 35.5 19.8 14.6 40.7 12.0 24.3 32.3 AIME24 Average Accuracy: 49.67% ± 1.20% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 50.00%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/OpenReasoning-Nemotron-1.5B_eval_8179.tabular10K<n<100K0 likes275 downloads1y agoHugging Face09kshitijthakkar /nemotron-sft-balanced-2b-v1 Nemotron SFT Dataset Overview This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets. Statistics Total Samples: 200,000 Total Tokens: 1,252,287,904 Average Tokens per Sample: 6261.4 Tokenizer: Qwen/Qwen3-0.6B Random Seed: 42 Strategy: balanced Subset Distribution Subset Samples Tokens Target Completion Avg Tokens/Sample Stage-1/math 20,000 151,546,125 20,000… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-balanced-2b-v1.tabular100K<n<1M0 likes240 downloads7mo agoHugging Face10mlfoundations-dev /AceReason-Nemotron-7B_eval_c64a mlfoundations-dev/AceReason-Nemotron-7B_eval_c64a Precomputed model outputs for evaluation. Evaluation Results LiveCodeBenchv5_v3 Average Accuracy: 41.67% ± 0.45% Number of Runs: 3 Run Accuracy Questions Solved Total Questions 1 41.04% 110 268 2 41.42% 111 268 3 42.54% 114 268 tabularn<1K0 likes201 downloads1y agoHugging Face11arpandeepk /generations-nemotron-nano-9b-v2-simnpo-gentle-igm-10btabular10K<n<100K0 likes182 downloads5mo agoHugging Face12kshitijthakkar /nemotron-sft-balanced-stage1-2 Nemotron SFT Dataset Overview This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets. Statistics Total Samples: 100,000 Total Tokens: 624,101,275 Average Tokens per Sample: 6241.0 Tokenizer: Qwen/Qwen3-0.6B Random Seed: 42 Strategy: balanced Subset Distribution Subset Samples Tokens Target Completion Avg Tokens/Sample Stage-1/math 10,000 75,402,505 10,000… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-balanced-stage1-2.tabular100K<n<1M0 likes145 downloads9mo agoHugging Face13unlearning-cleanslate /generations-nemotron-nano-9b-v2-simnpo-gentle-baselinetabular10K<n<100K0 likes140 downloads5mo agoHugging Face14jzinno /Ornith-1.5-35B-A3B-Nemotron-v2-100M Ornith 1.5 35B A3B Nemotron v2 100M This dataset contains 108,729 English conversations with 108,729 regenerated assistant turns and 100,014,884 generated assistant completion tokens. 100M refers to the completion-token target, not the number of examples. The prompt mix is a deterministic sample from nvidia/Nemotron-Post-Training-Dataset-v2. It covers the source dataset's chat, code, math, and STEM subsets. Every assistant turn was regenerated with ornith-ai/Ornith-1.5-35B-A3B;… See the full description on the dataset page: https://huggingface.co/datasets/jzinno/Ornith-1.5-35B-A3B-Nemotron-v2-100M.tabular100K<n<1M0 likes136 downloads1mo agoHugging Face15OALL /details_MarinaraSpaghetti__NemoReRemix-12B Dataset Card for Evaluation run of MarinaraSpaghetti/NemoReRemix-12B Dataset automatically created during the evaluation run of model MarinaraSpaghetti/NemoReRemix-12B. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_MarinaraSpaghetti__NemoReRemix-12B.tabular100K<n<1M0 likes133 downloads2y agoHugging Face16placeholderlabs /pretrain-nemotron-math-mixNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 22,927,812,461 (22.9B) Trainable tokens 22,927,812,461 (22.9B) Documents 21,377,358 Shards 180 UTF-8 bytes 77,994,866,327 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix.tabular10M<n<100M0 likes132 downloads13d agoHugging Face17unlearning-cleanslate /generations-nemotron-nano-9b-v2-simnpo-baselinetabular10K<n<100K0 likes126 downloads5mo agoHugging Face18Fern1221 /Nemotron-Reason2-GPTtabular100K<n<1M1 likes122 downloads1y agoHugging Face19unlearning-cleanslate /generations-nemotron-nano-9b-v2-pre_valtabular10K<n<100K0 likes121 downloads5mo agoHugging Face20nvidia /Nemotron-RL-math-advanced_calculations Dataset Description: The Nemotron-RL-math-advanced_calculations is a dataset designed to test a model's ability to solve complex, multi-step math problems in a multi-step agentic environment. It involves counterintuitive calculations with varying levels of function composition. This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection of training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-math-advanced_calculations.tabular1K<n<10K11 likes114 downloads10mo agoHugging Face21openeurollm /nemotron-cc-10K-sample-translated-judgedtabular1M<n<10M0 likes112 downloads1y agoHugging Face22placeholderlabs /exp-pool-nemotron-math-dolma2-tokenized Locus EXP Nemotron Math - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-nemotron-math-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes109 downloads1mo agoHugging Face23ringos /mistral_nemo_base-mmlu-valtabular10K<n<100K0 likes99 downloads2y agoHugging Face24jackyk02 /nemotron-cc-v2.1-hq-dqa-qwen3-tokens Nemotron-CC-v2.1 / High-Quality-DQA — tokenized with the Qwen3-8B tokenizer Question/answer pairs extracted from nvidia/Nemotron-CC-v2.1 (High-Quality-DQA subset) and tokenized with the Qwen/Qwen3-8B tokenizer. In the source data each row is a web document whose tail carries synthetic QA pairs marked Question: / Answer:. Here that document is split into its original prose (context) and the individual QA pairs, each tokenized separately. The Question: / Answer: marker keywords… See the full description on the dataset page: https://huggingface.co/datasets/jackyk02/nemotron-cc-v2.1-hq-dqa-qwen3-tokens.tabular100M<n<1B0 likes95 downloads2mo agoHugging Face25kshitijthakkar /nemotron-sft-code-focused-stage1-2-ChatML Nemotron SFT Dataset (Chat Template Formatted) Overview This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets. Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields. Statistics Total Samples: 50,000 Total Tokens: 415,605,764 Average Tokens per Sample: 8312.1 Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-code-focused-stage1-2-ChatML.tabular10K<n<100K0 likes94 downloads9mo agoHugging Face26unlearning-cleanslate /fsid-curated-nemotron-9b-target-100tabular10K<n<100K0 likes90 downloads5mo agoHugging Face27lillian039 /nemotron_cc_v2_hq_packed4096 Nemotron-CC-v2 High-Quality, packed to 4096 tokens 5% subset of nvidia/Nemotron-CC-v2 High-Quality documents, tokenized with google/t5gemma-2-270m-270m (add_special_tokens=False, EOS appended per document) and greedily packed into sequences of at most 4096 tokens. A document is never split across a pack boundary; documents longer than 4096 are truncated to their own pack. Every pack ends on an EOS/document boundary. Schema index (int64): running pack id input_ids… See the full description on the dataset page: https://huggingface.co/datasets/lillian039/nemotron_cc_v2_hq_packed4096.tabulartext-generation1M<n<10M0 likes87 downloads3mo agoHugging Face28OALL /details_cognitivecomputations__dolphin-2.9.3-mistral-nemo-12b Dataset Card for Evaluation run of cognitivecomputations/dolphin-2.9.3-mistral-nemo-12b Dataset automatically created during the evaluation run of model cognitivecomputations/dolphin-2.9.3-mistral-nemo-12b. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_cognitivecomputations__dolphin-2.9.3-mistral-nemo-12b.tabular100K<n<1M0 likes86 downloads2y agoHugging Face29kshitijthakkar /nemotron-sft-advanced-stage1-2-ChatML-V1 Nemotron SFT Dataset (Chat Template Formatted) Overview This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets. Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields. Statistics Total Samples: 50,000 Total Tokens: 306,946,919 Average Tokens per Sample: 6138.9 Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-advanced-stage1-2-ChatML-V1.tabular10K<n<100K0 likes84 downloads9mo agoHugging Face30kshitijthakkar /nemotron-sft-benchmark-focused-stage1-2-ChatML-V1 Nemotron SFT Dataset (Chat Template Formatted) Overview This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets. Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields. Statistics Total Samples: 50,000 Total Tokens: 428,330,639 Average Tokens per Sample: 8566.6 Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-benchmark-focused-stage1-2-ChatML-V1.tabular10K<n<100K0 likes82 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.