CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-ClimbMix ClimbMix Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠 HOMEPAGE Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models. Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.tabulartext-generation100M<n<1B128 likes6.6k downloads11mo agoHugging Face02fineinstructions /fineinstructions_nemotron ✨ Note: For all FineInstructions resources please visit: https://huggingface.co/fineinstructions This dataset is ~1B+ synthetic instruction-answer pairs or ~300B tokens created using the FineInstructions pipeline. The FineInstructions pipeline was run over the raw pre-training documents in the Nemotron-CC pre-training corpus (a subset of high-quality documents from CommonCrawl). See our paper for more details. Each .parquet file in the data folder has a corresponding judge-*.json file that… See the full description on the dataset page: https://huggingface.co/datasets/fineinstructions/fineinstructions_nemotron.tabular1B<n<10B28 likes2.6k downloads8mo agoHugging Face03nvidia /Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1 Dataset Description: We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838 different… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.tabular10K<n<100K32 likes1.4k downloads7mo agoHugging Face04nvidia /Nemotron-RL-Ultra-Training-Blends Dataset Description: This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Each subset is a separate blend; see the recipe for how the blends are used. The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Ultra-Training-Blends.tabulartext-generation10K<n<100K19 likes1.4k downloads2mo agoHugging Face05mlfoundations-dev /OpenReasoning-Nemotron-7B_eval_8179 mlfoundations-dev/OpenReasoning-Nemotron-7B_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 79.0 98.8 89.0 81.7 60.1 62.5 50.6 46.8 68.7 13.3 49.6 59.7 AIME24 Average Accuracy: 79.00% ± 1.42% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 70.00%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/OpenReasoning-Nemotron-7B_eval_8179.tabular10K<n<100K0 likes1.2k downloads1y agoHugging Face06nvidia /Nemotron-RL-Agentic-SWE-Pivot-v1 Dataset Description: The SWE-RL dataset provides GitHub issues for training and validating real-world software engineering agents using the OpenHands environment in NeMo Gym. The dataset is a refactored version of the SWE-Gym and R2E-Gym datasets to support the NeMo Gym input format. This dataset is released as part of NVIDIA NeMo Gym, a framework for building reinforcement learning environments to train large language models. NeMo Gym contains a growing collection of training… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1.tabular10K<n<100K15 likes1.1k downloads3mo agoHugging Face07mlfoundations-dev /Nemotron-Research-Reasoning-Qwen-1.5B_eval_569atabular1K<n<10K0 likes1.1k downloads1y agoHugging Face08nvidia /Nemotron-Cascade-2-RL-data Dataset Description: The Nemotron-Cascade-2-RL dataset is a curated reinforcement learning (RL) dataset blend used to train Nemotron-Cascade-2-30B-A3B model. It includes instruction-following RL, multi-domain RL, on-policy distillation, and software engineering RL (SWE-RL) data. This dataset is ready for commercial use. The dataset contains the following subset: IF-RL Contains 45,879 training samples for instruction-following RL. Our curation process mainly… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Cascade-2-RL-data.tabular10K<n<100K52 likes931 downloads6mo agoHugging Face09jamesdborin /Nemotron-SFT-Agentic-v2-prompt-only Nemotron-SFT-Agentic-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Agentic-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Agentic-v2-prompt-only.tabular100K<n<1M0 likes580 downloads3mo agoHugging Face10mlfoundations-dev /AceReason-Nemotron-7B_eval_118b mlfoundations-dev/AceReason-Nemotron-7B_eval_118b Precomputed model outputs for evaluation. Evaluation Results LiveCodeBenchv5_official Average Accuracy: 43.85% ± 0.24% Number of Runs: 3 Run Accuracy Questions Solved Total Questions 1 43.37% 121 279 2 44.09% 123 279 3 44.09% 123 279 tabularn<1K0 likes453 downloads1y agoHugging Face11SultanR /nemotron-mc-en-ar-midtrain nemotron-mc-en-ar-midtrain Arabic translation of the Nemotron-Pretraining-Multiple-Choice config of Nemotron-Pretraining-Specialized-v1.2 (pinned revision 807afc1). Translated with google/gemma-4-12B-it (bf16, greedy) on A100s. All 23,926,492 source rows are present, none dropped. English source and Arabic translation sit in the same row, so the dataset serves as a parallel corpus as well as an Arabic one. A sibling corpus from the same pipeline is available at… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-mc-en-ar-midtrain.tabulartext-generation10M<n<100M0 likes412 downloads1mo agoHugging Face12nvidia /Nemotron-RL-Instruction-Following-MultiTurnChat-v1 Dataset Description: The MultiChallenge Dataset is a rigorous benchmark designed to improve large language models in complex multi-turn conversations by explicitly targeting inference memory, instruction retention, version editing, and self-coherence. It employs a unique "model breaking" methodology where tasks are tested against advanced models (Nemotron-Nano-V2 and Qwen3-235B-A22B-Thinking-2507) to expose failure modes. A sample is only accepted into the dataset if the task is… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-MultiTurnChat-v1.tabular1K<n<10K4 likes407 downloads7mo agoHugging Face13SultanR /nemotron-r1-en-ar-midtrain nemotron-r1-en-ar-midtrain Arabic translation of the Llama_Nemotron_Post_Training_Dataset_reasoning_r1 split of smoltalk2 (config Mid, pinned revision fc6cc21): reasoning traces with <think> blocks in a conversational format. Translated with RedHatAI/gemma-4-26B-A4B-it-FP8-dynamic (greedy) on H100s. FP8 was verified lossless against its bf16 parent before the run (chrF 96.4, 0 of 510 chunks materially diverged). All 3,644,790 source rows are present, none dropped. Sibling… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/nemotron-r1-en-ar-midtrain.tabulartext-generation1M<n<10M0 likes336 downloads1mo agoHugging Face14lillian039 /nemotron_cc_v2_hq_packed4096_200shard Nemotron-CC-v2 High-Quality, packed to 4096 tokens (train) Documents from nvidia/Nemotron-CC-v2 High-Quality subset, tokenized with google/t5gemma-2-270m-270m (add_special_tokens=False, EOS appended per document) and greedily packed into sequences of at most 4096 tokens. A document is never split across a pack boundary; documents longer than 4096 are truncated to their own pack. Every pack ends on an EOS/document boundary. Schema index (int64): running pack id… See the full description on the dataset page: https://huggingface.co/datasets/lillian039/nemotron_cc_v2_hq_packed4096_200shard.tabulartext-generation10M<n<100M0 likes315 downloads2mo agoHugging Face15mlfoundations-dev /OpenReasoning-Nemotron-1.5B_eval_8179 mlfoundations-dev/OpenReasoning-Nemotron-1.5B_eval_8179 Precomputed model outputs for evaluation. Evaluation Results Summary Metric AIME24 AMC23 MATH500 JEEBench GPQADiamond LiveCodeBench CodeElo CodeForces AIME25 HLE LiveCodeBenchv5 HMMT Accuracy 49.7 83.0 78.0 49.4 31.0 35.5 19.8 14.6 40.7 12.0 24.3 32.3 AIME24 Average Accuracy: 49.67% ± 1.20% Number of Runs: 10 Run Accuracy Questions Solved Total Questions 1 50.00%… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/OpenReasoning-Nemotron-1.5B_eval_8179.tabular10K<n<100K0 likes274 downloads1y agoHugging Face16placeholderlabs /exp-pool-nemotron-math-dolma2-tokenized Locus EXP Nemotron Math - OLMo 2 tokenized Pretokenized experiment pool for reproducible proxy-training runs. MANIFEST.json is the authoritative schema, provenance, checksums, and train/holdout assignment. shards/shard-NNNNN/tokens.bin stores little-endian int32 token IDs. offsets.bin stores little-endian int64 document boundaries. index.parquet stores document IDs, offsets, and compact filter fields. metadata.parquet stores complete source metadata and is downloaded only for… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/exp-pool-nemotron-math-dolma2-tokenized.tabulartext-generation10M<n<100M0 likes268 downloads1mo agoHugging Face17lillian039 /nemotron_cc_v2_hq_packed4096 Nemotron-CC-v2 High-Quality, packed to 4096 tokens 5% subset of nvidia/Nemotron-CC-v2 High-Quality documents, tokenized with google/t5gemma-2-270m-270m (add_special_tokens=False, EOS appended per document) and greedily packed into sequences of at most 4096 tokens. A document is never split across a pack boundary; documents longer than 4096 are truncated to their own pack. Every pack ends on an EOS/document boundary. Schema index (int64): running pack id input_ids… See the full description on the dataset page: https://huggingface.co/datasets/lillian039/nemotron_cc_v2_hq_packed4096.tabulartext-generation1M<n<10M0 likes262 downloads2mo agoHugging Face18jamesdborin /Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only Prompt-only extraction from nvidia/Nemotron-SFT-Instruction-Following-Chat-v2. Files: prompts.csv: one prompt extraction record per source row. Records include prompt, separated system_prompt, and structured tools when the source row defines available tools. Nested values are JSON-encoded inside CSV cells. summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts. null_or_empty_rows.md: row… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Instruction-Following-Chat-v2-prompt-only.tabular1M<n<10M0 likes260 downloads3mo agoHugging Face19mlfoundations-dev /AceReason-Nemotron-7B_eval_c64a mlfoundations-dev/AceReason-Nemotron-7B_eval_c64a Precomputed model outputs for evaluation. Evaluation Results LiveCodeBenchv5_v3 Average Accuracy: 41.67% ± 0.45% Number of Runs: 3 Run Accuracy Questions Solved Total Questions 1 41.04% 110 268 2 41.42% 111 268 3 42.54% 114 268 tabularn<1K0 likes201 downloads1y agoHugging Face20kshitijthakkar /nemotron-sft-balanced-2b-v1 Nemotron SFT Dataset Overview This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets. Statistics Total Samples: 200,000 Total Tokens: 1,252,287,904 Average Tokens per Sample: 6261.4 Tokenizer: Qwen/Qwen3-0.6B Random Seed: 42 Strategy: balanced Subset Distribution Subset Samples Tokens Target Completion Avg Tokens/Sample Stage-1/math 20,000 151,546,125 20,000… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-balanced-2b-v1.tabular100K<n<1M0 likes188 downloads7mo agoHugging Face21nvidia /Nemotron-RLHF-GenRM-v1 Dataset Description: This dataset is designed to train Generative Reward Models (GenRMs). It leverages reinforcement learning at scale to train accurate and robust GenRMs that generalize better than traditional Bradley-Terry models and reduce the risk of reward hacking. The dataset is composed of: Preference data focused on diverse domains A synthetic safety blend The data follows a "meta-prompt" structure where the model is instructed to act as an expert evaluation judge. For… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RLHF-GenRM-v1.tabularreinforcement-learning100K<n<1M5 likes176 downloads7mo agoHugging Face22jzinno /Ornith-1.5-35B-A3B-Nemotron-v2-100M Ornith 1.5 35B A3B Nemotron v2 100M This dataset contains 108,729 English conversations with 108,729 regenerated assistant turns and 100,014,884 generated assistant completion tokens. 100M refers to the completion-token target, not the number of examples. The prompt mix is a deterministic sample from nvidia/Nemotron-Post-Training-Dataset-v2. It covers the source dataset's chat, code, math, and STEM subsets. Every assistant turn was regenerated with ornith-ai/Ornith-1.5-35B-A3B;… See the full description on the dataset page: https://huggingface.co/datasets/jzinno/Ornith-1.5-35B-A3B-Nemotron-v2-100M.tabular100K<n<1M0 likes171 downloads1mo agoHugging Face23arpandeepk /generations-nemotron-nano-9b-v2-simnpo-gentle-igm-10btabular10K<n<100K0 likes162 downloads5mo agoHugging Face24twinkle-ai /NVIDIA-Nemotron-3-Super-120B-A12B-FP8-eval-logs-and-scorestabular100K<n<1M0 likes161 downloads6mo agoHugging Face25twinkle-ai /nemotron-nano-eval-logs-and-scorestabular100K<n<1M0 likes157 downloads7mo agoHugging Face26unlearning-cleanslate /generations-nemotron-nano-9b-v2-simnpo-gentle-baselinetabular10K<n<100K0 likes139 downloads5mo agoHugging Face27kshitijthakkar /nemotron-sft-general-focused-stage1-2-ChatML-V3 Nemotron SFT Dataset (Chat Template Formatted) Overview This dataset is a curated supervised fine-tuning (SFT) dataset built from NVIDIA's Nemotron-Cascade-SFT-Stage-1 and Stage-2 datasets. Important: This dataset uses the tokenizer's apply_chat_template() method to properly format conversations from the original messages/conversations fields. Statistics Total Samples: 496,385 Total Tokens: 1,114,218,401 Average Tokens per Sample: 2244.7 Tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/kshitijthakkar/nemotron-sft-general-focused-stage1-2-ChatML-V3.tabular100K<n<1M0 likes132 downloads8mo agoHugging Face28jamesdborin /Nemotron-Agents-Tools-and-Structured-Task-Execution-prompt-only Agents, Tools and Structured Task Execution Prompt-Only This dataset combines prompt-only datasets by capability theme for distillation experiments. It contains 675,882 unique prompts from 808,884 raw rows; 133,002 exact canonical duplicates were removed. Rows retain the canonical prompt-extraction columns and add source_repo_id for provenance. Deduplication uses normalized system_prompt, prompt, tools, and schema_str, with the first row in manifest order retained. Original… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-Agents-Tools-and-Structured-Task-Execution-prompt-only.tabular100K<n<1M0 likes132 downloads2mo agoHugging Face29placeholderlabs /pretrain-nemotron-math-mixNormalized documents plus aligned Dolma-2 tokens and target masks. Size Tokens 22,927,812,461 (22.9B) Trainable tokens 22,927,812,461 (22.9B) Documents 21,377,358 Shards 180 UTF-8 bytes 77,994,866,327 Tokenizer allenai/dolma2-tokenizer@5292e5d6c0f4 documents.parquet - document_id, text, part_ends, part_trainable, must_not_split. The readable payload and the mask intent. metadata.parquet - one text-free row per document: token span, source, stratum… See the full description on the dataset page: https://huggingface.co/datasets/placeholderlabs/pretrain-nemotron-math-mix.tabular10M<n<100M0 likes129 downloads10d agoHugging Face30unlearning-cleanslate /generations-nemotron-nano-9b-v2-simnpo-baselinetabular10K<n<100K0 likes125 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.