CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-Terminal-Corpus Terminal-Corpus: Large-Scale SFT Dataset for Terminal Agents Terminal-Corpus is a large-scale Supervised Fine-Tuning (SFT) dataset designed to scale the terminal interaction capabilities of Large Language Models (LLMs). Developed by NVIDIA, this dataset was built using the Terminal-Task-Gen pipeline, which combines dataset adaptation with synthetic task generation across diverse domains. 🚀 Key Results & Performance The high-quality trajectories in Terminal-Corpus enable… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Terminal-Corpus.textquestion-answering100K<n<1M144 likes92k downloads7mo agoHugging Face02nvidia /Nemotron-CC-v2gated Nemotron-Pre-Training-Dataset-v1 Release Data Overview This pretraining dataset, for generative AI model training, preserves high-value math and code while enriching it with diverse multilingual Q&A, fueling the next generation of intelligent, globally-capable models. This dataset supports NVIDIA Nemotron Nano 2, a family of large language models (LLMs) that consists of the NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-9B-v2-Base, and… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.texttext-generation1B<n<10B142 likes58k downloads3mo agoHugging Face03Helsinki-NLP /nemotron-cc-translated Helsinki-NLP/nemotron-cc-translated nemotron-cc-tanslated is a collection of automatically translated documents from nemotron-cc taken out of the high-quality subset. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 156,431,999 documents with over 70 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 2.4 trillion tokens and the translated documents are aligned across all languages. v1.1… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/nemotron-cc-translated.texttranslation1B<n<10B5 likes55k downloads5mo agoHugging Face04nvidia /Nemotron-CC-Math-v1gated Nemotron-Pre-Training-Dataset-v1 Release 👩‍💻 Authors: Rabeeh Karimi Mahabadi, Sanjeev Satheesh 📘 Paper: Nemotron-cc-math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset 📝 Blog: Nemotron-cc-math blog Data Overview We’re excited to introduce Nemotron-CC-Math - a large-scale, high-quality math corpus extracted from Common Crawl which was used in nemotron pre-training. This dataset is built to preserve and surface high-value mathematical and code content… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-CC-Math-v1.texttext-generation100M<n<1B100 likes26k downloads9mo agoHugging Face05nvidia /Nemotron-Math-v2 Nemotron-Math-v2 This repository contains the dataset accompanying the paper Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision. Code: NeMo-Skills Documentation: NeMo-Skills Nemotron-Math-v2 Documentation Dataset Description Nemotron-Math-v2 is a large-scale mathematical reasoning dataset containing approximately 347K high-quality mathematical problems and 7M model-generated reasoning trajectories. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Math-v2.texttext-generation1M<n<10M192 likes24k downloads8mo agoHugging Face06nvidia /Nemotron-CC-v2.1gated Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-CC-v2.1.texttext-generation1B<n<10B139 likes23k downloads9mo agoHugging Face07nvidia /Nemotron-Post-Training-Dataset-v1 Nemotron-Post-Training-Dataset-v1 Release This dataset is a compilation of SFT data that supports improvements of math, code, stem, general reasoning, and tool calling capabilities of the original Llama instruct model Llama-3.3-Nemotron-Super-49B-v1.5. Llama-3.3-Nemotron-Super-49B-v1.5 is an LLM which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). Llama-3.3-Nemotron-Super-49B-v1.5 offers a great tradeoff between model accuracy and efficiency. Efficiency… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v1.text10M<n<100M194 likes14k downloads1y agoHugging Face08nvidia /Nemotron-ClimbLab ClimbLab Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠 HOMEPAGE Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models. Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbLab.text-generation1B<n<10B38 likes13k downloads1y agoHugging Face09nvidia /Nemotron-SFT-Math-v4 Nemotron-SFT-Math-v4 Dataset Description: Nemotron-SFT-Math-v4 is a large-scale mathematical reasoning dataset containing model-generated reasoning trajectories. Solutions in this version are generated using DeepSeek-V4-Pro on High inference mode. The problems in this dataset are sourced from nvidia/Nemotron-Math-v2, which contains high-quality mathematical problems derived from the Art of Problem Solving (AoPS) community and Math StackExchange/MathOverflow… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-Math-v4.texttext-generation100K<n<1M46 likes6.4k downloads1mo agoHugging Face10nvidia /Nemotron-Personas-Korea Nemotron-Personas-Korea 우리나라 실제 분포에 기반한 합성 페르소나를 위한 복합 AI 시스템 A compound AI approach to personas grounded in real-world distributions 데이터셋 개요 (Overview) Nemotron-Personas-Korea는 대한민국의 실제 인구통계학적·지리적·성격 특성 분포를 기반으로 합성된 오픈소스 페르소나 데이터셋(CC BY 4.0)으로, 우리나라 인구의 다양성과 특성을 폭넓게 반영하도록 설계되었습니다. 이는 최초의 대규모 우리말 페르소나 데이터셋이며, 이름, 성별, 나이, 혼인 상태, 교육 수준, 직업, 거주 지역 등의 속성을 실제 대한민국 국가데이터처 국가통계포털(KOSIS), 대법원, 국민건강보험공단, 농촌경제연구원, NAVER Cloud 통계 자료를 기반으로 합성하였습니다.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Korea.imagetext-generation1M<n<10M557 likes6.2k downloads3mo agoHugging Face11nvidia /Nemotron-Post-Training-Dataset-v2gated Nemotron-Post-Training-Dataset-v2 Release Data Overview This dataset adds to NVIDIA’s post-training dataset releases with an extension of SFT and RL data into five target languages: Spanish, French, German, Italian and Japanese. The data supports improvements of math, code, general reasoning, and instruction following capabilities of the NVIDIA-Nemotron-Nano-9B-v2-Base, in support of release of NVIDIA-Nemotron-Nano-8B-v2-Reasoning. NVIDIA-Nemotron-Nano-9B is a family of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Post-Training-Dataset-v2.text1M<n<10M154 likes5.6k downloads1y agoHugging Face12nvidia /Nemotron-Pretraining-Specialized-v1 Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.texttext-generation10M<n<100M87 likes5.6k downloads9mo agoHugging Face13nvidia /Nemotron-AIQ-Agentic-Safety-Dataset-1.0 Nemotron-AIQ Agentic Safety Dataset Dataset Summary Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.texttext-generation10K<n<100K18 likes4.5k downloads10mo agoHugging Face14nvidia /Nemotron-PII Nemotron-PII: Synthesized Data for Privacy-Preserving AI Dataset Description Nemotron‑PII is a synthetic, persona‑grounded dataset for training and evaluating detection of Personally Identifiable Information (PII) and Protected Health Information (PHI) in text at production quality. It contains 100,000 English records across 50+ industries with span‑level annotations for 55+ PII/PHI categories, generated with NVIDIA NeMo Data Designer using synthetic personas grounded in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-PII.texttoken-classification100K<n<1M114 likes4.1k downloads9mo agoHugging Face15nvidia /Nemotron-Pretraining-Code-v2gated Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v2.texttext-generation100M<n<1B135 likes4k downloads9mo agoHugging Face16nvidia /Nemotron-Pretraining-Code-v1gated Nemotron-Pre-Training-Dataset-v1 Release Data Overview This pretraining dataset, for generative AI model training, preserves high-value math and code while enriching it with diverse multilingual Q&A, fueling the next generation of intelligent, globally-capable models. This dataset supports NVIDIA Nemotron Nano 2, a family of large language models (LLMs) that consists of the NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-9B-v2-Base, and NVIDIA-Nemotron-Nano-12B-v2-Base… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Code-v1.texttext-generation100M<n<1B78 likes4k downloads9mo agoHugging Face17nvidia /Nemotron-Personas-Japan Nemotron-Personas-Japan 現実世界の分布に基づいたペルソナ生成のための複合AIアプローチ データセット概要 (Dataset Overview) Nemotron-Personas-Japan は、日本における人口の多様性と豊かさを捉えることを目的とし、実世界の人口統計、地理的分布、性格特性の分布に基づいて合成的に生成されたペルソナのオープンソースデータセットです。名前、性別、年齢、背景、婚姻状況、学歴、職業、居住地などの統計に基づいて生成した初のデータセットされた Nemotron-Personas の日本語版です。本バージョンでは、日本語における多様なモデリングユースケースに適した高品質のペルソナを提供します Nemotron-Personas-Japan は、日本のモデル開発者が重要な地域固有の人口統計や文化的背景を取り入れたソブリンAIシステムを開発することを支援します。本データセットは、日本の地理的・人口統計的な実分布を反映することで、合成データの多様性を高め、バイアスを軽減し、model… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Japan.imagetext-generation1M<n<10M130 likes3.5k downloads9mo agoHugging Face18nvidia /Nemotron-Personas-USA Nemotron-Personas-USA A compound AI approach to personas grounded in real-world distributions v1.1 Update The v1.1 update introduces the following changes: leverage openai/gpt-oss-120b model instead of mistralai/Mixtral-8x22B-v0.1 model to improve data quality and diversity increase the number of records from 100k to 1M, for a total of 0.94B tokens update the dataset name to Nemotron-Personas-USA in order to differentiate it from other region-specific datasets… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-USA.texttext-generation1M<n<10M350 likes3.4k downloads9mo agoHugging Face19nvidia /Nemotron-SFT-SWE-v3 Dataset Description: Nemotron-SFT-SWE-v3 is a software engineering instruction tuning dataset designed to advance the capabilities of LLMs on SWE-Bench style tasks. It includes agentic trajectories collected using a variety of agent harnesses, including the OpenHands, SWE-agent, and mini-SWE-agent frameworks. This dataset is ready for commercial use. Dataset Owner(s): NVIDIA Corporation Dataset Creation Date: Created on: 2026-06-04 Last Modified… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SFT-SWE-v3.texttext-generation100K<n<1M21 likes2.7k downloads4mo agoHugging Face20nvidia /Nemotron-Pretraining-Specialized-v1.1 Nemotron-Pretraining-Specialized-v1.1 Dataset Description: The Nemotron-Pretraining-Specialized-v1.1 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities in code concepts and algorithms, formal logic, economics, and multiple choice questions. The code concepts dataset is an instance of a general… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.1.texttext-generation10M<n<100M46 likes2.6k downloads7mo agoHugging Face21fineinstructions /fineinstructions_nemotron ✨ Note: For all FineInstructions resources please visit: https://huggingface.co/fineinstructions This dataset is ~1B+ synthetic instruction-answer pairs or ~300B tokens created using the FineInstructions pipeline. The FineInstructions pipeline was run over the raw pre-training documents in the Nemotron-CC pre-training corpus (a subset of high-quality documents from CommonCrawl). See our paper for more details. Each .parquet file in the data folder has a corresponding judge-*.json file that… See the full description on the dataset page: https://huggingface.co/datasets/fineinstructions/fineinstructions_nemotron.tabular1B<n<10B28 likes2.6k downloads8mo agoHugging Face22nvidia /Nemotron-Pretraining-Specialized-v1.2 Nemotron-Pretraining-Specialized-v1.2 Dataset Description: The Nemotron-Pretraining-Specialized-v1.2 dataset is part of the Nemotron Pretraining Data collection of pretraining datasets. Designed for the NVIDIA Nemotron 3 family of LLMs, this dataset contains a collection of synthetic datasets aimed to improve LLM capabilities on factual recall, moral scenarios, and diverse generative and multiple choice questions. Note: These are new datasets, not replacements.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-Specialized-v1.2.texttext-generation100M<n<1B16 likes2.6k downloads4mo agoHugging Face23ragrawal36 /nemotron-cc-v21-Parsed-QA4-Summarization-Qwen3-1.7Btext1M<n<10M0 likes2k downloads7mo agoHugging Face24laion /nemotron-terminal-corpus-unifiedtext100K<n<1M5 likes1.9k downloads6mo agoHugging Face25nvidia /Nemotron-Pretraining-SFT-v1gated Nemotron-Pre-Training-Dataset-v1 Release Data Overview This pretraining dataset, for generative AI model training, preserves high-value math and code while enriching it with diverse multilingual Q&A, fueling the next generation of intelligent, globally-capable models. This dataset supports NVIDIA Nemotron Nano 2, a family of large language models (LLMs) that consists of the NVIDIA-Nemotron-Nano-9B-v2, NVIDIA-Nemotron-Nano-9B-v2-Base, and NVIDIA-Nemotron-Nano-12B-v2-Base… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Pretraining-SFT-v1.texttext-generation100M<n<1B73 likes1.8k downloads9mo agoHugging Face26nvidia /Nemotron-Personas-India Nemotron-Personas-India A compound AI approach to personas grounded in real-world distributions वास्तविक दुनिया के वितरण पर आधारित व्यक्तित्वों के लिए एक मिश्रित AI दृष्टिकोण Dataset Overview (डेटासेट अवलोकन) Nemotron-Personas-India is an open-source (CC BY 4.0) dataset of synthetically-generated personas. This dataset is grounded in real-world demographic, geographic and personality trait distributions in India to capture the diversity and richness of… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-India.imagetext-generation1M<n<10M56 likes1.7k downloads9mo agoHugging Face27KORMo-VL /Nemotron-VLM-Dataset-v2from nvidia/Nemotron-VLM-Dataset-v2 samples are: visual7w_telling_cot: 435299 plotqa_cot: 295354 wiki_ko: 200000 wiki_en: 200000 mulberry_cot_1: 189378 mulberry_cot_2: 102279 sparsetables: 100000 mantis_instruct_cot: 67714 llava_cot_100k: 63019 visual_web_instruct_cot: 47800 chartqa_cot: 45710 docvqa_cot: 36333 tabmwp_cot: 20305 infographicsvqa_cot: 19548 hiertext: 514 image1M<n<10M0 likes1.6k downloads7mo agoHugging Face28nvidia /Nemotron-RL-knowledge-mcqa Dataset Description: The Nemotron-RL-knowledge-mcqa is a multi-domain synthetic multiple-choice question-answering (MCQA) dataset containing knowledge based questions. It combines and refines subsets of the [OpenScienceReasoning-2] (https://huggingface.co/datasets/nvidia/OpenScienceReasoning-2) dataset and other unstructured sources such as books and articles.The dataset was created using Qwen3-32B, [Qwen3-235B-A22B-Instruct-2507]… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-knowledge-mcqa.text100K<n<1M12 likes1.6k downloads10mo agoHugging Face29nvidia /Nemotron-CC-Code-v1gated Nemotron-Pre-Training-Dataset-v2.1 Dataset Description The Nemotron-Pre-Training-Dataset-v2.1 extends the previously released Nemotron pretraining datasets with refreshed, higher-quality, and more diverse data across math, code, English Common Crawl, and large-scale synthetic corpora. Designed for the NVIDIA Nemotron 3 family of LLMs, the dataset introduces new Common Crawl code extraction, 2.5T new English web tokens… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-CC-Code-v1.texttext-generation100M<n<1B30 likes1.6k downloads9mo agoHugging Face30SSBteam /nemotron_extra_sft_parquettext10K<n<100K0 likes1.5k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.