datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nemotron-nano-30b-miniswe-swebench-verified
Nemotron Nano 30B + mini-swe-agent SWE-bench Verified Trajectories
Agent trajectories from running NVIDIA Nemotron 3 Nano 30B A3B (MoE, 8B active params) on SWE-bench Verified using mini-swe-agent.
⚠️ Incomplete Run
This benchmark was terminated early due to poor performance. The model struggled with the agentic coding task.
Model Information
Attribute
Value
Model
NVIDIA Nemotron 3 Nano 30B A3B
Architecture
MoE (30B total, 8B active)
Serving
vLLM… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/nemotron-nano-30b-miniswe-swebench-verified.NanoData
Dataset Description
To facilitate researchers to use NanoLM for comparative analysis across different model designs, we build a curated pre-training dataset from those of existing large-scale models (i.e., Llama, Falcon, GPT-3). It covers diverse domains to improve the generalization capabilities of the resultant models.
Dataset Creation
The data is mainly post-processed and filtered from RedPajama and RedPajamaV2.
We develop a series of cleaning steps to remove redundant… See the full description on the dataset page: https://huggingface.co/datasets/CofeAI/NanoData.openpii-masking-nano-1k
OpenPII Nano: Multilingual PII Masking Sample
A nano-sized stratified sample of OpenPII 1.5M,
perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every
label that exists in the parent dataset is represented in proportion.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format
License
1,000
900
100
19
30
37
7… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-nano-1k.nanochat-rtx4070-sft-mixes
nanochat-rtx4070 SFT mixes
Eight SFT data mixes that were trained and evaluated on a single RTX 4070, and the results each one produced. Seven of them failed.
These are the actual independent variable behind the negative-results table in Bl4ckd09/nanochat-on-rtx4070. Every mix here was built deterministically, trained on the same frozen backbone with the same geometry and step count, and put through the same two-stage evaluation gate. Publishing only the winner would make the… See the full description on the dataset page: https://huggingface.co/datasets/Marcolini/nanochat-rtx4070-sft-mixes.pii-masking-nano-1k
PII Masking Nano: Multilingual Sample
A nano-sized stratified sample of pii-masking-openpii-1.5m,
the flagship release of the PII-Masking-3M family. Sampled proportionally by
(source_dataset, language) so every locale and label gets representation.
Asia Pacific rows appear first.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-nano-1k.nano_chat
Dataset Card for "nano_chat"
Dataset Summary
nano_chat is a synthetic dataset consisting of 2326 short dialogues in simple, learner-friendly English. It was generated using Google's Gemini 2.5 flash model and is designed for training tiny conversational language models in low-resource settings.
Each dialogue simulates a realistic conversation between two speakers (A and B), using short sentences, simple grammar, and occasional small mistakes to help models generalize… See the full description on the dataset page: https://huggingface.co/datasets/sixf0ur/nano_chat.nanoJEPA-base
nanoJEPA EN/ZH Ultra-FineWeb Dataset
This is a small pretraining dataset package for nanoJEPA. It is built by
streaming openbmb/Ultra-FineWeb split en and/or zh.
Files
train.jsonl: {"text": "...", "source": "...", "dataset": "...", "language": "en|zh", "score": 0.0}
valid.jsonl: same schema as train.jsonl
test.jsonl: same schema as train.jsonl
Generation Command
uv run python data/build_hf_dataset.py \
--out-dir dataset/nanojepa-small \
--languages en,zh… See the full description on the dataset page: https://huggingface.co/datasets/TerenceLau/nanoJEPA-base.BCE-Prettybird-Nano-Themis-v0.1
BCE-Prettybird-Nano-Themis-v0.1 Synthetic Multi- Law Dataset (400 Examples)
BCE-Prettybird-Nano-Themis-v0.1 Synthetic Multi-Law Dataset is a 400-example synthetic dataset developed by Prometech AŞ for experimentation with legal reasoning, instruction following, structured generation, and multi-dimensional response evaluation. Each example combines a task-specific instruction with structured reasoning and quality metadata, including BCE signals, truth and quality values… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Themis-v0.1.nemotron3-nano-kd-corpus
Nemotron 3 Nano KD Corpus
4,302 verified reasoning traces for coding problems, generated by
DeepSeek-V4-Flash (284B) and
filtered by executing the generated code against real test suites.
Built entirely on-premises on a single HP ZGX Fury (NVIDIA GB300 Grace Blackwell Ultra).
No cloud APIs were used at any stage.
Used to distill NVIDIA-Nemotron-3-Nano-30B-A3B-BF16,
raising held-out pass@1 from 23.0% to 63.4%. The corpus is model-agnostic and should be
usable for any student.… See the full description on the dataset page: https://huggingface.co/datasets/curtburk/nemotron3-nano-kd-corpus.BCE-Prettybird-Nano-OWL-v0.1
BCE-Prettybird-Nano-OWL-v0.1 - 630 Translates for Instruction-Based Learning
You can leverage our Hugging Face–ready nano translation dataset, which covers a diverse set of languages including Turkish, English, German, French, Spanish, Italian, Portuguese, Dutch, Russian, Ukrainian, Polish, Czech, Slovak, Hungarian, Romanian, Bulgarian, Greek, Arabic, Persian, Hebrew, Hindi, Bengali, Urdu, Tamil, Telugu, Kannada, Malayalam, Chinese, Japanese, Korean, Indonesian, Malay, Thai… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-OWL-v0.1.BCE-Prettybird-Nano-Hephaistos-v0.1
BCE-Prettybird-Nano-Hephaistos-v0.1 - 1390 Robotics for Instruction-Based Learning
BCE-Prettybird-Nano-Hephaistos-v0.1 – 1390 Robotics for Instruction-Based Learning is a bilingual Turkish–English math, sensor, robotics, and embedded-systems QA dataset designed for instruction-based learning, small language models, edge AI research, and robotics education. The dataset focuses on foundational and applied robotics topics such as linear and circular motion, forward and inverse… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Hephaistos-v0.1.BCE-Prettybird-Nano-Ulgen-v0.1
BCE-Prettybird-Nano-Ulgen-v0.1 Synthetic Multi- Trader Dataset (320 Examples)
BCE-Prettybird-Nano-Ulgen-v0.1 Synthetic Multi-Trader Dataset (320 Examples) is a bilingual Turkish-English synthetic financial reasoning dataset containing 320 instruction-response examples designed for training and evaluating AI systems on investment, portfolio management, corporate finance, risk management, market instruments, valuation, and algorithmic trading tasks. The dataset covers capital… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Ulgen-v0.1.nanoset
Sourceworks NanoSet
NanoSet is an experimental dataset where the main goal is to create a usable chatbot through less training data.
What is in NanoSet?
NanoSet is divded into 3 major sections, containg 36 entries divided into 6 sub-topics. The structure creates 108 total lines of training data, which may be subject to change in the future. The following is a visual on the structure:
108 entries total
3 Sections, each with 36 entries:
Chat Basics (Greetings, Jokes, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/srcworks-software/nanoset.BCE-Prettybird-Nano-Math-v0.1
BCE-Prettybird-Nano-Math-v0.1 - 500 Math Q&A Dataset for Instruction-Based Learning
We are excited to introduce a comprehensive math dataset containing 500 instruction-based question-answer pairs, designed to support research in mathematical reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced calculus, probability, and… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Math-v0.1.nemotron-nano-rl-math-22k
Nemotron Nano RL Math 22K
Nemotron Nano RL Math 22K is a 22,056-example English mathematical reasoning dataset prepared for reinforcement learning with verifiable rewards (RLVR). It contains the DAPO-Math and Skywork math profiling components identified in the metadata of NVIDIA's Nemotron-3-Nano-RL-Training-Blend, represented in a compact prompt / label / metadata JSONL schema.
Each example contains a ready-to-use user prompt, a reference answer, and empirical pass-rate… See the full description on the dataset page: https://huggingface.co/datasets/wflying/nemotron-nano-rl-math-22k.nano_wiki
Dataset Card for "nano_wiki"
Dataset Summary
nano_wiki is a synthetic encyclopedia-style text dataset generated using Google's Gemma 3 27B language model. It contains 9,107 articles in simple English, covering essential human knowledge based on the Wikipedia list of articles all languages should have (expanded version).
Each entry was generated using a consistent prompt designed to produce very simple, readable language suitable for small-scale language model pretraining.… See the full description on the dataset page: https://huggingface.co/datasets/sixf0ur/nano_wiki.BCE-Prettybird-Nano-Parrot-v0.2
BCE-Prettybird-Nano-Parrot-v0.2 - 700 Jokes for Instruction-Based Learning
This dataset is a bilingual (Turkish-English mixed) comedic text collection designed for training and fine-tuning conversational AI models with humor awareness, sarcasm detection, and cultural nuance understanding. It includes short joke-style prompts, observational comedy snippets, and absurd dialogue fragments that blend everyday Turkish expressions with English punchlines, reflecting real-world… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Parrot-v0.2.BCE-Prettybird-Nano-Science-v0.1
BCE-Prettybird-Nano-Science-v0.1 - 500 Science Q&A Dataset for Instruction-Based Learning
We are excited to introduce a comprehensive math-physics-chemistry-biology dataset containing 500 instruction-based question-answer pairs, designed to support research in science reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Science-v0.1.BCE-Prettybird-Nano-Apollo-v0.1
BCE-Prettybird-Nano-Apollo-v0.1 Synthetic Multi-Language Software Engineering & UI/UX Dataset (1,070 Examples)
This dataset contains 1,070 synthetic, high-quality examples covering a broad range of software engineering, architecture, database development, web design, UI/UX design, and design pattern implementations across multiple programming languages and frameworks.
The collection includes:
SOLID principle code examples in PHP, C#, Python, C++, Java, and JavaScript
Design… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Apollo-v0.1.nano-start-data
Nano-Start Learning Dataset
A small educational dataset for learning how to train language models from scratch.
Dataset Description
This dataset contains simple, factual examples designed to demonstrate LLM training concepts:
Completions: Factual statements the model learns to continue
Q&A: Question-answer pairs using chat special tokens
Chat: Multi-turn conversations with system prompts
The dataset is intentionally small (~276 examples) so models can be trained quickly… See the full description on the dataset page: https://huggingface.co/datasets/fs90/nano-start-data.BCE-Prettybird-Nano-Kangal-v0.1
BCE-Prettybird-Nano-Kangal-v0.1 - 525 LOVE Q&A Dataset for Instruction-Based Learning
The "BCE-Prettybird-Nano-Kangal-v0.1: Love Dataset" consists of 525 rows of insightful data, offering a comprehensive exploration of romantic relationships. Covering diverse aspects from sexuality and intimacy to romance, family life management, and tips on how to treat women, this dataset delves into the complexities of modern relationships. It aims to provide valuable perspectives for those… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Kangal-v0.1.BCE-Prettybird-Nano-Kayra-v0.1
BCE-Prettybird-Nano-Kayra-v0.1 - 200 AI Brain Mechanism Chat
Kayra is an experimental 200-sample chat dataset developed by PROMETECH A.Ş. for research on Behavioral Consciousness Engine-style control systems. The dataset was synthetically generated using Nemotron Super and is designed to go beyond standard conversation data by exposing layered behavioral signals such as trust scoring, risk level, ethical guardrails, ego–superego balance, KPI tracking, cognitive-level analysis… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Kayra-v0.1.BCE-Prettybird-Nano-Merkur-v0.1
BCE-Prettybird-Nano-Merkur-v0.1 - 4300 Chatting for Instruction-Based Learning
BCE-Prettybird-Nano-Merkur-v0.1 – This dataset contains 4,300 bilingual Turkish-English conversational dialogue samples designed for training and fine-tuning conversational AI systems, chatbots, and large language models. The dataset includes four carefully curated categories: Flirty General Chat, featuring playful and socially engaging conversations; Polite General Conversation, focused on respectful… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Merkur-v0.1.BCE-Prettybird-Nano-Thoth-v0.1
BCE-Prettybird-Nano-Thoth-v0.1 320 Latex Katex Math Q&A Dataset for Instruction-Based Learning
BCE-Prettybird-Nano-Thoth-v0.1 is a compact KaTeX/LaTeX-oriented BCE dataset released under pthinc/BCE-Prettybird-Nano-Thoth-v0.1, designed to explore how small behavioral-control datasets can teach structured mathematical expression, symbolic formatting, and explanation consistency in LLM workflows. The dataset contains 320 question–answer pairs focused on KaTeX and LaTeX usage, covering… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Thoth-v0.1.NanoCodeEval
NanoCodeEval-Nemotron-1K
NanoCodeEval-Nemotron-1K is a synthetic programming benchmark containing 1,000 coding tasks across Python, JavaScript, Java, C, and C++.
The dataset was generated with Nemotron and is designed to test whether a language model can understand a small programming request, produce a valid solution, and print the required result.
This repository contains a dataset, so this page is technically a Hugging Face dataset card rather than a model card.… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/NanoCodeEval.BCE-Prettybird-Nano-Parrot-v0.1
BCE-Prettybird-Nano-Parrot-v0.1 - 200 Jokes for Instruction-Based Learning
This dataset is a bilingual (Turkish-English mixed) comedic text collection designed for training and fine-tuning conversational AI models with humor awareness, sarcasm detection, and cultural nuance understanding. It includes short joke-style prompts, observational comedy snippets, and absurd dialogue fragments that blend everyday Turkish expressions with English punchlines, reflecting real-world… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Parrot-v0.1.nemotron-nano-rl-mcqa-19k
Nemotron Nano RL MCQA 19K
Nemotron Nano RL MCQA 19K is a 19,670-example English multiple-choice question answering dataset prepared for reinforcement learning with verifiable rewards (RLVR). Its nano_v3_sft_profiled_stem_mcqa identifier and schema correspond to the knowledge-MCQA component of NVIDIA's Nemotron-3-Nano-RL-Training-Blend, represented here in a compact prompt / label / metadata JSONL format.
Each record contains a formatted user prompt, the correct option identifier… See the full description on the dataset page: https://huggingface.co/datasets/wflying/nemotron-nano-rl-mcqa-19k.BCE-Prettybird-Nano-Apep-v0.1
BCE-Prettybird-Nano-Apep-v0.1 - 1025 Jokes for Instruction-Based Learning
This dataset is a bilingual (Turkish-English mixed) comedic text collection designed for training and fine-tuning conversational AI models with humor awareness, sarcasm detection, and cultural nuance understanding. It includes short joke-style prompts, observational comedy snippets, and absurd dialogue fragments that blend everyday Turkish expressions with English punchlines, reflecting real-world… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Apep-v0.1.nanoresearch-20topics
NanoResearch 20 Topics
This dataset contains 20 research-task specifications used to evaluate NanoResearch across multiple machine-learning domains. Each example describes a compact research problem, expected baselines, datasets, and user-facing requirements for generating an implementation-oriented research plan.
Schema
Each record contains:
question_id: unique task identifier.
domain: research domain, such as NLP, CV, Tabular ML, Time Series, Graph ML, Audio, or… See the full description on the dataset page: https://huggingface.co/datasets/xjh111/nanoresearch-20topics.nemotron-3-nano-30b-a3b-435xTrace of Nemotron 3 Nano 30B A3B LLM by NVidia.
Data is presented in ShareGPT format and each conversation split by newline. Ready to be used for fine-tuning.
Brought to you by sapbot from Romarchive
