datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Medical-Reasoning-SFT-Nemotron-Nano-30B
Medical-Reasoning-SFT-Nemotron-Nano-30B
A large-scale medical reasoning dataset generated using nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, containing over 444,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions.
Dataset Overview
Metric
Value
Model
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
Total Samples
444,544
Samples with Reasoning
444,544 (100%)
Estimated Tokens
~1.01 Billion
Content Tokens
~808 Million… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Nemotron-Nano-30B.nano-siem-dataset
NanoSIEM Dataset
NanoSIEM is a provenance-aware cybersecurity, SIEM and legal-source retrieval corpus. It contains normalized vulnerability and guidance records, official-source registries, synthetic redacted SIEM events and curated Turkish safety-oriented question-answer examples.
Data policy
Records retain source URLs, retrieval time, jurisdiction, identifiers and confidence. Legal records are informational and jurisdiction-sensitive. Synthetic SIEM events do… See the full description on the dataset page: https://huggingface.co/datasets/oytunistrator/nano-siem-dataset.browsecomp-plus-md-toc-gpt5.4-nano
BrowseComp-Plus Structured 100k Corpus
This dataset is a drop-in, official-format variant of the BrowseComp-Plus 100k corpus used in our RISE Agent experiments. It keeps the same row count, document ids, URLs, and column names as the original BrowseComp-Plus corpus, but replaces each document's text field with a structured version that adds a generated table of contents and section headings.
Files
data.parquet: the corpus in the same three-column schema as the… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus-md-toc-gpt5.4-nano.nanobubbleeval
NanoBubbleEval v1.0
⚠ For NeurIPS reviewers — use this Croissant URL
Please do NOT use the URL exposed by the "Use this dataset → Croissant"
button at the top-right of this page. That URL triggers a known bug in
mlcroissant==1.0.16 (the version pinned by the
NeurIPS Croissant validator Space)
and produces a FilterFiles error that does not reflect a problem with the
dataset itself.
Use this URL instead — copy the line below verbatim into the validator's
"URL Input" tab:… See the full description on the dataset page: https://huggingface.co/datasets/EliasHossain/nanobubbleeval.key_information_extractionnemotron-nano2-safety-distill-gptoss
Nemotron Nano 2 Safety Distill — GPT-OSS
A distilled safety dataset produced using the Nemotron Nano 2 recipe with GPT-OSS-20B and GPT-OSS-120B as teacher models.
⚠️ Content Warning: This dataset includes potentially harmful prompts. Use responsibly for research purposes only.
Overview
This safety-focused distilled dataset was created by following the Nemotron Nano 2 safety recipe, adapted to use GPT-OSS-20B and GPT-OSS-120B as teacher models. Due to resource limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ericwang/nemotron-nano2-safety-distill-gptoss.nano-hotpotqa-vn
NanoHotpotQA-VN
An MTEB dataset
Massive Text Embedding Benchmark
A translated dataset from HotpotQA is a question answering dataset featuring natural, multi-hop questions, with strong supervision for supporting facts to enable more explainable question answering systems. The process of creating the VN-MTEB (Vietnamese Massive Text Embedding Benchmark) from English samples involves a new automated system: - The system uses large language models (LLMs), specifically Coherence's Aya… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/nano-hotpotqa-vn.nanochat-depo-capability-data
Nanochat Depo Capability Pilot
This dataset is a deterministic natural-language rendering of the Depo directed-cycle
successor task. Each row contains shuffled operational records, one exact multi-hop
question, and its answer. Latent worlds are generated programmatically; no rows were
written or labeled by a language model.
Splits
Split
Worlds
Queries per world
Rows
Renderer family
train
32,768
4
131,072
incident handoff, six structural styles… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-depo-capability-data.BCE-Prettybird-Nano-Themis-v0.1
BCE-Prettybird-Nano-Themis-v0.1 Synthetic Multi- Law Dataset (400 Examples)
BCE-Prettybird-Nano-Themis-v0.1 Synthetic Multi-Law Dataset is a 400-example synthetic dataset developed by Prometech AŞ for experimentation with legal reasoning, instruction following, structured generation, and multi-dimensional response evaluation. Each example combines a task-specific instruction with structured reasoning and quality metadata, including BCE signals, truth and quality values… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Themis-v0.1.BCE-Prettybird-Nano-OWL-v0.1
BCE-Prettybird-Nano-OWL-v0.1 - 630 Translates for Instruction-Based Learning
You can leverage our Hugging Face–ready nano translation dataset, which covers a diverse set of languages including Turkish, English, German, French, Spanish, Italian, Portuguese, Dutch, Russian, Ukrainian, Polish, Czech, Slovak, Hungarian, Romanian, Bulgarian, Greek, Arabic, Persian, Hebrew, Hindi, Bengali, Urdu, Tamil, Telugu, Kannada, Malayalam, Chinese, Japanese, Korean, Indonesian, Malay, Thai… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-OWL-v0.1.nanochat-npu-stem-eval
nanochat-npu-stem-eval
Pre-processed STEM evaluation data for nanochat-npu, adapted from karpathy/nanochat for Huawei 910B3 NPU.
Tasks
Task
Type
Shot
Source
Examples
Description
gpqa_diamond
multiple_choice
0-shot
Idavidrein/gpqa
198
Graduate-level science QA (Diamond subset)
gsm8k_cot
generation
8-shot
openai/gsm8k
1319
Grade school math word problems (CoT)
math_cot
generation
4-shot
HuggingFaceH4/MATH-500
500
Competition mathematics (CoT)… See the full description on the dataset page: https://huggingface.co/datasets/Sexhuis/nanochat-npu-stem-eval.BCE-Prettybird-Nano-Hephaistos-v0.1
BCE-Prettybird-Nano-Hephaistos-v0.1 - 1390 Robotics for Instruction-Based Learning
BCE-Prettybird-Nano-Hephaistos-v0.1 – 1390 Robotics for Instruction-Based Learning is a bilingual Turkish–English math, sensor, robotics, and embedded-systems QA dataset designed for instruction-based learning, small language models, edge AI research, and robotics education. The dataset focuses on foundational and applied robotics topics such as linear and circular motion, forward and inverse… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Hephaistos-v0.1.nanochat-jp-eval-bundle
nanochat-jp-eval-bundle
nanochat の日本語フォーク nanochat-jp で使用する 日本語評価データ一式(eval bundle) です.
既存の公開日本語ベンチマークを nanochat の評価コードがそのまま読める形式へ変換し,設定ファイルとあわせて配布しています.
評価は2系統あります.
CORE: ベースモデル(事前学習直後)向け.few-shot の尤度比較(multiple choice)または継続生成(language modeling)で採点します.対象タスクと shot 数は core.yaml で定義されます.
Chat: SFT / RL 後のチャットモデル向け.few-shot の実例を user/assistant のターンとして与え,生成結果を採点します.タスクは nanochat-jp の nanochat/chat_eval_common.py に登録されています.
CORE
タスク
ファイル
形式
shot 数
件数
ランダム… See the full description on the dataset page: https://huggingface.co/datasets/tohoku-nlp/nanochat-jp-eval-bundle.BCE-Prettybird-Nano-Ulgen-v0.1
BCE-Prettybird-Nano-Ulgen-v0.1 Synthetic Multi- Trader Dataset (320 Examples)
BCE-Prettybird-Nano-Ulgen-v0.1 Synthetic Multi-Trader Dataset (320 Examples) is a bilingual Turkish-English synthetic financial reasoning dataset containing 320 instruction-response examples designed for training and evaluating AI systems on investment, portfolio management, corporate finance, risk management, market instruments, valuation, and algorithmic trading tasks. The dataset covers capital… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Ulgen-v0.1.nn-auto-bench-ds
nn-auto-bench-ds
nn-auto-bench-ds is a dataset designed for key information extraction (KIE) and serves as a benchmark dataset for nn-auto-bench.
Dataset Overview
The dataset comprises 1,000 documents, categorized into the following types:
Invoice
Receipt
Passport
Bank Statement
The documents are primarily available in English, with some also in German and Arabic. Each document is annotated for key information extraction and specific tasks. The dataset can be used to… See the full description on the dataset page: https://huggingface.co/datasets/nanonets/nn-auto-bench-ds.mot-dang-chiang-mai-chiang-rai
มดแดง Mot Dang — Chiang Mai & Chiang Rai city directory
88,161 places in and around Chiang Mai (62,772) and Chiang Rai (25,389),
in Thai and English, with coordinates, categories, opening hours, and the
channels a place actually answers on — phone, LINE, Facebook, a website that
still resolves.
The name is มดแดง, mot daeng, the red ant: the thing that knows every soi
because it has walked all of them. That is the ambition. The directory exists
because mainstream mapping is thin… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/mot-dang-chiang-mai-chiang-rai.nano-msmarco-vn
NanoMSMARCO-VN
An MTEB dataset
Massive Text Embedding Benchmark
A translated dataset from MS MARCO is a collection of datasets focused on deep learning in search The process of creating the VN-MTEB (Vietnamese Massive Text Embedding Benchmark) from English samples involves a new automated system: - The system uses large language models (LLMs), specifically Coherence's Aya model, for translation. - Applies advanced embedding models to filter the translations. - Use LLM-as-a-judge to… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/nano-msmarco-vn.trace-benchmark-dataset
TRACE Benchmark (100k tier)
TRACE — Task-Relevant Applied Constraint Execution: can a solver accomplish a task
correctly while automatically honoring the preferences and constraints that matter for that
task — even when those rules were stated once, in passing, and buried in a long prior
conversation?
Blog post: nanonets.com/research/trace
Each sample is a realistic enterprise (Record-to-Report / finance) conversation: a long
transcript where constraints are sprinkled throughout… See the full description on the dataset page: https://huggingface.co/datasets/nanonets/trace-benchmark-dataset.BCE-Prettybird-Nano-Math-v0.1
BCE-Prettybird-Nano-Math-v0.1 - 500 Math Q&A Dataset for Instruction-Based Learning
We are excited to introduce a comprehensive math dataset containing 500 instruction-based question-answer pairs, designed to support research in mathematical reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced calculus, probability, and… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Math-v0.1.nanochat-npu-stem-eval
nanochat-npu-stem-eval
Pre-processed STEM evaluation data for nanochat-npu, adapted from karpathy/nanochat for Huawei 910B3 NPU.
Tasks
Task
Type
Shot
Source
Examples
Description
gpqa_diamond
multiple_choice
0-shot
Idavidrein/gpqa
198
Graduate-level science QA (Diamond subset)
gsm8k_cot
generation
8-shot
openai/gsm8k
1319
Grade school math word problems (CoT)
math_cot
generation
4-shot
HuggingFaceH4/MATH-500
500
Competition mathematics (CoT)… See the full description on the dataset page: https://huggingface.co/datasets/liujin99/nanochat-npu-stem-eval.nano-nq-vn
NanoNQ-VN
An MTEB dataset
Massive Text Embedding Benchmark
A translated dataset from NFCorpus: A Full-Text Learning to Rank Dataset for Medical Information Retrieval The process of creating the VN-MTEB (Vietnamese Massive Text Embedding Benchmark) from English samples involves a new automated system: - The system uses large language models (LLMs), specifically Coherence's Aya model, for translation. - Applies advanced embedding models to filter the translations. - Use LLM-as-a-judge… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/nano-nq-vn.BCE-Prettybird-Nano-Parrot-v0.2
BCE-Prettybird-Nano-Parrot-v0.2 - 700 Jokes for Instruction-Based Learning
This dataset is a bilingual (Turkish-English mixed) comedic text collection designed for training and fine-tuning conversational AI models with humor awareness, sarcasm detection, and cultural nuance understanding. It includes short joke-style prompts, observational comedy snippets, and absurd dialogue fragments that blend everyday Turkish expressions with English punchlines, reflecting real-world… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Parrot-v0.2.nanochat-depo-retrieval-copy1-20260715
Nanochat Depo retrieval v1
Each latent 16-node graph yields eight independent, token-aligned, depth-one
query documents. This arm exposes 1 nested edge(s) per
document. Only the answer is supervised in every document; the terminal token
is supervised only for query ordinal 7. This source is separate from and does
not alter Depo-L0 v1.
obekt-question-answer-reasoning-nano-v0.1
Obekt Nano Reasoning Dataset (v0.1)
Dataset Description
This is a small "nano" dataset containing questions, answers, and reasoning traces. It is generated using the Xiaomi MiMo V2 Flash LLM and is intended for experimental purposes, quick prototyping, and fine-tuning trials where reasoning capability is a focus.
Source Model: xiaomi/mimo-v2-flash
Contains
obekt-question-answer-reasoning-nano-v0.1.csv: The main data file.
Columns:
question: The input query.… See the full description on the dataset page: https://huggingface.co/datasets/obekt/obekt-question-answer-reasoning-nano-v0.1.BCE-Prettybird-Nano-Science-v0.1
BCE-Prettybird-Nano-Science-v0.1 - 500 Science Q&A Dataset for Instruction-Based Learning
We are excited to introduce a comprehensive math-physics-chemistry-biology dataset containing 500 instruction-based question-answer pairs, designed to support research in science reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Science-v0.1.BCE-Prettybird-Nano-Apollo-v0.1
BCE-Prettybird-Nano-Apollo-v0.1 Synthetic Multi-Language Software Engineering & UI/UX Dataset (1,070 Examples)
This dataset contains 1,070 synthetic, high-quality examples covering a broad range of software engineering, architecture, database development, web design, UI/UX design, and design pattern implementations across multiple programming languages and frameworks.
The collection includes:
SOLID principle code examples in PHP, C#, Python, C++, Java, and JavaScript
Design… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Apollo-v0.1.nano-start-data
Nano-Start Learning Dataset
A small educational dataset for learning how to train language models from scratch.
Dataset Description
This dataset contains simple, factual examples designed to demonstrate LLM training concepts:
Completions: Factual statements the model learns to continue
Q&A: Question-answer pairs using chat special tokens
Chat: Multi-turn conversations with system prompts
The dataset is intentionally small (~276 examples) so models can be trained quickly… See the full description on the dataset page: https://huggingface.co/datasets/fs90/nano-start-data.BCE-Prettybird-Nano-Kangal-v0.1
BCE-Prettybird-Nano-Kangal-v0.1 - 525 LOVE Q&A Dataset for Instruction-Based Learning
The "BCE-Prettybird-Nano-Kangal-v0.1: Love Dataset" consists of 525 rows of insightful data, offering a comprehensive exploration of romantic relationships. Covering diverse aspects from sexuality and intimacy to romance, family life management, and tips on how to treat women, this dataset delves into the complexities of modern relationships. It aims to provide valuable perspectives for those… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Kangal-v0.1.BCE-Prettybird-Nano-Kayra-v0.1
BCE-Prettybird-Nano-Kayra-v0.1 - 200 AI Brain Mechanism Chat
Kayra is an experimental 200-sample chat dataset developed by PROMETECH A.Ş. for research on Behavioral Consciousness Engine-style control systems. The dataset was synthetically generated using Nemotron Super and is designed to go beyond standard conversation data by exposing layered behavioral signals such as trust scoring, risk level, ethical guardrails, ego–superego balance, KPI tracking, cognitive-level analysis… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Kayra-v0.1.nanochat-depo-composition-depth2-w4-retry-20260715
Nanochat Depo composition v1
Each 16-node single-cycle graph yields eight independent one-query documents:
four base starts paired across query depths (1, 2).
This source contains train and validation splits only. Phase depth is
2; the materialized context width is 4.
