datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Pluto-Nano-1.0-Pretrain-v2
ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2)
Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI).
v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.nanochat-brevo-capability-data-10x
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for the complete recursive
prerequisite closure in a valid leaf-first order. Training uses project-planning
language; validation uses evidence synthesis; test uses build manifests. Eleven
deterministic structural styles vary wording, layout, and record order.
The latent graph generator and exact validator label every row. No language model
generated or labeled the data. Alternative valid orders… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data-10x.nanochat-brevo-capability-data
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for the complete recursive
prerequisite closure in a valid leaf-first order. Training uses project-planning
language; validation uses evidence synthesis; test uses build manifests. Six
deterministic structural styles vary wording, layout, and record order.
The latent graph generator and exact validator label every row. No language model
generated or labeled the data. Alternative valid orders are… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data.BCE-Prettybird-Nano-Kayra-v0.1
BCE-Prettybird-Nano-Kayra-v0.1 - 200 AI Brain Mechanism Chat
Kayra is an experimental 200-sample chat dataset developed by PROMETECH A.Ş. for research on Behavioral Consciousness Engine-style control systems. The dataset was synthetically generated using Nemotron Super and is designed to go beyond standard conversation data by exposing layered behavioral signals such as trust scoring, risk level, ethical guardrails, ego–superego balance, KPI tracking, cognitive-level analysis… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Kayra-v0.1.docvqa-nanochat
DocVQA for Nanochat
Single-page document QA dataset processed for nanochat fine-tuning.
Description
This dataset is derived from pixparse/docvqa-single-page-questions and has been processed for efficient fine-tuning of small language models with limited context windows.
Modifications from Source
OCR truncation: Answer-priority truncation ensures the answer is always present in the truncated context. Lines containing the answer are prioritized, then surrounding… See the full description on the dataset page: https://huggingface.co/datasets/morgan/docvqa-nanochat.nanochat-tool-routing-v1-65k-20260714
Nanochat Tool Routing and Continuation
This deterministic corpus teaches a decoder to choose among four declared
functions, answer directly when the request already contains the answer, ask for
missing required arguments, and continue after a masked tool result. Because
this is pretraining rather than SFT, the natural system and user text remains
ordinary supervised language-model data. Only external tool results are visible
context excluded from causal-LM loss.
Train examples:… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-tool-routing-v1-65k-20260714.nanochat-brevo-capability-data-v2
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for complete recursive
prerequisite closures in valid leaf-first orders. Each compact training document
reuses one graph for 4 worked questions, increasing answer
supervision without repeating the graph. Training uses project-planning language;
validation uses evidence synthesis; test uses build manifests. Eleven deterministic
structural styles vary wording, layout, and record order. Per-world… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data-v2.NanoCodeEval
NanoCodeEval-Nemotron-1K
NanoCodeEval-Nemotron-1K is a synthetic programming benchmark containing 1,000 coding tasks across Python, JavaScript, Java, C, and C++.
The dataset was generated with Nemotron and is designed to test whether a language model can understand a small programming request, produce a valid solution, and print the required result.
This repository contains a dataset, so this page is technically a Hugging Face dataset card rather than a model card.… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/NanoCodeEval.nanochat-brevo-protocol-probe-v3
Nanochat Brevo Protocol Probe v3
This is a bounded memorization/generalization probe, not a scaling corpus. Each
document contains one project-plan query in the notes style, has no distractor
edges, and ends with a supervised <|assistant_end|> token supplied by the
whole-document loader. The four depth-width cells (1x1, 1x2, 2x1, 2x2) are
balanced. Complete three-word label combinations are hash-partitioned; individual
components remain shared.
Train worlds: 4096
Validation… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-protocol-probe-v3.nanochat-brevo-capability-v4-49k-20260714
Nanochat Brevo Capability Pilot
Brevo presents shuffled dependency records and asks for complete recursive
prerequisite closures in valid leaf-first orders. Each compact training document
reuses one graph for 4 worked questions, increasing answer
supervision without repeating the graph. Training balances
project_plan, build_manifest language; validation uses held-out
evidence synthesis; test uses build manifests. Eleven deterministic
structural styles vary wording, layout, and… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-v4-49k-20260714.Pluto-Nano-1.0-Pretrain
ASTRAI Pluto Nano 1.0 — Pretrain Mix
Curated multilingual pretraining corpus (~37 GB parquet, ~10 B tokens after tokenization) used for the base pretrain of ASTRAI Pluto Nano 1.0, a 1 B-total / 47 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI).
Source mix
Weight
Category
Source
License
Local subdir
30 %
EN Web
HuggingFaceFW/fineweb-edu (sample-350BT)
ODC-BY 1.0
en_fineweb_edu/
10 %
EN Web
HuggingFaceFW/fineweb-edu… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain.glossapi-greek-nanochat-pretraining-dataset
Glossapi Greek Nanochat Pretraining Dataset
This repository contains the source-separated Greek corpus used to build nanochat Greek pretraining mixtures. It is intentionally not a pre-split train/validation/test export: builders load data/*.parquet, preserve source_dataset, and create deterministic experiment-specific mixes and splits downstream.
Current Snapshot
Total rows: 49474947
Total characters: 248276390721
Included source datasets: 19
Data files: 273… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/glossapi-greek-nanochat-pretraining-dataset.
