CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ASTRAI-labs /Pluto-Nano-1.0-Pretrain-v2 ASTRAI Pluto Nano 1.0 — Pretrain Mix (v2) Curated multilingual pretraining corpus (~50 GB parquet, ~12 B tokens after tokenization) used for ASTRAI Pluto Nano 1.0, a 1 B-total / 50 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI). v2 additions vs v1: OpenThoughts3 (CoT reasoning), openstax textbooks + peS2o (science), and reweighting for better balance. NOTE: factsense (openbmb) was used at training time but is not redistributed here due to its… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain-v2.tabulartext-generation10M<n<100M2 likes565 downloads3mo agoHugging Face02SolidSnake123 /nanochat-brevo-capability-data-10x Nanochat Brevo Capability Pilot Brevo presents shuffled dependency records and asks for the complete recursive prerequisite closure in a valid leaf-first order. Training uses project-planning language; validation uses evidence synthesis; test uses build manifests. Eleven deterministic structural styles vary wording, layout, and record order. The latent graph generator and exact validator label every row. No language model generated or labeled the data. Alternative valid orders… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data-10x.tabulartext-generation100K<n<1M0 likes54 downloads3mo agoHugging Face03SolidSnake123 /nanochat-brevo-capability-data Nanochat Brevo Capability Pilot Brevo presents shuffled dependency records and asks for the complete recursive prerequisite closure in a valid leaf-first order. Training uses project-planning language; validation uses evidence synthesis; test uses build manifests. Six deterministic structural styles vary wording, layout, and record order. The latent graph generator and exact validator label every row. No language model generated or labeled the data. Alternative valid orders are… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data.tabulartext-generation10K<n<100K0 likes29 downloads3mo agoHugging Face04pthinc /BCE-Prettybird-Nano-Kayra-v0.1 BCE-Prettybird-Nano-Kayra-v0.1 - 200 AI Brain Mechanism Chat Kayra is an experimental 200-sample chat dataset developed by PROMETECH A.Ş. for research on Behavioral Consciousness Engine-style control systems. The dataset was synthetically generated using Nemotron Super and is designed to go beyond standard conversation data by exposing layered behavioral signals such as trust scoring, risk level, ethical guardrails, ego–superego balance, KPI tracking, cognitive-level analysis… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Kayra-v0.1.tabulartext-classificationn<1K0 likes26 downloads4mo agoHugging Face05morgan /docvqa-nanochat DocVQA for Nanochat Single-page document QA dataset processed for nanochat fine-tuning. Description This dataset is derived from pixparse/docvqa-single-page-questions and has been processed for efficient fine-tuning of small language models with limited context windows. Modifications from Source OCR truncation: Answer-priority truncation ensures the answer is always present in the truncated context. Lines containing the answer are prioritized, then surrounding… See the full description on the dataset page: https://huggingface.co/datasets/morgan/docvqa-nanochat.tabularquestion-answering10K<n<100K0 likes24 downloads9mo agoHugging Face06SolidSnake123 /nanochat-tool-routing-v1-65k-20260714 Nanochat Tool Routing and Continuation This deterministic corpus teaches a decoder to choose among four declared functions, answer directly when the request already contains the answer, ask for missing required arguments, and continue after a masked tool result. Because this is pretraining rather than SFT, the natural system and user text remains ordinary supervised language-model data. Only external tool results are visible context excluded from causal-LM loss. Train examples:… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-tool-routing-v1-65k-20260714.tabulartext-generation10K<n<100K0 likes22 downloads2mo agoHugging Face07SolidSnake123 /nanochat-brevo-capability-data-v2 Nanochat Brevo Capability Pilot Brevo presents shuffled dependency records and asks for complete recursive prerequisite closures in valid leaf-first orders. Each compact training document reuses one graph for 4 worked questions, increasing answer supervision without repeating the graph. Training uses project-planning language; validation uses evidence synthesis; test uses build manifests. Eleven deterministic structural styles vary wording, layout, and record order. Per-world… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data-v2.tabulartext-generation10K<n<100K0 likes20 downloads3mo agoHugging Face08exnivo /NanoCodeEval NanoCodeEval-Nemotron-1K NanoCodeEval-Nemotron-1K is a synthetic programming benchmark containing 1,000 coding tasks across Python, JavaScript, Java, C, and C++. The dataset was generated with Nemotron and is designed to test whether a language model can understand a small programming request, produce a valid solution, and print the required result. This repository contains a dataset, so this page is technically a Hugging Face dataset card rather than a model card.… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/NanoCodeEval.tabulartext-generation1K<n<10K1 likes20 downloads2mo agoHugging Face09SolidSnake123 /nanochat-brevo-protocol-probe-v3 Nanochat Brevo Protocol Probe v3 This is a bounded memorization/generalization probe, not a scaling corpus. Each document contains one project-plan query in the notes style, has no distractor edges, and ends with a supervised <|assistant_end|> token supplied by the whole-document loader. The four depth-width cells (1x1, 1x2, 2x1, 2x2) are balanced. Complete three-word label combinations are hash-partitioned; individual components remain shared. Train worlds: 4096 Validation… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-protocol-probe-v3.tabulartext-generation1K<n<10K0 likes14 downloads3mo agoHugging Face10SolidSnake123 /nanochat-brevo-capability-v4-49k-20260714 Nanochat Brevo Capability Pilot Brevo presents shuffled dependency records and asks for complete recursive prerequisite closures in valid leaf-first orders. Each compact training document reuses one graph for 4 worked questions, increasing answer supervision without repeating the graph. Training balances project_plan, build_manifest language; validation uses held-out evidence synthesis; test uses build manifests. Eleven deterministic structural styles vary wording, layout, and… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-v4-49k-20260714.tabulartext-generation10K<n<100K0 likes14 downloads2mo agoHugging Face11ASTRAI-labs /Pluto-Nano-1.0-Pretrain ASTRAI Pluto Nano 1.0 — Pretrain Mix Curated multilingual pretraining corpus (~37 GB parquet, ~10 B tokens after tokenization) used for the base pretrain of ASTRAI Pluto Nano 1.0, a 1 B-total / 47 M-active MoE model with 64 k vocabulary and 5 target languages (EN, PT, ES, ZH, HI). Source mix Weight Category Source License Local subdir 30 % EN Web HuggingFaceFW/fineweb-edu (sample-350BT) ODC-BY 1.0 en_fineweb_edu/ 10 % EN Web HuggingFaceFW/fineweb-edu… See the full description on the dataset page: https://huggingface.co/datasets/ASTRAI-labs/Pluto-Nano-1.0-Pretrain.tabulartext-generation10M<n<100M1 likes10 downloads3mo agoHugging Face12fffoivos /glossapi-greek-nanochat-pretraining-datasetgated Glossapi Greek Nanochat Pretraining Dataset This repository contains the source-separated Greek corpus used to build nanochat Greek pretraining mixtures. It is intentionally not a pre-split train/validation/test export: builders load data/*.parquet, preserve source_dataset, and create deterministic experiment-specific mixes and splits downstream. Current Snapshot Total rows: 49474947 Total characters: 248276390721 Included source datasets: 19 Data files: 273… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/glossapi-greek-nanochat-pretraining-dataset.tabulartext-generation10M<n<100M0 likes5 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.