CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01pankajmathur /nemotron-nano-30b-miniswe-swebench-verified Nemotron Nano 30B + mini-swe-agent SWE-bench Verified Trajectories Agent trajectories from running NVIDIA Nemotron 3 Nano 30B A3B (MoE, 8B active params) on SWE-bench Verified using mini-swe-agent. ⚠️ Incomplete Run This benchmark was terminated early due to poor performance. The model struggled with the agentic coding task. Model Information Attribute Value Model NVIDIA Nemotron 3 Nano 30B A3B Architecture MoE (30B total, 8B active) Serving vLLM… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/nemotron-nano-30b-miniswe-swebench-verified.texttext-generationn<1K0 likes357 downloads9mo agoHugging Face02CofeAI /NanoData Dataset Description To facilitate researchers to use NanoLM for comparative analysis across different model designs, we build a curated pre-training dataset from those of existing large-scale models (i.e., Llama, Falcon, GPT-3). It covers diverse domains to improve the generalization capabilities of the resultant models. Dataset Creation The data is mainly post-processed and filtered from RedPajama and RedPajamaV2. We develop a series of cleaning steps to remove redundant… See the full description on the dataset page: https://huggingface.co/datasets/CofeAI/NanoData.texttext-generation1M<n<10M3 likes240 downloads2y agoHugging Face03ai4privacy /openpii-masking-nano-1k OpenPII Nano: Multilingual PII Masking Sample A nano-sized stratified sample of OpenPII 1.5M, perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every label that exists in the parent dataset is represented in proportion. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format License 1,000 900 100 19 30 37 7… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-nano-1k.texttoken-classification1K<n<10K3 likes107 downloads4mo agoHugging Face04Marcolini /nanochat-rtx4070-sft-mixes nanochat-rtx4070 SFT mixes Eight SFT data mixes that were trained and evaluated on a single RTX 4070, and the results each one produced. Seven of them failed. These are the actual independent variable behind the negative-results table in Bl4ckd09/nanochat-on-rtx4070. Every mix here was built deterministically, trained on the same frozen backbone with the same geometry and step count, and put through the same two-stage evaluation gate. Publishing only the winner would make the… See the full description on the dataset page: https://huggingface.co/datasets/Marcolini/nanochat-rtx4070-sft-mixes.texttext-generation1K<n<10K0 likes91 downloads17d agoHugging Face05ai4privacy /pii-masking-nano-1k PII Masking Nano: Multilingual Sample A nano-sized stratified sample of pii-masking-openpii-1.5m, the flagship release of the PII-Masking-3M family. Sampled proportionally by (source_dataset, language) so every locale and label gets representation. Asia Pacific rows appear first. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-nano-1k.texttoken-classificationn<1K0 likes76 downloads4mo agoHugging Face06sixf0ur /nano_chat Dataset Card for "nano_chat" Dataset Summary nano_chat is a synthetic dataset consisting of 2326 short dialogues in simple, learner-friendly English. It was generated using Google's Gemini 2.5 flash model and is designed for training tiny conversational language models in low-resource settings. Each dialogue simulates a realistic conversation between two speakers (A and B), using short sentences, simple grammar, and occasional small mistakes to help models generalize… See the full description on the dataset page: https://huggingface.co/datasets/sixf0ur/nano_chat.texttext-classification1K<n<10K1 likes65 downloads1y agoHugging Face07TerenceLau /nanoJEPA-base nanoJEPA EN/ZH Ultra-FineWeb Dataset This is a small pretraining dataset package for nanoJEPA. It is built by streaming openbmb/Ultra-FineWeb split en and/or zh. Files train.jsonl: {"text": "...", "source": "...", "dataset": "...", "language": "en|zh", "score": 0.0} valid.jsonl: same schema as train.jsonl test.jsonl: same schema as train.jsonl Generation Command uv run python data/build_hf_dataset.py \ --out-dir dataset/nanojepa-small \ --languages en,zh… See the full description on the dataset page: https://huggingface.co/datasets/TerenceLau/nanoJEPA-base.texttext-generation1M<n<10M0 likes65 downloads4mo agoHugging Face08pthinc /BCE-Prettybird-Nano-Themis-v0.1 BCE-Prettybird-Nano-Themis-v0.1 Synthetic Multi- Law Dataset (400 Examples) BCE-Prettybird-Nano-Themis-v0.1 Synthetic Multi-Law Dataset is a 400-example synthetic dataset developed by Prometech AŞ for experimentation with legal reasoning, instruction following, structured generation, and multi-dimensional response evaluation. Each example combines a task-specific instruction with structured reasoning and quality metadata, including BCE signals, truth and quality values… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Themis-v0.1.texttext-generationn<1K0 likes64 downloads6d agoHugging Face09curtburk /nemotron3-nano-kd-corpus Nemotron 3 Nano KD Corpus 4,302 verified reasoning traces for coding problems, generated by DeepSeek-V4-Flash (284B) and filtered by executing the generated code against real test suites. Built entirely on-premises on a single HP ZGX Fury (NVIDIA GB300 Grace Blackwell Ultra). No cloud APIs were used at any stage. Used to distill NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, raising held-out pass@1 from 23.0% to 63.4%. The corpus is model-agnostic and should be usable for any student.… See the full description on the dataset page: https://huggingface.co/datasets/curtburk/nemotron3-nano-kd-corpus.texttext-generation1K<n<10K0 likes62 downloads28d agoHugging Face10pthinc /BCE-Prettybird-Nano-OWL-v0.1 BCE-Prettybird-Nano-OWL-v0.1 - 630 Translates for Instruction-Based Learning You can leverage our Hugging Face–ready nano translation dataset, which covers a diverse set of languages including Turkish, English, German, French, Spanish, Italian, Portuguese, Dutch, Russian, Ukrainian, Polish, Czech, Slovak, Hungarian, Romanian, Bulgarian, Greek, Arabic, Persian, Hebrew, Hindi, Bengali, Urdu, Tamil, Telugu, Kannada, Malayalam, Chinese, Japanese, Korean, Indonesian, Malay, Thai… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-OWL-v0.1.texttext-classificationn<1K0 likes58 downloads5mo agoHugging Face11pthinc /BCE-Prettybird-Nano-Hephaistos-v0.1 BCE-Prettybird-Nano-Hephaistos-v0.1 - 1390 Robotics for Instruction-Based Learning BCE-Prettybird-Nano-Hephaistos-v0.1 – 1390 Robotics for Instruction-Based Learning is a bilingual Turkish–English math, sensor, robotics, and embedded-systems QA dataset designed for instruction-based learning, small language models, edge AI research, and robotics education. The dataset focuses on foundational and applied robotics topics such as linear and circular motion, forward and inverse… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Hephaistos-v0.1.texttext-classification1K<n<10K0 likes55 downloads4mo agoHugging Face12pthinc /BCE-Prettybird-Nano-Ulgen-v0.1 BCE-Prettybird-Nano-Ulgen-v0.1 Synthetic Multi- Trader Dataset (320 Examples) BCE-Prettybird-Nano-Ulgen-v0.1 Synthetic Multi-Trader Dataset (320 Examples) is a bilingual Turkish-English synthetic financial reasoning dataset containing 320 instruction-response examples designed for training and evaluating AI systems on investment, portfolio management, corporate finance, risk management, market instruments, valuation, and algorithmic trading tasks. The dataset covers capital… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Ulgen-v0.1.texttext-generationn<1K0 likes48 downloads5d agoHugging Face13srcworks-software /nanoset Sourceworks NanoSet NanoSet is an experimental dataset where the main goal is to create a usable chatbot through less training data. What is in NanoSet? NanoSet is divded into 3 major sections, containg 36 entries divided into 6 sub-topics. The structure creates 108 total lines of training data, which may be subject to change in the future. The following is a visual on the structure: 108 entries total 3 Sections, each with 36 entries: Chat Basics (Greetings, Jokes, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/srcworks-software/nanoset.texttext-generationn<1K0 likes38 downloads1y agoHugging Face14pthinc /BCE-Prettybird-Nano-Math-v0.1 BCE-Prettybird-Nano-Math-v0.1 - 500 Math Q&A Dataset for Instruction-Based Learning We are excited to introduce a comprehensive math dataset containing 500 instruction-based question-answer pairs, designed to support research in mathematical reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced calculus, probability, and… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Math-v0.1.texttext-classificationn<1K0 likes37 downloads6mo agoHugging Face15wflying /nemotron-nano-rl-math-22k Nemotron Nano RL Math 22K Nemotron Nano RL Math 22K is a 22,056-example English mathematical reasoning dataset prepared for reinforcement learning with verifiable rewards (RLVR). It contains the DAPO-Math and Skywork math profiling components identified in the metadata of NVIDIA's Nemotron-3-Nano-RL-Training-Blend, represented in a compact prompt / label / metadata JSONL schema. Each example contains a ready-to-use user prompt, a reference answer, and empirical pass-rate… See the full description on the dataset page: https://huggingface.co/datasets/wflying/nemotron-nano-rl-math-22k.texttext-generation10K<n<100K0 likes37 downloads2mo agoHugging Face16sixf0ur /nano_wiki Dataset Card for "nano_wiki" Dataset Summary nano_wiki is a synthetic encyclopedia-style text dataset generated using Google's Gemma 3 27B language model. It contains 9,107 articles in simple English, covering essential human knowledge based on the Wikipedia list of articles all languages should have (expanded version). Each entry was generated using a consistent prompt designed to produce very simple, readable language suitable for small-scale language model pretraining.… See the full description on the dataset page: https://huggingface.co/datasets/sixf0ur/nano_wiki.texttext-generation1K<n<10K0 likes34 downloads1y agoHugging Face17pthinc /BCE-Prettybird-Nano-Parrot-v0.2 BCE-Prettybird-Nano-Parrot-v0.2 - 700 Jokes for Instruction-Based Learning This dataset is a bilingual (Turkish-English mixed) comedic text collection designed for training and fine-tuning conversational AI models with humor awareness, sarcasm detection, and cultural nuance understanding. It includes short joke-style prompts, observational comedy snippets, and absurd dialogue fragments that blend everyday Turkish expressions with English punchlines, reflecting real-world… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Parrot-v0.2.texttext-classificationn<1K0 likes32 downloads4mo agoHugging Face18pthinc /BCE-Prettybird-Nano-Science-v0.1 BCE-Prettybird-Nano-Science-v0.1 - 500 Science Q&A Dataset for Instruction-Based Learning We are excited to introduce a comprehensive math-physics-chemistry-biology dataset containing 500 instruction-based question-answer pairs, designed to support research in science reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Science-v0.1.texttext-classificationn<1K0 likes29 downloads6mo agoHugging Face19pthinc /BCE-Prettybird-Nano-Apollo-v0.1 BCE-Prettybird-Nano-Apollo-v0.1 Synthetic Multi-Language Software Engineering & UI/UX Dataset (1,070 Examples) This dataset contains 1,070 synthetic, high-quality examples covering a broad range of software engineering, architecture, database development, web design, UI/UX design, and design pattern implementations across multiple programming languages and frameworks. The collection includes: SOLID principle code examples in PHP, C#, Python, C++, Java, and JavaScript Design… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Apollo-v0.1.texttext-generation1K<n<10K0 likes29 downloads4mo agoHugging Face20fs90 /nano-start-data Nano-Start Learning Dataset A small educational dataset for learning how to train language models from scratch. Dataset Description This dataset contains simple, factual examples designed to demonstrate LLM training concepts: Completions: Factual statements the model learns to continue Q&A: Question-answer pairs using chat special tokens Chat: Multi-turn conversations with system prompts The dataset is intentionally small (~276 examples) so models can be trained quickly… See the full description on the dataset page: https://huggingface.co/datasets/fs90/nano-start-data.texttext-generationn<1K0 likes28 downloads10mo agoHugging Face21pthinc /BCE-Prettybird-Nano-Kangal-v0.1 BCE-Prettybird-Nano-Kangal-v0.1 - 525 LOVE Q&A Dataset for Instruction-Based Learning The "BCE-Prettybird-Nano-Kangal-v0.1: Love Dataset" consists of 525 rows of insightful data, offering a comprehensive exploration of romantic relationships. Covering diverse aspects from sexuality and intimacy to romance, family life management, and tips on how to treat women, this dataset delves into the complexities of modern relationships. It aims to provide valuable perspectives for those… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Kangal-v0.1.texttext-classificationn<1K0 likes26 downloads23d agoHugging Face22pthinc /BCE-Prettybird-Nano-Kayra-v0.1 BCE-Prettybird-Nano-Kayra-v0.1 - 200 AI Brain Mechanism Chat Kayra is an experimental 200-sample chat dataset developed by PROMETECH A.Ş. for research on Behavioral Consciousness Engine-style control systems. The dataset was synthetically generated using Nemotron Super and is designed to go beyond standard conversation data by exposing layered behavioral signals such as trust scoring, risk level, ethical guardrails, ego–superego balance, KPI tracking, cognitive-level analysis… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Kayra-v0.1.tabulartext-classificationn<1K0 likes26 downloads4mo agoHugging Face23pthinc /BCE-Prettybird-Nano-Merkur-v0.1 BCE-Prettybird-Nano-Merkur-v0.1 - 4300 Chatting for Instruction-Based Learning BCE-Prettybird-Nano-Merkur-v0.1 – This dataset contains 4,300 bilingual Turkish-English conversational dialogue samples designed for training and fine-tuning conversational AI systems, chatbots, and large language models. The dataset includes four carefully curated categories: Flirty General Chat, featuring playful and socially engaging conversations; Polite General Conversation, focused on respectful… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Merkur-v0.1.texttext-classification1K<n<10K0 likes23 downloads4mo agoHugging Face24pthinc /BCE-Prettybird-Nano-Thoth-v0.1 BCE-Prettybird-Nano-Thoth-v0.1 320 Latex Katex Math Q&A Dataset for Instruction-Based Learning BCE-Prettybird-Nano-Thoth-v0.1 is a compact KaTeX/LaTeX-oriented BCE dataset released under pthinc/BCE-Prettybird-Nano-Thoth-v0.1, designed to explore how small behavioral-control datasets can teach structured mathematical expression, symbolic formatting, and explanation consistency in LLM workflows. The dataset contains 320 question–answer pairs focused on KaTeX and LaTeX usage, covering… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Thoth-v0.1.texttext-classificationn<1K0 likes20 downloads4mo agoHugging Face25exnivo /NanoCodeEval NanoCodeEval-Nemotron-1K NanoCodeEval-Nemotron-1K is a synthetic programming benchmark containing 1,000 coding tasks across Python, JavaScript, Java, C, and C++. The dataset was generated with Nemotron and is designed to test whether a language model can understand a small programming request, produce a valid solution, and print the required result. This repository contains a dataset, so this page is technically a Hugging Face dataset card rather than a model card.… See the full description on the dataset page: https://huggingface.co/datasets/exnivo/NanoCodeEval.tabulartext-generation1K<n<10K1 likes20 downloads2mo agoHugging Face26pthinc /BCE-Prettybird-Nano-Parrot-v0.1 BCE-Prettybird-Nano-Parrot-v0.1 - 200 Jokes for Instruction-Based Learning This dataset is a bilingual (Turkish-English mixed) comedic text collection designed for training and fine-tuning conversational AI models with humor awareness, sarcasm detection, and cultural nuance understanding. It includes short joke-style prompts, observational comedy snippets, and absurd dialogue fragments that blend everyday Turkish expressions with English punchlines, reflecting real-world… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Parrot-v0.1.texttext-classificationn<1K1 likes19 downloads4mo agoHugging Face27wflying /nemotron-nano-rl-mcqa-19k Nemotron Nano RL MCQA 19K Nemotron Nano RL MCQA 19K is a 19,670-example English multiple-choice question answering dataset prepared for reinforcement learning with verifiable rewards (RLVR). Its nano_v3_sft_profiled_stem_mcqa identifier and schema correspond to the knowledge-MCQA component of NVIDIA's Nemotron-3-Nano-RL-Training-Blend, represented here in a compact prompt / label / metadata JSONL format. Each record contains a formatted user prompt, the correct option identifier… See the full description on the dataset page: https://huggingface.co/datasets/wflying/nemotron-nano-rl-mcqa-19k.textquestion-answering10K<n<100K0 likes19 downloads2mo agoHugging Face28pthinc /BCE-Prettybird-Nano-Apep-v0.1 BCE-Prettybird-Nano-Apep-v0.1 - 1025 Jokes for Instruction-Based Learning This dataset is a bilingual (Turkish-English mixed) comedic text collection designed for training and fine-tuning conversational AI models with humor awareness, sarcasm detection, and cultural nuance understanding. It includes short joke-style prompts, observational comedy snippets, and absurd dialogue fragments that blend everyday Turkish expressions with English punchlines, reflecting real-world… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Apep-v0.1.texttext-classification1K<n<10K0 likes15 downloads5mo agoHugging Face29xjh111 /nanoresearch-20topics NanoResearch 20 Topics This dataset contains 20 research-task specifications used to evaluate NanoResearch across multiple machine-learning domains. Each example describes a compact research problem, expected baselines, datasets, and user-facing requirements for generating an implementation-oriented research plan. Schema Each record contains: question_id: unique task identifier. domain: research domain, such as NLP, CV, Tabular ML, Time Series, Graph ML, Audio, or… See the full description on the dataset page: https://huggingface.co/datasets/xjh111/nanoresearch-20topics.texttext-generationn<1K2 likes11 downloads5mo agoHugging Face30sapbot /nemotron-3-nano-30b-a3b-435xTrace of Nemotron 3 Nano 30B A3B LLM by NVidia. Data is presented in ShareGPT format and each conversation split by newline. Ready to be used for fine-tuning. Brought to you by sapbot from Romarchive texttext-generationn<1K0 likes6 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.