CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OpenMed /Medical-Reasoning-SFT-Nemotron-Nano-30B Medical-Reasoning-SFT-Nemotron-Nano-30B A large-scale medical reasoning dataset generated using nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, containing over 444,000 samples with detailed chain-of-thought reasoning for medical and healthcare questions. Dataset Overview Metric Value Model nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 Total Samples 444,544 Samples with Reasoning 444,544 (100%) Estimated Tokens ~1.01 Billion Content Tokens ~808 Million… See the full description on the dataset page: https://huggingface.co/datasets/OpenMed/Medical-Reasoning-SFT-Nemotron-Nano-30B.texttext-generation100K<n<1M46 likes323 downloads8mo agoHugging Face02oytunistrator /nano-siem-dataset NanoSIEM Dataset NanoSIEM is a provenance-aware cybersecurity, SIEM and legal-source retrieval corpus. It contains normalized vulnerability and guidance records, official-source registries, synthetic redacted SIEM events and curated Turkish safety-oriented question-answer examples. Data policy Records retain source URLs, retrieval time, jurisdiction, identifiers and confidence. Legal records are informational and jurisdiction-sensitive. Synthetic SIEM events do… See the full description on the dataset page: https://huggingface.co/datasets/oytunistrator/nano-siem-dataset.text-retrieval0 likes226 downloads9d agoHugging Face03Tevatron /browsecomp-plus-md-toc-gpt5.4-nano BrowseComp-Plus Structured 100k Corpus This dataset is a drop-in, official-format variant of the BrowseComp-Plus 100k corpus used in our RISE Agent experiments. It keeps the same row count, document ids, URLs, and column names as the original BrowseComp-Plus corpus, but replaces each document's text field with a structured version that adds a generated table of contents and section headings. Files data.parquet: the corpus in the same three-column schema as the… See the full description on the dataset page: https://huggingface.co/datasets/Tevatron/browsecomp-plus-md-toc-gpt5.4-nano.textquestion-answering100K<n<1M0 likes135 downloads3mo agoHugging Face04EliasHossain /nanobubbleeval NanoBubbleEval v1.0 ⚠ For NeurIPS reviewers — use this Croissant URL Please do NOT use the URL exposed by the "Use this dataset → Croissant" button at the top-right of this page. That URL triggers a known bug in mlcroissant==1.0.16 (the version pinned by the NeurIPS Croissant validator Space) and produces a FilterFiles error that does not reflect a problem with the dataset itself. Use this URL instead — copy the line below verbatim into the validator's "URL Input" tab:… See the full description on the dataset page: https://huggingface.co/datasets/EliasHossain/nanobubbleeval.tabularquestion-answering10K<n<100K0 likes123 downloads5mo agoHugging Face05nanonets /key_information_extractiontextquestion-answeringn<1K6 likes113 downloads1y agoHugging Face06Ericwang /nemotron-nano2-safety-distill-gptoss Nemotron Nano 2 Safety Distill — GPT-OSS A distilled safety dataset produced using the Nemotron Nano 2 recipe with GPT-OSS-20B and GPT-OSS-120B as teacher models. ⚠️ Content Warning: This dataset includes potentially harmful prompts. Use responsibly for research purposes only. Overview This safety-focused distilled dataset was created by following the Nemotron Nano 2 safety recipe, adapted to use GPT-OSS-20B and GPT-OSS-120B as teacher models. Due to resource limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ericwang/nemotron-nano2-safety-distill-gptoss.texttext-generation10K<n<100K2 likes92 downloads11mo agoHugging Face07GreenNode /nano-hotpotqa-vn NanoHotpotQA-VN An MTEB dataset Massive Text Embedding Benchmark A translated dataset from HotpotQA is a question answering dataset featuring natural, multi-hop questions, with strong supervision for supporting facts to enable more explainable question answering systems. The process of creating the VN-MTEB (Vietnamese Massive Text Embedding Benchmark) from English samples involves a new automated system: - The system uses large language models (LLMs), specifically Coherence's Aya… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/nano-hotpotqa-vn.texttext-retrieval100K<n<1M0 likes70 downloads9mo agoHugging Face08SolidSnake123 /nanochat-depo-capability-data Nanochat Depo Capability Pilot This dataset is a deterministic natural-language rendering of the Depo directed-cycle successor task. Each row contains shuffled operational records, one exact multi-hop question, and its answer. Latent worlds are generated programmatically; no rows were written or labeled by a language model. Splits Split Worlds Queries per world Rows Renderer family train 32,768 4 131,072 incident handoff, six structural styles… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-depo-capability-data.tabularquestion-answering100K<n<1M0 likes70 downloads3mo agoHugging Face09pthinc /BCE-Prettybird-Nano-Themis-v0.1 BCE-Prettybird-Nano-Themis-v0.1 Synthetic Multi- Law Dataset (400 Examples) BCE-Prettybird-Nano-Themis-v0.1 Synthetic Multi-Law Dataset is a 400-example synthetic dataset developed by Prometech AŞ for experimentation with legal reasoning, instruction following, structured generation, and multi-dimensional response evaluation. Each example combines a task-specific instruction with structured reasoning and quality metadata, including BCE signals, truth and quality values… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Themis-v0.1.texttext-generationn<1K0 likes64 downloads6d agoHugging Face10pthinc /BCE-Prettybird-Nano-OWL-v0.1 BCE-Prettybird-Nano-OWL-v0.1 - 630 Translates for Instruction-Based Learning You can leverage our Hugging Face–ready nano translation dataset, which covers a diverse set of languages including Turkish, English, German, French, Spanish, Italian, Portuguese, Dutch, Russian, Ukrainian, Polish, Czech, Slovak, Hungarian, Romanian, Bulgarian, Greek, Arabic, Persian, Hebrew, Hindi, Bengali, Urdu, Tamil, Telugu, Kannada, Malayalam, Chinese, Japanese, Korean, Indonesian, Malay, Thai… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-OWL-v0.1.texttext-classificationn<1K0 likes58 downloads5mo agoHugging Face11Sexhuis /nanochat-npu-stem-eval nanochat-npu-stem-eval Pre-processed STEM evaluation data for nanochat-npu, adapted from karpathy/nanochat for Huawei 910B3 NPU. Tasks Task Type Shot Source Examples Description gpqa_diamond multiple_choice 0-shot Idavidrein/gpqa 198 Graduate-level science QA (Diamond subset) gsm8k_cot generation 8-shot openai/gsm8k 1319 Grade school math word problems (CoT) math_cot generation 4-shot HuggingFaceH4/MATH-500 500 Competition mathematics (CoT)… See the full description on the dataset page: https://huggingface.co/datasets/Sexhuis/nanochat-npu-stem-eval.question-answering1K<n<10K0 likes58 downloads2mo agoHugging Face12pthinc /BCE-Prettybird-Nano-Hephaistos-v0.1 BCE-Prettybird-Nano-Hephaistos-v0.1 - 1390 Robotics for Instruction-Based Learning BCE-Prettybird-Nano-Hephaistos-v0.1 – 1390 Robotics for Instruction-Based Learning is a bilingual Turkish–English math, sensor, robotics, and embedded-systems QA dataset designed for instruction-based learning, small language models, edge AI research, and robotics education. The dataset focuses on foundational and applied robotics topics such as linear and circular motion, forward and inverse… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Hephaistos-v0.1.texttext-classification1K<n<10K0 likes55 downloads4mo agoHugging Face13tohoku-nlp /nanochat-jp-eval-bundle nanochat-jp-eval-bundle nanochat の日本語フォーク nanochat-jp で使用する 日本語評価データ一式(eval bundle) です. 既存の公開日本語ベンチマークを nanochat の評価コードがそのまま読める形式へ変換し,設定ファイルとあわせて配布しています. 評価は2系統あります. CORE: ベースモデル(事前学習直後)向け.few-shot の尤度比較(multiple choice)または継続生成(language modeling)で採点します.対象タスクと shot 数は core.yaml で定義されます. Chat: SFT / RL 後のチャットモデル向け.few-shot の実例を user/assistant のターンとして与え,生成結果を採点します.タスクは nanochat-jp の nanochat/chat_eval_common.py に登録されています. CORE タスク ファイル 形式 shot 数 件数 ランダム… See the full description on the dataset page: https://huggingface.co/datasets/tohoku-nlp/nanochat-jp-eval-bundle.multiple-choice0 likes51 downloads1mo agoHugging Face14pthinc /BCE-Prettybird-Nano-Ulgen-v0.1 BCE-Prettybird-Nano-Ulgen-v0.1 Synthetic Multi- Trader Dataset (320 Examples) BCE-Prettybird-Nano-Ulgen-v0.1 Synthetic Multi-Trader Dataset (320 Examples) is a bilingual Turkish-English synthetic financial reasoning dataset containing 320 instruction-response examples designed for training and evaluating AI systems on investment, portfolio management, corporate finance, risk management, market instruments, valuation, and algorithmic trading tasks. The dataset covers capital… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Ulgen-v0.1.texttext-generationn<1K0 likes48 downloads5d agoHugging Face15nanonets /nn-auto-bench-ds nn-auto-bench-ds nn-auto-bench-ds is a dataset designed for key information extraction (KIE) and serves as a benchmark dataset for nn-auto-bench. Dataset Overview The dataset comprises 1,000 documents, categorized into the following types: Invoice Receipt Passport Bank Statement The documents are primarily available in English, with some also in German and Arabic. Each document is annotated for key information extraction and specific tasks. The dataset can be used to… See the full description on the dataset page: https://huggingface.co/datasets/nanonets/nn-auto-bench-ds.question-answering1K<n<10K5 likes46 downloads2y agoHugging Face16NaNoBotCo /mot-dang-chiang-mai-chiang-rai มดแดง Mot Dang — Chiang Mai & Chiang Rai city directory 88,161 places in and around Chiang Mai (62,772) and Chiang Rai (25,389), in Thai and English, with coordinates, categories, opening hours, and the channels a place actually answers on — phone, LINE, Facebook, a website that still resolves. The name is มดแดง, mot daeng, the red ant: the thing that knows every soi because it has walked all of them. That is the ambition. The directory exists because mainstream mapping is thin… See the full description on the dataset page: https://huggingface.co/datasets/NaNoBotCo/mot-dang-chiang-mai-chiang-rai.geospatialtext-retrieval10K<n<100K0 likes46 downloads17d agoHugging Face17GreenNode /nano-msmarco-vn NanoMSMARCO-VN An MTEB dataset Massive Text Embedding Benchmark A translated dataset from MS MARCO is a collection of datasets focused on deep learning in search The process of creating the VN-MTEB (Vietnamese Massive Text Embedding Benchmark) from English samples involves a new automated system: - The system uses large language models (LLMs), specifically Coherence's Aya model, for translation. - Applies advanced embedding models to filter the translations. - Use LLM-as-a-judge to… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/nano-msmarco-vn.texttext-retrieval100K<n<1M0 likes43 downloads9mo agoHugging Face18nanonets /trace-benchmark-dataset TRACE Benchmark (100k tier) TRACE — Task-Relevant Applied Constraint Execution: can a solver accomplish a task correctly while automatically honoring the preferences and constraints that matter for that task — even when those rules were stated once, in passing, and buried in a long prior conversation? Blog post: nanonets.com/research/trace Each sample is a realistic enterprise (Record-to-Report / finance) conversation: a long transcript where constraints are sprinkled throughout… See the full description on the dataset page: https://huggingface.co/datasets/nanonets/trace-benchmark-dataset.textquestion-answeringn<1K0 likes38 downloads2mo agoHugging Face19pthinc /BCE-Prettybird-Nano-Math-v0.1 BCE-Prettybird-Nano-Math-v0.1 - 500 Math Q&A Dataset for Instruction-Based Learning We are excited to introduce a comprehensive math dataset containing 500 instruction-based question-answer pairs, designed to support research in mathematical reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced calculus, probability, and… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Math-v0.1.texttext-classificationn<1K0 likes37 downloads6mo agoHugging Face20liujin99 /nanochat-npu-stem-eval nanochat-npu-stem-eval Pre-processed STEM evaluation data for nanochat-npu, adapted from karpathy/nanochat for Huawei 910B3 NPU. Tasks Task Type Shot Source Examples Description gpqa_diamond multiple_choice 0-shot Idavidrein/gpqa 198 Graduate-level science QA (Diamond subset) gsm8k_cot generation 8-shot openai/gsm8k 1319 Grade school math word problems (CoT) math_cot generation 4-shot HuggingFaceH4/MATH-500 500 Competition mathematics (CoT)… See the full description on the dataset page: https://huggingface.co/datasets/liujin99/nanochat-npu-stem-eval.question-answering1K<n<10K0 likes35 downloads2mo agoHugging Face21GreenNode /nano-nq-vn NanoNQ-VN An MTEB dataset Massive Text Embedding Benchmark A translated dataset from NFCorpus: A Full-Text Learning to Rank Dataset for Medical Information Retrieval The process of creating the VN-MTEB (Vietnamese Massive Text Embedding Benchmark) from English samples involves a new automated system: - The system uses large language models (LLMs), specifically Coherence's Aya model, for translation. - Applies advanced embedding models to filter the translations. - Use LLM-as-a-judge… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/nano-nq-vn.texttext-retrieval100K<n<1M0 likes33 downloads9mo agoHugging Face22pthinc /BCE-Prettybird-Nano-Parrot-v0.2 BCE-Prettybird-Nano-Parrot-v0.2 - 700 Jokes for Instruction-Based Learning This dataset is a bilingual (Turkish-English mixed) comedic text collection designed for training and fine-tuning conversational AI models with humor awareness, sarcasm detection, and cultural nuance understanding. It includes short joke-style prompts, observational comedy snippets, and absurd dialogue fragments that blend everyday Turkish expressions with English punchlines, reflecting real-world… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Parrot-v0.2.texttext-classificationn<1K0 likes32 downloads4mo agoHugging Face23SolidSnake123 /nanochat-depo-retrieval-copy1-20260715 Nanochat Depo retrieval v1 Each latent 16-node graph yields eight independent, token-aligned, depth-one query documents. This arm exposes 1 nested edge(s) per document. Only the answer is supervised in every document; the terminal token is supervised only for query ordinal 7. This source is separate from and does not alter Depo-L0 v1. tabularquestion-answering10K<n<100K0 likes32 downloads2mo agoHugging Face24obekt /obekt-question-answer-reasoning-nano-v0.1 Obekt Nano Reasoning Dataset (v0.1) Dataset Description This is a small "nano" dataset containing questions, answers, and reasoning traces. It is generated using the Xiaomi MiMo V2 Flash LLM and is intended for experimental purposes, quick prototyping, and fine-tuning trials where reasoning capability is a focus. Source Model: xiaomi/mimo-v2-flash Contains obekt-question-answer-reasoning-nano-v0.1.csv: The main data file. Columns: question: The input query.… See the full description on the dataset page: https://huggingface.co/datasets/obekt/obekt-question-answer-reasoning-nano-v0.1.texttext-generation1K<n<10K0 likes30 downloads9mo agoHugging Face25pthinc /BCE-Prettybird-Nano-Science-v0.1 BCE-Prettybird-Nano-Science-v0.1 - 500 Science Q&A Dataset for Instruction-Based Learning We are excited to introduce a comprehensive math-physics-chemistry-biology dataset containing 500 instruction-based question-answer pairs, designed to support research in science reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Science-v0.1.texttext-classificationn<1K0 likes29 downloads6mo agoHugging Face26pthinc /BCE-Prettybird-Nano-Apollo-v0.1 BCE-Prettybird-Nano-Apollo-v0.1 Synthetic Multi-Language Software Engineering & UI/UX Dataset (1,070 Examples) This dataset contains 1,070 synthetic, high-quality examples covering a broad range of software engineering, architecture, database development, web design, UI/UX design, and design pattern implementations across multiple programming languages and frameworks. The collection includes: SOLID principle code examples in PHP, C#, Python, C++, Java, and JavaScript Design… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Apollo-v0.1.texttext-generation1K<n<10K0 likes29 downloads4mo agoHugging Face27fs90 /nano-start-data Nano-Start Learning Dataset A small educational dataset for learning how to train language models from scratch. Dataset Description This dataset contains simple, factual examples designed to demonstrate LLM training concepts: Completions: Factual statements the model learns to continue Q&A: Question-answer pairs using chat special tokens Chat: Multi-turn conversations with system prompts The dataset is intentionally small (~276 examples) so models can be trained quickly… See the full description on the dataset page: https://huggingface.co/datasets/fs90/nano-start-data.texttext-generationn<1K0 likes28 downloads10mo agoHugging Face28pthinc /BCE-Prettybird-Nano-Kangal-v0.1 BCE-Prettybird-Nano-Kangal-v0.1 - 525 LOVE Q&A Dataset for Instruction-Based Learning The "BCE-Prettybird-Nano-Kangal-v0.1: Love Dataset" consists of 525 rows of insightful data, offering a comprehensive exploration of romantic relationships. Covering diverse aspects from sexuality and intimacy to romance, family life management, and tips on how to treat women, this dataset delves into the complexities of modern relationships. It aims to provide valuable perspectives for those… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Kangal-v0.1.texttext-classificationn<1K0 likes26 downloads23d agoHugging Face29pthinc /BCE-Prettybird-Nano-Kayra-v0.1 BCE-Prettybird-Nano-Kayra-v0.1 - 200 AI Brain Mechanism Chat Kayra is an experimental 200-sample chat dataset developed by PROMETECH A.Ş. for research on Behavioral Consciousness Engine-style control systems. The dataset was synthetically generated using Nemotron Super and is designed to go beyond standard conversation data by exposing layered behavioral signals such as trust scoring, risk level, ethical guardrails, ego–superego balance, KPI tracking, cognitive-level analysis… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Nano-Kayra-v0.1.tabulartext-classificationn<1K0 likes26 downloads4mo agoHugging Face30SolidSnake123 /nanochat-depo-composition-depth2-w4-retry-20260715 Nanochat Depo composition v1 Each 16-node single-cycle graph yields eight independent one-query documents: four base starts paired across query depths (1, 2). This source contains train and validation splits only. Phase depth is 2; the materialized context width is 4. tabularquestion-answering10K<n<100K0 likes25 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.