CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Nanbeige /ToolMind ToolMind: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset ToolMind is a large-scale, high-quality tool-agentic dataset with 160k synthetic data instances generated using over 20k tools and 200k augmented open-source data instances. Our data synthesis pipeline first constructs a function graph based on parameter correlations and then uses a multi-agent framework to simulate realistic user–assistant–tool interactions. Beyond trajectory-level validation, we employ fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/Nanbeige/ToolMind.documenttext-generation100K<n<1M173 likes3.4k downloads9mo agoHugging Face02facebook /natural_reasoningNaturalReasoning is a large-scale dataset for general reasoning tasks. It consists of high-quality challenging reasoning questions backtranslated from pretraining corpora DCLM and FineMath. The questions have been deduplicated and decontaminated from popular reasoning benchmarks including MATH, GPQA, MMLU-Pro, MMLU-STEM. For each question, we extract the reference final answer from the original document from the pretraining corpora if possible. We also provide a model-generated response from… See the full description on the dataset page: https://huggingface.co/datasets/facebook/natural_reasoning.texttext-generation1M<n<10M585 likes2.7k downloads2y agoHugging Face03nassimjp /pashto-emoji-dataset Pashto Emoji Dataset This dataset is a Pashto translation of the KomeijiForce/Text2Emoji dataset. It is designed for tasks involving the translation of text into emoji sequences and understanding the sentiment or topic of a given text. The dataset contains over 504,000 rows, each consisting of a text passage in Pashto, a corresponding emoji sequence, and a topic label. Dataset Structure The dataset is provided in the following format: text: A string containing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-emoji-dataset.texttext-to-image100K<n<1M0 likes494 downloads26d agoHugging Face04xywang1 /NaturalConv NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation Introduction This dataset is described in the paper NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation. The entire dataset contains 5 data files. 1. dialog_release.json: It is a json file containing a list of dictionaries. After loading in python this way: import json import codecs dialog_list = json.loads(codecs.open("dialog_release.json"… See the full description on the dataset page: https://huggingface.co/datasets/xywang1/NaturalConv.texttext-generation10K<n<100K22 likes447 downloads2y agoHugging Face05pankajmathur /nemotron-nano-30b-miniswe-swebench-verified Nemotron Nano 30B + mini-swe-agent SWE-bench Verified Trajectories Agent trajectories from running NVIDIA Nemotron 3 Nano 30B A3B (MoE, 8B active params) on SWE-bench Verified using mini-swe-agent. ⚠️ Incomplete Run This benchmark was terminated early due to poor performance. The model struggled with the agentic coding task. Model Information Attribute Value Model NVIDIA Nemotron 3 Nano 30B A3B Architecture MoE (30B total, 8B active) Serving vLLM… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/nemotron-nano-30b-miniswe-swebench-verified.texttext-generationn<1K0 likes354 downloads9mo agoHugging Face06nasrellahkharroubi /DarijaDz DarijaDZ DarijaDZ is a large-scale corpus of user-generated text collected from Algerian YouTube and TikTok channels. The corpus contains approximately 22.5 million comment documents and 259.22 million word-level tokens, with content written primarily in Algerian Darija script alongside Latin/Arabizi writing and mixed-script content. Dataset Description Motivation Algerian Darija is an under-resourced language variety with comparatively limited… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDz.texttext-generation10M<n<100M10 likes324 downloads10d agoHugging Face07narrative-io /narrative-function-calling-v1 Narrative Function Calling v1 Welcome to Narrative Function Calling v1! This dataset is purpose-built for training (or fine-tuning) models that produce consistent, structured function calls in conversation-like settings. The dataset integrates and normalizes data from both Glaive Function Calling v2 (Apache License 2.0) and Salesforce XLAM function calling data (CC-BY-4.0)[^liu2024apigen]. It provides a clean, rich, and comprehensive set of examples that guide large language models… See the full description on the dataset page: https://huggingface.co/datasets/narrative-io/narrative-function-calling-v1.texttext-generation100K<n<1M0 likes315 downloads2y agoHugging Face08rojagtap /natural_questions_cleantextquestion-answering100K<n<1M10 likes300 downloads3y agoHugging Face09naimulislam /reasoning-math-advanced-1m 🧠 Reasoning Math Advanced 1M 📖 Dataset Summary Reasoning Math Advanced 1M is a large-scale, synthetic dataset designed to enhance the reasoning capabilities of Large Language Models (LLMs). Comprising 1,000,000 unique samples, this dataset focuses on Math, Logic, and Common Sense reasoning tasks. A unique feature of this dataset is its adaptive reasoning structure, where the presence of Chain-of-Thought (CoT) reasoning scales with difficulty. All reasoning traces are… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/reasoning-math-advanced-1m.texttext-generation1M<n<10M0 likes276 downloads9mo agoHugging Face10CofeAI /NanoData Dataset Description To facilitate researchers to use NanoLM for comparative analysis across different model designs, we build a curated pre-training dataset from those of existing large-scale models (i.e., Llama, Falcon, GPT-3). It covers diverse domains to improve the generalization capabilities of the resultant models. Dataset Creation The data is mainly post-processed and filtered from RedPajama and RedPajamaV2. We develop a series of cleaning steps to remove redundant… See the full description on the dataset page: https://huggingface.co/datasets/CofeAI/NanoData.texttext-generation1M<n<10M3 likes246 downloads2y agoHugging Face11Nanbeige /SSR-RCoT-16K SSR-RCoT-16K: Turning answer-only data into high-quality reasoning-supervision data SSR-RCoT-16K is a public 16k subset derived from the data construction pipeline introduced in Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation. This dataset is designed for powerful reasoning on general tasks, especially for the realistic setting where high-quality responses are available but chain-of-thought annotations are missing. In such answer-rich but… See the full description on the dataset page: https://huggingface.co/datasets/Nanbeige/SSR-RCoT-16K.documenttext-generation10K<n<100K16 likes245 downloads6mo agoHugging Face12nassimjp /Pashto-Free-Hand-Reasoning-Dataset Pashto Free-Hand Reasoning SFT Dataset 🧠♻️ This dataset contains high-quality, long-form SFT (Supervised Fine-Tuning) conversational data in Pashto, featuring unconstrained, natural model reasoning (<think> blocks) paired with standardized chat responses. 🔄 The 3R Approach (Recycle, Reuse, Reason) Instead of discarding legacy QA pairs, this dataset follows a 3R data philosophy: Recycle: Taking older, simple, or raw legacy Pashto questions. Reuse: Re-processing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Free-Hand-Reasoning-Dataset.texttext-generation1K<n<10K0 likes215 downloads8d agoHugging Face13nagarhimanshu37 /brain-memory 🧠 NIFTY AI Agent: Memory OS Cloud Snapshot Cloud backup repository for the NIFTY 50 Autonomous AI Agent Memory OS. • Repository: nagarhimanshu37/brain-memory• Total Stored Records: 231• Last Synchronized: 2026-09-24 12:55:53 UTC 📊 Partition Statistics Partition Records Description conversation_memory 78 Multi-turn trader dialogues & intent logs episodic_memory 50 Trading day episodes (facts vs interpretations) experience_memory 50 Crystallized… See the full description on the dataset page: https://huggingface.co/datasets/nagarhimanshu37/brain-memory.texttext-generationn<1K0 likes172 downloads9h agoHugging Face14MaxAcand /NaiAD NaiAD: Native Advertising in Dialogue Dataset Dataset Summary NaiAD is a specialized dataset designed for research in Native Advertising Insertion within conversational or long-form text generation. It explores how promotional content (ads) can be seamlessly integrated into user-requested content (queries' responses) using various strategies. This dataset is being submitted as part of the research for NeurIPS 2026. (Left) We define the task and the four decoupled… See the full description on the dataset page: https://huggingface.co/datasets/MaxAcand/NaiAD.texttext-generation10K<n<100K3 likes155 downloads5mo agoHugging Face15namhop88 /AD-GEN AD-GEN: Evidence-Preserving Generation of Validated ATT&CK-Aligned Narratives from Large-Scale Endpoint Telemetry LLM-ready SOC narratives derived from large-scale Windows endpoint telemetry Overview Modern endpoint telemetry datasets contain rich behavioral evidence but are often difficult to use directly for large language model (LLM) reasoning. Raw Sysmon logs are fragmented across individual events, contain sensitive identifiers, suffer from process… See the full description on the dataset page: https://huggingface.co/datasets/namhop88/AD-GEN.texttext-classification100K<n<1M2 likes145 downloads4mo agoHugging Face16Nart /parallel_ab-ru Dataset Summary The Abkhaz Russian parallel corpus dataset is a collection of 205,665 sentences/words extracted from different sources; e-books, web scrapping. Dataset Creation Source Data Here is a link to the source on github Considerations for Using the Data Other Known Limitations The accuracy of the dataset is around 95% (gramatical, arthographical errors) texttext-generationn<1K1 likes144 downloads2y agoHugging Face17Naholav /cukurova_university_chatbot Çukurova University Computer Engineering Chatbot Dataset 📊 Dataset Overview This dataset contains 22,524 high-quality question-answer pairs specifically designed for training an AI chatbot that serves the Computer Engineering Department at Çukurova University. The dataset is part of the CengBot project, a sophisticated multilingual Telegram chatbot that provides automated assistance to students regarding courses, programs, and departmental information. 🔢… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/cukurova_university_chatbot.textquestion-answering10K<n<100K0 likes136 downloads1y agoHugging Face18namakoo /idfu-verified-code IDFU Code Negative Dataset — Free Preview A curated dataset of Python code samples that failed execution-based validation, designed for training reward models, DPO rejected-side pairs, and error-detection classifiers. Free 100-sample preview; paid full versions available separately. What's inside this preview 100 unique Python samples, all AST-validated 19 CS domains represented (MCMC, FFT, distributed consensus, ZKP, formal methods, HFT microstructure, and more)… See the full description on the dataset page: https://huggingface.co/datasets/namakoo/idfu-verified-code.texttext-generationn<1K1 likes136 downloads5mo agoHugging Face19OpenMLRL /BFCL-V4-Parallel-Native BFCL V4 Parallel Native Native BFCL v4 single-turn parallel function-calling rows for decentralized multi-agent collaboration. Source data comes from the official Berkeley Function Calling Leaderboard v4 data and possible-answer files. Fields id official_category task_type user_prompt function ground_truth Categories live_parallel live_parallel_multiple parallel parallel_multiple Counts train: 352 rows eval: 88 rows total: 440… See the full description on the dataset page: https://huggingface.co/datasets/OpenMLRL/BFCL-V4-Parallel-Native.texttext-generationn<1K1 likes131 downloads3mo agoHugging Face20llama-farm /drone-router-dataset-navlink-v2 NAVLINK Drone Router Dataset — Round 2 Snapshot This dataset is the exact training/eval snapshot used for the best reviewed overnight FunctionGemma/NAVLINK run (“10-epoch run 2”), which reached: Tool-call exact-match accuracy (line 1): 95.2% 476 / 500 correct 0 safety violations This is the dataset snapshot before the later waypoint-copy-heavy augmentation that regressed performance. Files navlink_train_run2.jsonl — 4,307 training examples navlink_test_run2.jsonl —… See the full description on the dataset page: https://huggingface.co/datasets/llama-farm/drone-router-dataset-navlink-v2.texttext-generation1K<n<10K1 likes126 downloads6mo agoHugging Face21Nanhang /ctf-dataset ctf-dataset CTF 与网络安全知识的 ShareGPT/ChatML 风格 SFT 数据集,可用于 LoRA 微调。 数据格式 每行是一个 JSON 对象,核心字段如下: 字段 说明 id 样本唯一 ID dataset 数据集名称,当前为 ctf-dataset category 来源主题或 CTF/安全类别 ctf_task_type 任务类型标签 messages ShareGPT 消息数组,包含 system / user / assistant metadata 来源路径、章节、字符数、chunk 等溯源信息 LLaMA-Factory 接入 将 ctf-dataset.jsonl 放入 LLaMA-Factory 的 data/ 目录后,在 data/dataset_info.json 中添加: { "ctf_dataset": { "file_name": "ctf-dataset.jsonl"… See the full description on the dataset page: https://huggingface.co/datasets/Nanhang/ctf-dataset.texttext-generation1K<n<10K0 likes115 downloads2mo agoHugging Face22naklecha /minecraft-question-answer-700k minecraft-question-answer-700k Introducing the largest synthetic Minecraft Q&A dataset, covering every topic, game mechanic, item and craft in Minecraft. The dataset was generated by extracting over 18,000 Minecraft wiki pages, and using glaive.ai's synthetic data generation pipeline. about the dataset rows - 694,814 tokens - 47,133,624 source - https://minecraft.wiki/ Hit me up on twitter if you see a bug or need a synthetic dataset for your company:… See the full description on the dataset page: https://huggingface.co/datasets/naklecha/minecraft-question-answer-700k.textquestion-answering100K<n<1M46 likes113 downloads2y agoHugging Face23ai4privacy /openpii-masking-nano-1k OpenPII Nano: Multilingual PII Masking Sample A nano-sized stratified sample of OpenPII 1.5M, perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every label that exists in the parent dataset is represented in proportion. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format License 1,000 900 100 19 30 37 7… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-nano-1k.texttoken-classification1K<n<10K3 likes111 downloads4mo agoHugging Face24NagaYu /isotope-bench Isotope Bench An indirect-prompt-injection benchmark for tool-calling agents, plus the complete audit trail of one recorded run: 438 influence certificates, one for every action an agent attempted across five defence conditions. Built for Isotope, which tracks untrusted influence inside the forward pass. The corpus is independent of that method and usable with any defence. 💻 Code: https://github.com/NagaYu/isotope 🤗 Demo: https://huggingface.co/spaces/NagaYu/isotope 🤗… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/isotope-bench.tabulartext-generationn<1K1 likes107 downloads19d agoHugging Face25rs545837 /entity-native-agent-sessions Entity-Native vs File-Native Agent Sessions on SWE-bench Verified Full session logs from a controlled A/B experiment measuring how a coding agent's retrieval substrate changes its behaviour, cost, and success rate on real software-engineering tasks. Both arms run the same model (Claude Sonnet 4.5), on the same tasks, from the same repository state. The only difference is how the agent is allowed to find code. Arm Label Tools available A file-native Bash, Read, Grep… See the full description on the dataset page: https://huggingface.co/datasets/rs545837/entity-native-agent-sessions.tabulartext-generationn<1K0 likes106 downloads23d agoHugging Face26Jackrong /Natural-Reasoning-gpt-oss-120B-S1 Dataset Card: Natural-Reasoning-gpt-oss-120B-S1 📜 Dataset Overview This is a meticulously curated instruction fine-tuning dataset designed specifically for efficient knowledge distillation tasks. Built upon the first 100,000 questions from the large-scale reasoning corpus facebook/natural_reasoning (s1, I will process the remaining parts later), it aims to transfer the advanced, multi-step reasoning capabilities of the teacher model gpt-oss-120-high to a student model… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Natural-Reasoning-gpt-oss-120B-S1.texttext-generation10K<n<100K24 likes104 downloads1y agoHugging Face27amd /InstructGpt-NaturalQa LuminaSFT LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities: UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following. InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy. CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-NaturalQa.texttext-generation1M<n<10M1 likes104 downloads7mo agoHugging Face28Naholav /CodeGen-Diverse-5K CodeGen-Diverse-5K: Broad Coverage for Competitive Programming Part of the CodeGen suite | CodeGen-Deep-5K (sister dataset) Dataset Description CodeGen-Diverse-5K is a broad coverage dataset designed for training code generation models across a wide variety of competitive programming problems. This dataset prioritizes problem diversity over solution diversity, covering 5,000 unique problems with consistent, high-quality solutions. Key Statistics Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Diverse-5K.tabulartext-generation1K<n<10K0 likes100 downloads10mo agoHugging Face295CD-AI /Vietnamese-nampdn-ai-tiny-webtext-gg-translatedtextquestion-answering1M<n<10M10 likes99 downloads3y agoHugging Face30Nanthasit /github-docs GitHub Docs Corpus A dataset containing only information from GitHub — the official github/docs repository, i.e. the source of docs.github.com. Dataset Structure Files: data/train.jsonl Format: JSONL, one chunk per line Columns: text (cleaned doc chunk), metadata (source, title) Rows: 3,336 Composition Source: github/docs (main branch), content/ tree only — 3,734 Markdown files covering GitHub features, workflows, webhooks, REST/GraphQL API docs… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/github-docs.texttext-generation1K<n<10K0 likes99 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.