datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ToolMind
ToolMind: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset
ToolMind is a large-scale, high-quality tool-agentic dataset with 160k synthetic data instances generated using over 20k tools and 200k augmented open-source data instances.
Our data synthesis pipeline first constructs a function graph based on parameter correlations and then uses a multi-agent framework to simulate realistic user–assistant–tool interactions.
Beyond trajectory-level validation, we employ fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/Nanbeige/ToolMind.natural_reasoningNaturalReasoning is a large-scale dataset for general reasoning tasks. It consists of high-quality challenging reasoning questions backtranslated from pretraining corpora DCLM and FineMath. The questions have been deduplicated and decontaminated from popular reasoning benchmarks including MATH, GPQA, MMLU-Pro, MMLU-STEM. For each question, we extract the reference final answer from the original document from the pretraining corpora if possible. We also provide a model-generated response from… See the full description on the dataset page: https://huggingface.co/datasets/facebook/natural_reasoning.pashto-emoji-dataset
Pashto Emoji Dataset
This dataset is a Pashto translation of the KomeijiForce/Text2Emoji dataset. It is designed for tasks involving the translation of text into emoji sequences and understanding the sentiment or topic of a given text.
The dataset contains over 504,000 rows, each consisting of a text passage in Pashto, a corresponding emoji sequence, and a topic label.
Dataset Structure
The dataset is provided in the following format:
text: A string containing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-emoji-dataset.NaturalConv
NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation
Introduction
This dataset is described in the paper NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation. The entire dataset contains 5 data files.
1. dialog_release.json:
It is a json file containing a list of dictionaries.
After loading in python this way:
import json
import codecs
dialog_list = json.loads(codecs.open("dialog_release.json"… See the full description on the dataset page: https://huggingface.co/datasets/xywang1/NaturalConv.nemotron-nano-30b-miniswe-swebench-verified
Nemotron Nano 30B + mini-swe-agent SWE-bench Verified Trajectories
Agent trajectories from running NVIDIA Nemotron 3 Nano 30B A3B (MoE, 8B active params) on SWE-bench Verified using mini-swe-agent.
⚠️ Incomplete Run
This benchmark was terminated early due to poor performance. The model struggled with the agentic coding task.
Model Information
Attribute
Value
Model
NVIDIA Nemotron 3 Nano 30B A3B
Architecture
MoE (30B total, 8B active)
Serving
vLLM… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/nemotron-nano-30b-miniswe-swebench-verified.DarijaDz
DarijaDZ
DarijaDZ is a large-scale corpus of user-generated text collected from Algerian YouTube and TikTok channels. The corpus contains approximately 22.5 million comment documents and 259.22 million word-level tokens, with content written primarily in Algerian Darija script alongside Latin/Arabizi writing and mixed-script content.
Dataset Description
Motivation
Algerian Darija is an under-resourced language variety with comparatively limited… See the full description on the dataset page: https://huggingface.co/datasets/nasrellahkharroubi/DarijaDz.narrative-function-calling-v1
Narrative Function Calling v1
Welcome to Narrative Function Calling v1! This dataset is purpose-built for training (or fine-tuning) models that produce consistent, structured function calls in conversation-like settings. The dataset integrates and normalizes data from both Glaive Function Calling v2 (Apache License 2.0) and Salesforce XLAM function calling data (CC-BY-4.0)[^liu2024apigen]. It provides a clean, rich, and comprehensive set of examples that guide large language models… See the full description on the dataset page: https://huggingface.co/datasets/narrative-io/narrative-function-calling-v1.natural_questions_cleanreasoning-math-advanced-1m
🧠 Reasoning Math Advanced 1M
📖 Dataset Summary
Reasoning Math Advanced 1M is a large-scale, synthetic dataset designed to enhance the reasoning capabilities of Large Language Models (LLMs). Comprising 1,000,000 unique samples, this dataset focuses on Math, Logic, and Common Sense reasoning tasks.
A unique feature of this dataset is its adaptive reasoning structure, where the presence of Chain-of-Thought (CoT) reasoning scales with difficulty. All reasoning traces are… See the full description on the dataset page: https://huggingface.co/datasets/naimulislam/reasoning-math-advanced-1m.NanoData
Dataset Description
To facilitate researchers to use NanoLM for comparative analysis across different model designs, we build a curated pre-training dataset from those of existing large-scale models (i.e., Llama, Falcon, GPT-3). It covers diverse domains to improve the generalization capabilities of the resultant models.
Dataset Creation
The data is mainly post-processed and filtered from RedPajama and RedPajamaV2.
We develop a series of cleaning steps to remove redundant… See the full description on the dataset page: https://huggingface.co/datasets/CofeAI/NanoData.SSR-RCoT-16K
SSR-RCoT-16K: Turning answer-only data into high-quality reasoning-supervision data
SSR-RCoT-16K is a public 16k subset derived from the data construction pipeline introduced in
Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation.
This dataset is designed for powerful reasoning on general tasks, especially for the realistic setting where high-quality responses are available but chain-of-thought annotations are missing.
In such answer-rich but… See the full description on the dataset page: https://huggingface.co/datasets/Nanbeige/SSR-RCoT-16K.Pashto-Free-Hand-Reasoning-Dataset
Pashto Free-Hand Reasoning SFT Dataset 🧠♻️
This dataset contains high-quality, long-form SFT (Supervised Fine-Tuning) conversational data in Pashto, featuring unconstrained, natural model reasoning (<think> blocks) paired with standardized chat responses.
🔄 The 3R Approach (Recycle, Reuse, Reason)
Instead of discarding legacy QA pairs, this dataset follows a 3R data philosophy:
Recycle: Taking older, simple, or raw legacy Pashto questions.
Reuse: Re-processing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/Pashto-Free-Hand-Reasoning-Dataset.brain-memory
🧠 NIFTY AI Agent: Memory OS Cloud Snapshot
Cloud backup repository for the NIFTY 50 Autonomous AI Agent Memory OS.
• Repository: nagarhimanshu37/brain-memory• Total Stored Records: 231• Last Synchronized: 2026-09-24 12:55:53 UTC
📊 Partition Statistics
Partition
Records
Description
conversation_memory
78
Multi-turn trader dialogues & intent logs
episodic_memory
50
Trading day episodes (facts vs interpretations)
experience_memory
50
Crystallized… See the full description on the dataset page: https://huggingface.co/datasets/nagarhimanshu37/brain-memory.NaiAD
NaiAD: Native Advertising in Dialogue Dataset
Dataset Summary
NaiAD is a specialized dataset designed for research in Native Advertising Insertion within conversational or long-form text generation. It explores how promotional content (ads) can be seamlessly integrated into user-requested content (queries' responses) using various strategies.
This dataset is being submitted as part of the research for NeurIPS 2026.
(Left) We define the task and the four decoupled… See the full description on the dataset page: https://huggingface.co/datasets/MaxAcand/NaiAD.AD-GEN
AD-GEN: Evidence-Preserving Generation of Validated ATT&CK-Aligned Narratives from Large-Scale Endpoint Telemetry
LLM-ready SOC narratives derived from large-scale Windows endpoint telemetry
Overview
Modern endpoint telemetry datasets contain rich behavioral evidence but are often difficult to use directly for large language model (LLM) reasoning. Raw Sysmon logs are fragmented across individual events, contain sensitive identifiers, suffer from process… See the full description on the dataset page: https://huggingface.co/datasets/namhop88/AD-GEN.parallel_ab-ru
Dataset Summary
The Abkhaz Russian parallel corpus dataset is a collection of 205,665 sentences/words extracted from different sources; e-books, web scrapping.
Dataset Creation
Source Data
Here is a link to the source on github
Considerations for Using the Data
Other Known Limitations
The accuracy of the dataset is around 95% (gramatical, arthographical errors)
cukurova_university_chatbot
Çukurova University Computer Engineering Chatbot Dataset
📊 Dataset Overview
This dataset contains 22,524 high-quality question-answer pairs specifically designed for training an AI chatbot that serves the Computer Engineering Department at Çukurova University. The dataset is part of the CengBot project, a sophisticated multilingual Telegram chatbot that provides automated assistance to students regarding courses, programs, and departmental information.
🔢… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/cukurova_university_chatbot.idfu-verified-code
IDFU Code Negative Dataset — Free Preview
A curated dataset of Python code samples that failed execution-based
validation, designed for training reward models, DPO rejected-side pairs,
and error-detection classifiers. Free 100-sample preview; paid full versions
available separately.
What's inside this preview
100 unique Python samples, all AST-validated
19 CS domains represented (MCMC, FFT, distributed consensus, ZKP,
formal methods, HFT microstructure, and more)… See the full description on the dataset page: https://huggingface.co/datasets/namakoo/idfu-verified-code.BFCL-V4-Parallel-Native
BFCL V4 Parallel Native
Native BFCL v4 single-turn parallel function-calling rows for decentralized multi-agent collaboration.
Source data comes from the official Berkeley Function Calling Leaderboard v4 data and possible-answer files.
Fields
id
official_category
task_type
user_prompt
function
ground_truth
Categories
live_parallel
live_parallel_multiple
parallel
parallel_multiple
Counts
train: 352 rows
eval: 88 rows
total: 440… See the full description on the dataset page: https://huggingface.co/datasets/OpenMLRL/BFCL-V4-Parallel-Native.drone-router-dataset-navlink-v2
NAVLINK Drone Router Dataset — Round 2 Snapshot
This dataset is the exact training/eval snapshot used for the best reviewed overnight FunctionGemma/NAVLINK run (“10-epoch run 2”), which reached:
Tool-call exact-match accuracy (line 1): 95.2%
476 / 500 correct
0 safety violations
This is the dataset snapshot before the later waypoint-copy-heavy augmentation that regressed performance.
Files
navlink_train_run2.jsonl — 4,307 training examples
navlink_test_run2.jsonl —… See the full description on the dataset page: https://huggingface.co/datasets/llama-farm/drone-router-dataset-navlink-v2.ctf-dataset
ctf-dataset
CTF 与网络安全知识的 ShareGPT/ChatML 风格 SFT 数据集,可用于 LoRA 微调。
数据格式
每行是一个 JSON 对象,核心字段如下:
字段
说明
id
样本唯一 ID
dataset
数据集名称,当前为 ctf-dataset
category
来源主题或 CTF/安全类别
ctf_task_type
任务类型标签
messages
ShareGPT 消息数组,包含 system / user / assistant
metadata
来源路径、章节、字符数、chunk 等溯源信息
LLaMA-Factory 接入
将 ctf-dataset.jsonl 放入 LLaMA-Factory 的 data/ 目录后,在 data/dataset_info.json 中添加:
{
"ctf_dataset": {
"file_name": "ctf-dataset.jsonl"… See the full description on the dataset page: https://huggingface.co/datasets/Nanhang/ctf-dataset.minecraft-question-answer-700k
minecraft-question-answer-700k
Introducing the largest synthetic Minecraft Q&A dataset, covering every topic, game mechanic, item and craft in Minecraft. The dataset was generated by extracting over 18,000 Minecraft wiki pages, and using glaive.ai's synthetic data generation pipeline.
about the dataset
rows - 694,814
tokens - 47,133,624
source - https://minecraft.wiki/
Hit me up on twitter if you see a bug or need a synthetic dataset for your company:… See the full description on the dataset page: https://huggingface.co/datasets/naklecha/minecraft-question-answer-700k.openpii-masking-nano-1k
OpenPII Nano: Multilingual PII Masking Sample
A nano-sized stratified sample of OpenPII 1.5M,
perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every
label that exists in the parent dataset is represented in proportion.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format
License
1,000
900
100
19
30
37
7… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-nano-1k.isotope-bench
Isotope Bench
An indirect-prompt-injection benchmark for tool-calling agents, plus the
complete audit trail of one recorded run: 438 influence certificates, one for
every action an agent attempted across five defence conditions.
Built for Isotope, which tracks
untrusted influence inside the forward pass. The corpus is independent of that
method and usable with any defence.
💻 Code: https://github.com/NagaYu/isotope
🤗 Demo: https://huggingface.co/spaces/NagaYu/isotope
🤗… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/isotope-bench.entity-native-agent-sessions
Entity-Native vs File-Native Agent Sessions on SWE-bench Verified
Full session logs from a controlled A/B experiment measuring how a coding agent's
retrieval substrate changes its behaviour, cost, and success rate on real
software-engineering tasks.
Both arms run the same model (Claude Sonnet 4.5), on the same tasks, from the
same repository state. The only difference is how the agent is allowed to find code.
Arm
Label
Tools available
A
file-native
Bash, Read, Grep… See the full description on the dataset page: https://huggingface.co/datasets/rs545837/entity-native-agent-sessions.Natural-Reasoning-gpt-oss-120B-S1
Dataset Card: Natural-Reasoning-gpt-oss-120B-S1
📜 Dataset Overview
This is a meticulously curated instruction fine-tuning dataset designed specifically for efficient knowledge distillation tasks. Built upon the first 100,000 questions from the large-scale reasoning corpus facebook/natural_reasoning (s1, I will process the remaining parts later), it aims to transfer the advanced, multi-step reasoning capabilities of the teacher model gpt-oss-120-high to a student model… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Natural-Reasoning-gpt-oss-120B-S1.InstructGpt-NaturalQa
LuminaSFT
LuminaSFT is a synthetic SFT dataset suite specifically designed to improve both general-purpose and task-specific SLMs. LuminaSFT consists of multiple curated splits that target diverse capabilities:
UltraChat200K-DeepSeek - A regenerated base SFT dataset for broad instruction following.
InstructGPT-NaturalQA and InstructGPT-TriviaQA - Factual question answering datasets to strengthen knowledge recall and answer accuracy.
CoT-Drop - A reading comprehension dataset with… See the full description on the dataset page: https://huggingface.co/datasets/amd/InstructGpt-NaturalQa.CodeGen-Diverse-5K
CodeGen-Diverse-5K: Broad Coverage for Competitive Programming
Part of the CodeGen suite | CodeGen-Deep-5K (sister dataset)
Dataset Description
CodeGen-Diverse-5K is a broad coverage dataset designed for training code generation models across a wide variety of competitive programming problems. This dataset prioritizes problem diversity over solution diversity, covering 5,000 unique problems with consistent, high-quality solutions.
Key Statistics
Total samples:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/CodeGen-Diverse-5K.Vietnamese-nampdn-ai-tiny-webtext-gg-translatedgithub-docs
GitHub Docs Corpus
A dataset containing only information from GitHub — the official
github/docs repository, i.e. the source of docs.github.com.
Dataset Structure
Files: data/train.jsonl
Format: JSONL, one chunk per line
Columns: text (cleaned doc chunk), metadata (source, title)
Rows: 3,336
Composition
Source: github/docs (main branch), content/ tree only — 3,734
Markdown files covering GitHub features, workflows, webhooks, REST/GraphQL
API docs… See the full description on the dataset page: https://huggingface.co/datasets/Nanthasit/github-docs.
