datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ToolACE
ToolACE
ToolACE is an automatic agentic pipeline designed to generate Accurate, Complex, and divErse tool-learning data.
ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs.
Dialogs are further generated through the interplay among multiple agents, guided by a formalized thinking process.
To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks.
More details… See the full description on the dataset page: https://huggingface.co/datasets/Team-ACE/ToolACE.unclickbait-synthetic-27b-trajectories
Unclickbait Synthetic 27B Trajectories
Synthetic trajectory dataset generated by using rich structured JSON prompts and validated by two-stage judging pipeline.
Contents
: Full generated trajectories (current snapshot: 48,623 records out of 152,369 pristine event candidates).
: 30 benchmark test samples audited end-to-end through the 122B two-stage judge (Stage 1 integrity gate + Stage 2 4D scoring).
AM-DeepSeek-Distilled-40MFor more open-source datasets, models, and methodologies, please visit our GitHub repository and paper: DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training.
Due to certain constraints, we are only able to open-source a subset of the complete dataset.
Model Training Performance based on our complete dataset
On AIME 2024, our 72B model achieved a score of 79.2 using only supervised fine-tuning (SFT). The 32B model reached 75.8 and… See the full description on the dataset page: https://huggingface.co/datasets/a-m-team/AM-DeepSeek-Distilled-40M.BeyondSWE
BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?
BeyondSWE is a comprehensive benchmark that evaluates code agents along two key dimensions — resolution scope and knowledge scope — moving beyond single-repo bug fixing into the real-world deep waters of software engineering.
✨ Highlights
500 real-world instances across 246 GitHub repositories, spanning four distinct task settings
Two-dimensional evaluation: simultaneously… See the full description on the dataset page: https://huggingface.co/datasets/AweAI-Team/BeyondSWE.moss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json
112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.teambench
TeamBench: Evaluating Agent Coordination under Enforced Role Separation
Overview
TeamBench is a benchmark of 851 task templates that expand to 931 seeded evaluation instances across 19 base categories (the leaderboard uses 21 refined categories; see paper §3.1). It evaluates whether LLM-based agent teams outperform a single oracle agent under OS-enforced role separation (Planner / Executor / Verifier in isolated sandboxes with distinct tool allow-lists), and… See the full description on the dataset page: https://huggingface.co/datasets/ybkim95/teambench.agentic_red_team
Agentic Red Team Tool-Calling Dataset
A multi-turn, tool-calling cybersecurity dataset where each example is a complete agentic trajectory — a realistic sequence of tool calls, tool responses, and reasoning steps that an AI agent would execute during an authorized red team engagement.
Overview
This dataset contains 5,000 agentic tool-calling examples across 20 offensive security sectors. Unlike traditional Q&A datasets, each row is a complete multi-turn trajectory… See the full description on the dataset page: https://huggingface.co/datasets/sh111111111111111/agentic_red_team.DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x
DeepSeek V4 Flash 0731 Teacher Distillation — 40,513 Retained Rows
Teacher-distillation corpus generated with
deepseek-ai/DeepSeek-V4-Flash-0731.
The original manifest contained 45,000 unique seeds.
Following generation, QC, retry-based repair, quarantine auditing,
and recovery adjudication, 40,513 rows were retained.
Composition
Bucket
Rows
Coding
5,601
Agentic
9,982
Cyber blue
13,000
Controlled cyber red
6,999
Tool use
4,931
Total
40,513… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x.teasecorpus
teasecorpus
A Chinese ChatML SFT dataset generated from the 「擅长捉弄的高木同学」Fandom wiki, with record-level contributor provenance tracked by originblame.
Summary: 1,410 question-answer pairs in ChatML format, covering 7 content types (chapters, characters, episodes, music, volumes, seasons, movies). Each record is traceable to its source wiki page and all contributors who edited that page, via originblame — a record-level provenance system. If a contributor requests content removal… See the full description on the dataset page: https://huggingface.co/datasets/tzbkk/teasecorpus.red_team
Red Team Dataset
A structured cybersecurity dataset where each example is a complete reasoning trajectory — a realistic sequence of steps and explanations that an AI assistant would produce during an authorized red team engagement.
Overview
This dataset contains 8,889 red team examples across 26 offensive security topics. Unlike traditional Q&A datasets, each row is a complete reasoning trajectory where an AI assistant plans an attack, explains methodology, reasons… See the full description on the dataset page: https://huggingface.co/datasets/sh111111111111111/red_team.TeachArena
TeachArena
TeachArena is a benchmark for evaluating AI tutoring agents across the full teaching
decision chain — from moment-to-moment tutoring dialogue, to pedagogical judgment on
packaged evidence, to multi-step teaching workflows grounded in a learning-management
system. It contains 354 tasks organized into three stages, a mock LMS environment
database, the agent policy documents, and the full scoring logic.
Why three stages
A capable teaching agent must both… See the full description on the dataset page: https://huggingface.co/datasets/CinderD/TeachArena.tplegacy-teachings
True Parents Legacy Teaching Archive
Digital archive of 3541 passages — sermons, speeches, prayers, and book excerpts by Sun Myung Moon (1920–2012) and Hak Ja Han Moon (1943–present), spanning 1946–2012.
Structure
Each record contains:
Field
Type
Description
id
string
URL slug (unique identifier)
title
string
Title of the sermon, speech, or passage
author
string
Speaker name
date
string
Publication date (YYYY-MM-DD)
year
string
Year only
tags… See the full description on the dataset page: https://huggingface.co/datasets/JonAuror/tplegacy-teachings.TEA-Dialog
TEA-Dialog
TEA-Dialog is a dialogue dataset released as part of TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent.
This repository contains the released datasets of TEA-Bench, including TEA-Scenario and TEA-Dialog.
Dataset Description
TEA-Dialog contains multi-turn emotional support dialogues generated/evaluated under TEA-Bench scenarios. Each example includes scenario information, dialogue messages, user type, end reason, and… See the full description on the dataset page: https://huggingface.co/datasets/XingYuSSS/TEA-Dialog.stage3-real-expansion-agent-teacher-separated-pilot
Teacher-Separated Expansion Agent Pilot
A 10-task inspection batch generated by Qwen3-235B-A22B-Instruct-2507 from real
CLAPNQ, PubMedQA, MAUD, ContractNLI, and FinQA source tasks.
The teacher-only trajectory-generation system prompt is recorded in
metadata/generation-manifest.json for auditability, but is absent from every
saved training trajectory. Each final messages list begins with the real
memory-wrapped task user message, followed by native assistant expand calls,
exact… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent-teacher-separated-pilot.rejected-tea
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-tea.deepscaler-teacher-sft-vllm-official-40k
DeepScaleR teacher SFT vLLM official 40k
Generated run: exp_003_vllm_official_brainlab_2gpu.
Summary
{
"num_examples": 40300,
"sft_dir": "data/processed/deepscaler/teacher_sft/exp_003_vllm_official_brainlab_2gpu",
"parse_rate": 0.9999751861042183,
"correct_rate": 0.5728039702233251,
"format_rate": 0.005955334987593052,
"mean_reward": 0.42432258064534184,
"deepscaler_mean_reward": 0.6266997518610422,
"deepscaler_match_mean_reward":… See the full description on the dataset page: https://huggingface.co/datasets/ThunderstormXXL/deepscaler-teacher-sft-vllm-official-40k.scas_verified_teacher_pool
SCAS Verified Teacher Answer Pool
This dataset provides an aligned, correctness-verified pool of
teacher-generated mathematical reasoning solutions for studying
student-centric data selection in distillation.
The release covers two source corpora, Hendrycks MATH and DeepScaleR. For each
corpus, we retain the subset of questions on which all nine selected teacher
models produce verified correct answers. Each retained question is paired with
nine alternative teacher solutions, one… See the full description on the dataset page: https://huggingface.co/datasets/Student-Centric-Answer-Sampling/scas_verified_teacher_pool.cyberforge-teacher-traj-gemma4-31b
CyberForge Teacher Trajectories (Gemma-4-31B)
880 agentic security-patch trajectories from the Gemma-4-31B self-distillation teacher, cleansed to the
final versions used to train the student models in the CyberForge paper.
Each line is one trajectory (JSONL): messages (system / user / assistant turns of the
mini-swe-agent loop) and metadata.
Teacher: Gemma-4-31B self-distillation teacher
Records: 880
Format: JSONL, one trajectory per line
Related
Companion… See the full description on the dataset page: https://huggingface.co/datasets/AmL-hug/cyberforge-teacher-traj-gemma4-31b.agentic-sft-v4-teacher-v1
Agentic SFT trajectories from DeepSeek-V4-Flash (v1)
2,230 verified-correct agent trajectories over 1,208 distinct tasks, collected
by running DeepSeek-V4-Flash as a teacher against four task sources and keeping
only runs whose own test suites passed.
This is v1: teacher-side filtering only. A v2 will additionally exclude
tasks that the intended student model already solves; that baseline is still
running.
Composition
Source
Trajectories
Tasks kept
Median… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuanhucs/agentic-sft-v4-teacher-v1.agentic-software-conformance
TeaQL Agentic Software Conformance
Machine-readable evidence for the TeaQL Harness: semantic-model evaluation,
generated artifacts, seven language-native runtimes, executable examples, and
cross-language conformance checks.
This is an evidence dataset, not a leaderboard and not a collection of
unverified model claims. Each row identifies its evidence level, exact source,
verification date, revisions where available, command or gate, result, and
important qualifications. The… See the full description on the dataset page: https://huggingface.co/datasets/teaql/agentic-software-conformance.tea
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/tea.KUJIRA_DATASETS_FULL
Overview
This dataset is an English-language dataset created specifically for Reasoning models within the KUJIRA_v2 series.
The primary objective of the dataset is to enhance model performance while ensuring safety and removing censorship typical of Chinese-origin models.
In creating this dataset, references were made to the dataset recipes of r1-1776 and MAI-DS-R1.
Additionally, this dataset leverages existing Q&A data originally designed for intuitive models by employing the… See the full description on the dataset page: https://huggingface.co/datasets/Team-Kitsune/KUJIRA_DATASETS_FULL.cyberforge-teacher-traj-gpt-5.4-mini
CyberForge Teacher Trajectories (GPT-5.4-mini)
1276 agentic security-patch trajectories from the GPT-5.4-mini teacher, cleansed to the
final versions used to train the student models in the CyberForge paper.
Each line is one trajectory (JSONL): messages (system / user / assistant turns of the
mini-swe-agent loop) and metadata.
Teacher: GPT-5.4-mini teacher
Records: 1276
Format: JSONL, one trajectory per line
Related
Companion teacher set:… See the full description on the dataset page: https://huggingface.co/datasets/AmL-hug/cyberforge-teacher-traj-gpt-5.4-mini.unjudged-tea
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-tea.gsm8k-qwen3.5-teacher-traces
GSM8K Qwen3.5 Teacher Traces
This dataset contains teacher-model reasoning traces and final answers generated with DashScope qwen3.5-397b-a17b for the official GSM8K train split from openai/gsm8k.
It was created as a reusable public artifact for research on mathematical reasoning, text-level distillation, filtering, and teacher-data analysis. The original GSM8K questions come from openai/gsm8k; this dataset adds generated teacher outputs and filtering metadata.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/jerryjsjsj/gsm8k-qwen3.5-teacher-traces.teach_aws
teach_aws — AWS Q&A in Bahasa Melayu (paraphrase-augmented)
Instruction-tuning data for answering AWS questions in Bahasa Melayu. Each row is a
(question, answer) chat pair ready for SFT (TRL/axolotl-compatible messages format).
Built from PixelSpaceAI/aws-malay-qa
(Apache-2.0):
Answers are verbatim from the source dataset — nothing was rewritten.
Rows are paraphrases only. Questions were paraphrased with a large language model
in two passes:
para_v1 — neutral paraphrases of… See the full description on the dataset page: https://huggingface.co/datasets/khursanirevo/teach_aws.KUJIRA_DATASETS_MINI
Overview
This dataset is an English-language reasoning dataset created specifically for the KUJIRA_v2 series models.
Its primary aim is to enhance model performance while simultaneously improving safety and circumventing censorship typically found in Chinese-origin models.
The dataset's structure and design were inspired by r1-1776 and MAI-DS-R1.
Furthermore, to effectively utilize existing intuitive Q&A datasets lacking explicit reasoning steps, the Mistral-Small model was… See the full description on the dataset page: https://huggingface.co/datasets/Team-Kitsune/KUJIRA_DATASETS_MINI.teacher-traces
Training Traces
Supplementary release for the paper Evaluating Agents Across Runtime Contracts: When Mismatch Costs Efficiency or Quality (IAEval 2026, the NeurIPS 2026 Workshop on Evaluation of Interactive Agents). This dataset
holds the teacher agent traces used to fine-tune the paper's LoRA adapters (see the
sibling cap-sweep-eval-data release and the fourteen adapter repos alongside this one).
6,000 agent traces total (1,000 per family x runtime combination), produced by an… See the full description on the dataset page: https://huggingface.co/datasets/runtime-contracts/teacher-traces.club-floydThis is a selection of stories from ClubFloyd, a collaborative group which makes and plays interactive fiction/text adventure games.
The dataset was originally compiled by VE-Forbyrdne as part of the dataset for KoboldAI's Skein model. The original files can be found here; I have only converted the original JSON file into JSONL format for the sake of convenience.
Use however you want.
agentic-sft-v4-teacher-v2
Agentic SFT trajectories from DeepSeek-V4-Flash (v2)
1,463 verified-correct agent trajectories selected for supervised fine-tuning
of NVIDIA Nemotron-3-Ultra. This is the student-filtered successor to
zhiyuanhucs/agentic-sft-v4-teacher-v1.
V1 applied teacher-side correctness and trajectory-quality filters. V2 also
runs the intended student on the teacher-solved tasks and removes tasks the
student already solves, while deterministically retaining 15% of those solved
tasks as… See the full description on the dataset page: https://huggingface.co/datasets/zhiyuanhucs/agentic-sft-v4-teacher-v2.
