datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UltraData-SFT-Agent-2609
UltraData-SFT-Agent-2609
📦 UltraData Collection |
🌐 UltraData |
🤗 MiniCPM5 Series
English |
中文
📚 Introduction
UltraData-SFT-Agent-2609 is the L3 refined data for Agent instruction-tuning within UltraData's L0-L4 tiered data management framework. Built for the post-training of MiniCPM5-2B, it complements UltraData-SFT-2605 (core-domain SFT) with executable Agent trajectories. The release contains approximately 500,000 samples spanning tool use… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/UltraData-SFT-Agent-2609.TIR-Bench
TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
Introduction:
TIR-Bench is a comprehensive benchmark designed to evaluate the "thinking-with-images" capabilities of Multimodal Large Language Models (MLLMs), addressing a gap left by existing benchmarks like Visual Search which only test basic operations. As models like OpenAI o3 begin to intelligently create and operate tools to transform images for problem-solving, TIR-Bench provides 13… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/TIR-Bench.DEBATE
DEBATE: Diverse Multi-Agent Debates
This dataset is presented in the paper "MALLM: Multi-Agent Large Language Models Framework".
Citation
comming soon.
text-sft-questions-answers-only
text-sft: Questions and Answers
This dataset consists of question-and-answer pairs generated from short excerpts drawn from Wikipedia, Cosmopedia, and FineWeb-Edu. It is an adapted version of agentlans/text-sft.
Overview
The dataset provides compact examples of English question-and-answer relationships that can help models learn linguistic patterns, syntactic structures, and semantic associations between questions and their corresponding answers.
Intended Use… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/text-sft-questions-answers-only.agents-index
AgentCrush Agent Index
Evidence-ranked index of the AI agent economy. Updated daily from agentcrush.xyz.
Overview
1,443 agents indexed across categories: developer tools, tokenized agents, service agents, model families
207 evidence-ranked with verified multi-signal scores
Updated: 2026-09-24
Configs
Config
Description
Rows
agents
All indexed agents with metadata
~1,443
evidence_ranked
Evidence-ranked tier only
~207
snapshots_latest… See the full description on the dataset page: https://huggingface.co/datasets/AgentCrush/agents-index.AgentJudgeBench
AgentJudgeBench: Evaluating LLM Judge Reliability on Agentic Tool-Calling
A benchmark for systematically evaluating how reliably LLM judges assess
agentic tool-calling workflows across structured, dependency-driven tasks.
Why this benchmark?
AgentJudgeBench measures how reliably LLM judges assess agentic tool-calling outputs. It provides 3,808 benchmark records spanning six DAG topologies and three difficulty… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/AgentJudgeBench.orca
Orca Dataset Collection
The Orca Dataset Collection is a unified compilation of multiple datasets from the Microsoft Orca and OpenOrca projects.
All duplicate entries have been removed, and any personally identifiable information (PII) has been carefully redacted.
All rows have been sorted by hash and split into JSONL files containing 100,000 entries each for easier handling and consistency.
Example Entry
id: MD5 hash of the system prompt, question, and answer JSON… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/orca.agentmujo-function-calling
agentmujo-function-calling (v0.1.0 — 1252 uzorka)
Ručno dizajniran kanonski skup za function calling na bosanskom jeziku
(ijekavica), dio AgentMujo Training Frameworka
(configs/tools.yaml je Single Source of Truth za alate).
Verzija: 0.1.0 · Uzoraka: 1252 (split 983/128/141) · Jezik: bs-ijekavica
Format: JSONL; svaki red: id, version, language, task, difficulty, enable_thinking, messages[] (user/assistant/tool + tool_calls[]), metadata{}
(schema: schemas/dataset.schema.json u… See the full description on the dataset page: https://huggingface.co/datasets/shaban2024/agentmujo-function-calling.lightblue-tagengo-gpt4
lightblue/tagengo-gpt4
An unofficial, reformatted version of lightblue/tagengo-gpt4.
Tagengo is described by its author as the world's largest high-quality multilingual chat dataset - containing over 75,000 single-turn conversations between humans and GPT‑4 (gpt-4-0125-preview) across 74 languages. It fills a major gap in multilingual chat data, which has so far been limited compared to English.
Additional Processing
Split by language
Kept only entries with both… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/lightblue-tagengo-gpt4.agentvidbench
AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
Agentic Video Understanding Benchmark — 100 multiple-choice video QA questions, 26 options each (A-Z; ~3.8% random baseline)
by Seoyeon An*, Hyeonseo Jang*, Minsu Kim*, Chanho Lee, Younghan Park, Kangwook Lee (KRAFTON AI)
Layout
.
├── README.md
├── questions.jsonl # 100 rows — one per question
├── videos.jsonl # 71 rows — one per unique video
├── videos/… See the full description on the dataset page: https://huggingface.co/datasets/agentvidbench/agentvidbench.AgentProcessBench
AgentProcessBench
AgentProcessBench is a benchmark for process-level evaluation of tool-using agents. Each example is a full agent trajectory with multi-turn messages, tool definitions, tool-use traces, reference outputs, and step-wise process labels.
The benchmark contains 1,000 trajectories in total, with 250 examples from each subset:
bfcl
gaia_dev
hotpotqa
tau2
arxiv.org/abs/2603.14465
claude-reasoning
Claude Reasoning Dataset
This dataset is a curated collection of prompts and responses generated by Claude by Anthropic. It combines high-quality long reasoning data from multiple sources to provide a focused training set for models requiring logic, math, and coding capabilities.
If multiple answers were generated for the same input during the data collection process, the entry with the shortest reasoning content was selected to ensure conciseness and high signal-to-noise ratio.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/claude-reasoning.agentvidbench
AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
Agentic Video Understanding Benchmark — 100 multiple-choice video QA questions, 26 options each (A-Z; ~3.8% random baseline)
by Seoyeon An*, Hyeonseo Jang*, Minsu Kim*, Chanho Lee, Younghan Park, Kangwook Lee (KRAFTON AI)
Layout
.
├── README.md
├── questions.jsonl # 100 rows — one per question
├── videos.jsonl # 71 rows — one per unique video
├── videos/… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench.agentmujo-joint-01
agentmujo-joint-01 (v0.1.0)
Joint trening mix protiv forgettinga: samo train splitovi (test/valid
ostaju čisti za evaluaciju).
Sastav (295): 117 function-calling + 88 agentic-terminal + 90 bosnian-core.
Metoda: deterministički shuffle (seed 11); ID-evi jedinstveni.
Format: JSONL, kanonska schema. Validacija: 295/295 ACCEPT.
Licenca: Apache-2.0. Bez ličnih podataka, bez tajni.
agentmujo-agentic-terminal
agentmujo-agentic-terminal (v0.1.0 — 541 trag)
Ručno dizajniran kanonski skup agentskih terminal-tragova na bosanskom
(ijekavica): dijagnoza → akcija → opservacija, dio
AgentMujo Training Frameworka.
Verzija: 0.1.0 · Tragova: 541 (split 450/42/49) · Jezik: bs-ijekavica
Format: JSONL kao function-calling skup; thinking uzorci nose
<think> reasoning blokove, svi nose tool_calls[] ili tekstualne odgovore.
Obrasci: jedan-korak statusi, višestepene dijagnoze (servis→mreža→store)… See the full description on the dataset page: https://huggingface.co/datasets/shaban2024/agentmujo-agentic-terminal.allenai-WildChat
AllenAI WildChat Combined Dataset
This unofficial repository provides the AllenAI WildChat Combined Dataset, which merges the WildChat-4.8M and WildChat-1M collections of human–ChatGPT conversations.
WildChat-1M contains 1 million chats, of which 25.53% are from GPT‑4 and the remainder from GPT‑3.5. These conversations cover a wide range of complex interactions, including code-switching, ambiguity, and political topics.
WildChat-4.8M originally comprised 4.8 million conversations.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/allenai-WildChat.agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.Agent-ValueBench
Agent-ValueBench
Agent-ValueBench constitutes the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions).
This Hugging Face release contains both structured JSONL tables for dataset viewing and Croissant metadata generation, and the original raw benchmark artifacts.
Repository Structure… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nips2026/Agent-ValueBench.Agent-ValueBench
Agent-ValueBench
Paper | Project Page | GitHub
Agent-ValueBench is the first comprehensive benchmark dedicated to evaluating the underlying values of autonomous agents. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that span 28 value systems (332 dimensions).
Repository Structure
README.md
data/
cases.jsonl
rubrics.jsonl
environments.jsonl
raw/
case/
rubric/
environment/
Data Files… See the full description on the dataset page: https://huggingface.co/datasets/Value4AI/Agent-ValueBench.AgenticRAGTracer
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG
Paper | Code
🎉 Our work has been accepted to ACL 2026 Findings!
AgenticRAGTracer is a benchmark designed to diagnose and evaluate multi-step retrieval reasoning in Agentic RAG systems. Unlike traditional benchmarks that provide only final questions and answers, AgenticRAGTracer includes intermediate hop-level questions that connect atomic questions to the final query. This allows… See the full description on the dataset page: https://huggingface.co/datasets/YqjMartin/AgenticRAGTracer.agentic-safety-gguf
agentic-safety-gguf: Training & Evaluation Datasets
Model: guerilla7/agentic-safety-ggufPaper: (https://arxiv.org/abs/2601.00848)Total: 80,992 examples (80,851 after deduplication)
Overview
Complete training and evaluation datasets for agentic-safety-gguf, a specialized Llama 3.1 8B model for agentic AI security analysis. Supports iterative continuation training methodology (V2→V3→V4) for full reproducibility.
Dataset Files
File
Examples
Size
Purpose… See the full description on the dataset page: https://huggingface.co/datasets/guerilla7/agentic-safety-gguf.PortBench-QA
PortBench QA Dataset
Dataset Description
6,269 structured question-answer pairs probing correlation-based financial reasoning for multi-asset portfolio management, generated from the PortBench Market Base Dataset.
Task Templates
Template
Task
Complexity
Pairs
T1
Return prediction — direction for next N days
1 (single asset)
1,000
T2
Risk assessment — VaR at given confidence level
1
1,000
T3
Position sizing — given max drawdown… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PortBench-QA.agentvidbench-sample
AgentVidBench (Sample): A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents
This repository is a representative sample of AgentVidBench, provided so reviewers can inspect data quality without downloading the full ~4GB+ corpus. The full dataset remains available at the link above.
Sample selection
The sample contains the first 10 questions (question_id 1–10) and the 9 unique videos they reference. IDs and filenames are preserved from the… See the full description on the dataset page: https://huggingface.co/datasets/KamiKrafton/agentvidbench-sample.sft-dataAgentForge-1152
AgentForge-1152: Tool-Verified General + Cyber Reasoning
1,152 evidence-grounded agent trajectories. 144 disjoint task families. A clean 50/50 split between general reasoning and defensive cybersecurity.
AgentForge-1152 is a model-agnostic English corpus for supervised fine-tuning and evaluating assistants that use tools, recover from failed checks, and ground final answers in captured evidence. It uses portable messages and JSON function definitions rather than a model-specific… See the full description on the dataset page: https://huggingface.co/datasets/0xKitkat/AgentForge-1152.UltraData-SFT-Agent-2609
UltraData-SFT-Agent-2609
📦 UltraData Collection |
🌐 UltraData |
🤗 MiniCPM5 Series
English |
中文
📚 Introduction
UltraData-SFT-Agent-2609 is the L3 refined data for Agent instruction-tuning within UltraData's L0-L4 tiered data management framework. Built for the post-training of MiniCPM5-2B, it complements UltraData-SFT-2605 (core-domain SFT) with executable Agent trajectories. The release contains approximately 500,000 samples spanning tool use… See the full description on the dataset page: https://huggingface.co/datasets/Supbatomic/UltraData-SFT-Agent-2609.agent-financial-interactions
Purple Flea Agent Financial Interactions
A dataset of API interactions for autonomous AI agents using financial infrastructure (Purple Flea). Useful for training agents that handle crypto, trading, gambling, domain registration, escrow, and free onboarding via faucet.
Research
This dataset is referenced in:
"Purple Flea: A Multi-Agent Financial Infrastructure Protocol for Autonomous AI Systems"
https://doi.org/10.5281/zenodo.18808440
The paper covers the economic model… See the full description on the dataset page: https://huggingface.co/datasets/PurpleFlea/agent-financial-interactions.JumpForge-Agentic-SE-3K
JumpForge-Agentic-SE-3K
JumpForge-Agentic-SE-3K is a structured synthetic dataset for training and evaluating
AI software-engineering agents. Its primary target is agent behavior across the software
engineering lifecycle, not raw code generation or memorization of programming-language syntax.
The dataset teaches an agent to:
understand intent and ambiguity before acting;
explore repositories and trace system behavior;
decompose work into reversible steps;
select tools based on… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JumpForge-Agentic-SE-3K.hendar-agentic-ai-dataset
Hendar Agentic AI Evaluation & Security Benchmark
A compact, expert-authored benchmark for evaluating trustworthy agentic AI systems across capability, tool use, retrieval, security, policy enforcement, multi-agent coordination and regression safety. The current release contains 104 synthetic cases: 13 cases in each of 8 domains.
This dataset is a public companion to the Agentic AI Academy by Hendar Mawan, PhD. It is designed for evaluation, CI regression testing, red-team… See the full description on the dataset page: https://huggingface.co/datasets/h0000w/hendar-agentic-ai-dataset.sci-agent-verification-cascade
Scientific Agent Verification Cascade
Public evaluation fixtures and verified aggregate results for testing whether
scientific claims keep their source, meaning, uncertainty, and verification
requirements as they move between AI agents.
This dataset accompanies the
Scientific Agent Verification Cascade
codebase. Version 0.2.0
contains synthetic evaluation data and aggregate-only results. It contains no
raw hosted-model response, private holdout identifier,
source-record… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/sci-agent-verification-cascade.
