datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.deepseek-v4-pro-0813-agentic
DeepSeek-V4-Pro 0813 Agentic (DS4)
A standalone, verifiable-first agentic training corpus: 19,072 training traces
plus 2,135 held-out evaluation rows (validation 1,070 / test 1,065), generated by
DeepSeek-V4-Pro 0813 (deepseek-v4-pro-0813, official API, thinking mode) across 13 verifiable task families,
each row admitted only after passing a deterministic programmatic verifier. The corpus is
designed to be directly usable for SFT, GRPO/RLVR, and NeMo Gym / NeMo RL
(verified… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/deepseek-v4-pro-0813-agentic.Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Creative-Professionals-Agentic-Tasks-1M.Agentic-Long-Context-Understanding-QA 📖 Agentic Long Context Understanding 📖
Self-Taught Agentic Long Context Understanding (Arxiv).
AgenticLU refines complex, long-context queries through self-clarifications and contextual grounding, enabling robust long-document understanding in a single pass.
Installation Requirements
This codebase is largely based on OpenRLHF and Helmet, kudos to them.
The requirements are the same
pip install openrlhf
pip install -r ./HELMET/requirements.txt… See the full description on the dataset page: https://huggingface.co/datasets/yzhuang/Agentic-Long-Context-Understanding-QA.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/rAVEUK/Creative-Professionals-Agentic-Tasks-1M.Creative-Professionals-Agentic-Tasks-1M
Creative Professionals Agentic Tasks (1M)
Abstract
A massive-scale, high-fidelity synthetic task dataset comprising 1,070,917 agentic command operations across 36 creative, technical, and engineering software environments. This dataset is engineered exclusively to stress-test, evaluate, and fine-tune multimodal AI agents designed for Agent Environment operation, complex software interaction, and multi-step reasoning within deep software infrastructures.… See the full description on the dataset page: https://huggingface.co/datasets/kryp1234/Creative-Professionals-Agentic-Tasks-1M.agentmujo-agentic-terminal
agentmujo-agentic-terminal (v0.1.0 — MVP)
Ručno dizajnirani višekoračni agentic/terminal tragovi na bosanskom
(ijekavica), dio AgentMujo Training Frameworka:
problem → dijagnoza → tool call → opservacija → analiza →
akcija → verifikacija → finalni odgovor. Agent nikada ne pretpostavlja
da je akcija uspjela.
Verzija: 0.1.0 · Tragova: 105 · Jezik: bs-ijekavica
Format: JSONL, ista schema kao function-calling
(schemas/dataset.schema.json u framework repou).
Obrasci: nginx… See the full description on the dataset page: https://huggingface.co/datasets/shaban2024/agentmujo-agentic-terminal.muse12-nemo-agentic
Muse Spark 1.2 High-Reasoning NeMo Agentic Dataset
A reproducible, verified 24,000-row synthetic agentic dataset generated with Meta Muse Spark 1.2, NeMo Gym, and deterministic task-family verifiers.
The project is a quality-focused successor to r0b0tlab/deepseek-v4-pro-0813-agentic. It keeps rollout prompts separate from reference trajectories and offline-training views, records usage and provenance, and does not publish private chain-of-thought.
[!IMPORTANT]
Status:… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/muse12-nemo-agentic.agentic-llm-pretraining-1.7b
Agentic LLM Pretraining Dataset
A pretraining corpus for small language models (1-3B parameters) optimized for agentic tasks. The corpus emphasizes learning to comprehend language, reason, follow instructions, and use tools over memorizing factual knowledge — the assumption is that domain knowledge will be provided at runtime via RAG. The idea is that this could enable much smaller pretraining corpora by omitting the large volumes of text typically needed to memorize facts.… See the full description on the dataset page: https://huggingface.co/datasets/visionscaper/agentic-llm-pretraining-1.7b.Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/yuqing1207/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.AgenticRAGTracer
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG
Paper | Code
🎉 Our work has been accepted to ACL 2026 Findings!
AgenticRAGTracer is a benchmark designed to diagnose and evaluate multi-step retrieval reasoning in Agentic RAG systems. Unlike traditional benchmarks that provide only final questions and answers, AgenticRAGTracer includes intermediate hop-level questions that connect atomic questions to the final query. This allows… See the full description on the dataset page: https://huggingface.co/datasets/YqjMartin/AgenticRAGTracer.daily-oracle
Daily Oracle
📰 Project Website📝 Paper - Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
Daily Oracle is a continuous evaluation benchmark using automatically generated QA pairs from daily news to assess how the future prediction capabilities of LLMs evolve over time.
Dataset Details
Question Type: True/False (TF) & Multiple Choice (MC)
Current Version*
Time Span: 2020.01.01 - 2026.07.18
Size: 20,376 TF questions and 18,557 MC… See the full description on the dataset page: https://huggingface.co/datasets/agentic-learning-ai-lab/daily-oracle.PortBench-QA
PortBench QA Dataset
Dataset Description
6,269 structured question-answer pairs probing correlation-based financial reasoning for multi-asset portfolio management, generated from the PortBench Market Base Dataset.
Task Templates
Template
Task
Complexity
Pairs
T1
Return prediction — direction for next N days
1 (single asset)
1,000
T2
Risk assessment — VaR at given confidence level
1
1,000
T3
Position sizing — given max drawdown… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PortBench-QA.agentic-safety-gguf
agentic-safety-gguf: Training & Evaluation Datasets
Model: guerilla7/agentic-safety-ggufPaper: (https://arxiv.org/abs/2601.00848)Total: 80,992 examples (80,851 after deduplication)
Overview
Complete training and evaluation datasets for agentic-safety-gguf, a specialized Llama 3.1 8B model for agentic AI security analysis. Supports iterative continuation training methodology (V2→V3→V4) for full reproducibility.
Dataset Files
File
Examples
Size
Purpose… See the full description on the dataset page: https://huggingface.co/datasets/guerilla7/agentic-safety-gguf.JumpForge-Agentic-SE-3K
JumpForge-Agentic-SE-3K
JumpForge-Agentic-SE-3K is a structured synthetic dataset for training and evaluating
AI software-engineering agents. Its primary target is agent behavior across the software
engineering lifecycle, not raw code generation or memorization of programming-language syntax.
The dataset teaches an agent to:
understand intent and ambiguity before acting;
explore repositories and trace system behavior;
decompose work into reversible steps;
select tools based on… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/JumpForge-Agentic-SE-3K.agentic-publication-protocol-dataset
APP compare-app benchmark
Paired reader conversations and blinded evaluations comparing an Agentic
Publication Protocol (APP) paper agent against a general repository-aware
agent, on 11 quantum-physics papers.
For each paper, a neutral reader asks the same scripted questions to both agents;
the two transcripts are anonymized and scored by a blinded evaluator on
accuracy, informativeness, grounding, and honesty (1-10).
Evaluator: Codex CLI, gpt-5.5, reasoning effort xhigh… See the full description on the dataset page: https://huggingface.co/datasets/phynics/agentic-publication-protocol-dataset.deepseek-v4-pro-0813-agentic
DeepSeek-V4-Pro 0813 Agentic (DS4)
A standalone, verifiable-first agentic training corpus: 19,072 training traces
plus 2,135 held-out evaluation rows (validation 1,070 / test 1,065), generated by
DeepSeek-V4-Pro 0813 (deepseek-v4-pro-0813, official API, thinking mode) across 13 verifiable task families,
each row admitted only after passing a deterministic programmatic verifier. The corpus is
designed to be directly usable for SFT, GRPO/RLVR, and NeMo Gym / NeMo RL
(verified… See the full description on the dataset page: https://huggingface.co/datasets/ArkhAngelLifeJiggy/deepseek-v4-pro-0813-agentic.json-mode-agentic-reasoningASQ
🤖 ASQ: Agentic Search Queryset
A dataset capturing RAG agents' search behaviours.
📖 Dataset Description
ASQ (Agentic Search Queryset) is a dataset designed to capture the search behaviors of the RAG agents.
It collects intermediate synthetic queries, retrieved documents, and thoughts (reasoning descriptions) produced or consumed by agents.
📊 Dataset Statistics
615k traces (0.12% incomplete)
614k answers
680k synthetic queries
680k retrieved ranked… See the full description on the dataset page: https://huggingface.co/datasets/AgenticSearchQueryset/ASQ.agentic-reasoning-benchmark
Agentic & Reasoning Benchmark (ARB) – Expanded
Ein synthetischer Benchmark mit 2.550 Fragen und Lösungen, optimiert für die Evaluation von Agentic Capabilities und Reasoning.
Überblick
Eigenschaft
Wert
Anzahl Beispiele
2.550
Kategorien
8
Schwierigkeitsgrade
easy / medium / hard
Formate
CSV + JSON
Reproduzierbarkeit
Generator-Skript (seed=42) enthalten
Lizenz
CC-BY-4.0
Kategorien
Kategorie
Anzahl
Beschreibung… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/agentic-reasoning-benchmark.PyFi-600K
Dataset Card for PyFi-600K
This dataset card aims to be a introduction for PyFi-600K, A financial VLM dataset containing 600K question-answer pairs generated via Adversarial agents.
AgenticFinLab/PyFi-600K/
├── README.md # Dataset documentation and description
├── images.zip # Compressed image files
├── PyFi-600K-dataset.csv # Q&A pairs in CSV format
├── PyFi-600K-dataset.json # Q&A pairs in JSON format
├── PyFi-600K-chain-dataset.json # Chain of Thought Q&A pairs dataset
└──… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PyFi-600K.hendar-agentic-ai-dataset
Hendar Agentic AI Evaluation & Security Benchmark
A compact, expert-authored benchmark for evaluating trustworthy agentic AI systems across capability, tool use, retrieval, security, policy enforcement, multi-agent coordination and regression safety.
This dataset is a public companion to the Agentic AI Academy by Hendar Mawan, PhD. It is designed for evaluation, CI regression testing, red-team exercises and engineering education—not as a generic instruction-tuning corpus.… See the full description on the dataset page: https://huggingface.co/datasets/h0000w/hendar-agentic-ai-dataset.eu-tenders-with-questions-for-agentic-checklist-filling
eu-tenders-with-questions-for-agentic-checklist-filling
Dataset Description
This dataset contains questions and answers for evaluating Retrieval-Augmented Generation (RAG) systems in the context of generative agentic checklist-filling. The dataset is designed to benchmark various RAG architectures (Hybrid RAG, Graph RAG, Multi-Hop/Agentic RAG) on document analysis tasks.
Dataset Summary
Total Questions: 97
Document Families: 7
Languages: EN
Domain: Procurement… See the full description on the dataset page: https://huggingface.co/datasets/tmskss/eu-tenders-with-questions-for-agentic-checklist-filling.betterwright-agentic-browser-50k
BetterWright Agentic Browser — 6,093-row stopped checkpoint
This is the public checkpoint of a generation run originally planned for 50,000 rows. Generation was stopped at the account owner's request and the exact 6,093 accepted rows were packaged. It is synthetic training data, not live browser recordings.
Contents
5,971 BetterWright demonstrations and 122 Playwright demonstrations.
32 task domains and 21 browser feature categories.
Harness-shaped conversations… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/betterwright-agentic-browser-50k.OWASP-Agentic-AI-Threats
OWASP Agentic AI Threats and Mitigations
Dataset Summary
The OWASP Agentic AI Threats and Mitigations Question Answering Dataset is a
synthetic instruction-style question-answering dataset derived from the OWASP
Agentic AI - Threats and Mitigations report.
The dataset is designed to support training, fine-tuning, retrieval evaluation,
and domain-specific question-answering use cases related to agentic AI security,
large language model agents, multi-agent… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/OWASP-Agentic-AI-Threats.multiagent-entropy-rawdata
When Does Multi-Agent Collaboration Help? An Entropy Perspective
📄 Paper · 💻 Code · 🌐 Project Page
This is the raw experimental data behind every figure and claim in paper. Use it to reproduce all results and conclusions reported in the paper.
Data Overview
Size: ~5 GB, 237 filesFormat: CSV (aggregated metrics) and JSON (entropy distributions, evaluation metrics)
The data is organized as follows:
1. Merged Dataset
merged_datasets/master.csv… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/multiagent-entropy-rawdata.agentic_coding_dataset
Agentic Coding Dataset
This dataset is a compilation of various coding and instruction-following datasets, designed to train agentic coding models.
Sources
This dataset aggregates samples from the following sources:
CodeAlpaca-20k
Instruction-following coding tasks.
Evol-CodeAlpaca-v1
Complex evolved coding instructions (WizardCoder style).
Code Review Instruct
Python code review, critique, and revision examples.
APPS (Automated Programming Progress Standard)… See the full description on the dataset page: https://huggingface.co/datasets/ethanker/agentic_coding_dataset.agentic-lightweight-envs-runtime-20260528
Lightweight Agentic RL Runtime Environments
This dataset contains prebuilt runtime environments for lightweight agentic reinforcement learning. It is intended to be used with the accompanying environment server/runtime code. The LLM synthesis pipeline used to create these environments is not required for serving this dataset.
Files
runtime_catalog.json.gz: prebuilt runtime task catalog consumed by the env server.
task_drafts.json.gz: source task drafts and… See the full description on the dataset page: https://huggingface.co/datasets/quantumfr/agentic-lightweight-envs-runtime-20260528.Agentic-Coding-Tessa
Agentic Coding Dataset for Tessa
A comprehensive dataset for training coding agents with tool-use, reasoning, and software engineering capabilities.
Dataset Composition
This dataset combines multiple high-quality sources:
hermes_reasoning (20.0%): Tool-use and reasoning dataset - interstellarninja/hermes_reasoning_tool_use
search_arena (15.0%): Search and retrieval tasks - lmarena-ai/search-arena-24k
arena_human_pref (15.0%): Human preference data for alignment -… See the full description on the dataset page: https://huggingface.co/datasets/smirki/Agentic-Coding-Tessa.
