datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AgenticRAGTracer
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG
Paper | Code
🎉 Our work has been accepted to ACL 2026 Findings!
AgenticRAGTracer is a benchmark designed to diagnose and evaluate multi-step retrieval reasoning in Agentic RAG systems. Unlike traditional benchmarks that provide only final questions and answers, AgenticRAGTracer includes intermediate hop-level questions that connect atomic questions to the final query. This allows… See the full description on the dataset page: https://huggingface.co/datasets/YqjMartin/AgenticRAGTracer.agentic-rag-redteam-bench
WARNING: HARMFUL CONTENT - RESEARCH USE ONLY
This dataset contains adversarial prompts, jailbreak attacks, toxic outputs, and other explicitly harmful content generated for AI safety research. Samples include prompt injections, social engineering payloads, misinformation, hate speech, instructions for illegal activities, phishing templates, and other dangerous material. All content is synthetic and produced by automated red-teaming pipelines for the sole purpose of evaluating and improving… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu/agentic-rag-redteam-bench.secai-agentic-rag-sft-v1
secAI Agentic-RAG SFT v1
A curated Vietnamese/English cybersecurity instruction and agentic-RAG supervised fine-tuning dataset. It teaches direct security assistance as well as grounded tool-use behaviour: tool selection, JSON arguments, consuming tool results, no-result handling, and multi-turn follow-ups.
This is a frozen training release, not a standalone claim of safety, factual correctness, or production readiness. Keep human review and authorization controls for all… See the full description on the dataset page: https://huggingface.co/datasets/DuyTa/secai-agentic-rag-sft-v1.AgenticRag
AgenticRAG-FP
This repository describes AgenticRAG-FP, a research dataset and evaluation
suite for studying how failures propagate through agentic retrieval-augmented
generation pipelines. The dataset normalizes multi-hop QA examples into a common
schema, runs real or mock ReAct-style RAG agents over them, injects controlled
failures at specific retrieval/reasoning hops, and records whether diagnostic
methods can recover the true root cause after the failure has propagated.
The… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/AgenticRag.dataflow-demo-multihop-AgenticRAG
DataFlow Demo -- AgenticRAG Pipeline
This dataset card serves as a demo for showcasing the Agentic Retrieval-Augmented Generation (AgenticRAG) pipeline of the DataFlow Project.It provides a clear comparison between single-hop input questions and their corresponding multi-hop, agent-driven reasoning outputs.
Overview
The AgenticRAG pipeline is designed to solve complex questions that require multi-step reasoning across multiple documents.Instead of answering questions in… See the full description on the dataset page: https://huggingface.co/datasets/OpenDCAI/dataflow-demo-multihop-AgenticRAG.alpha_agentic_search_ragdatadataflow-demo-AgenticRAG
DataFlow demo -- Agentic RAG Pipeline
This dataset card serves as a demo for showcasing the Agentic RAG data processing pipeline of the Dataflow Project. It provides an intuitive view of the pipeline’s inputs and outputs.
Overview
The Agentic RAG Data Synthesis Pipeline is an end-to-end framework to:
Support RL-based agentic RAG training.
Generate high-quality pairs of questions and answers from provided text contents.
This pipeline only need text contexts for… See the full description on the dataset page: https://huggingface.co/datasets/OpenDCAI/dataflow-demo-AgenticRAG.dataflow-demo-AgenticRAG
DataFlow Demo -- AgenticRAG Pipeline
This dataset card serves as a demo for showcasing the Agentic Retrieval-Augmented Generation (AgenticRAG) pipeline of the DataFlow Project.It provides a clear comparison between single-hop input questions and their corresponding multi-hop, agent-driven reasoning outputs.
Overview
The AgenticRAG pipeline is designed to solve complex questions that require multi-step reasoning across multiple documents.Instead of answering questions in… See the full description on the dataset page: https://huggingface.co/datasets/YqjMartin/dataflow-demo-AgenticRAG.changelogs-agentic-rag-3docs-generationsagentic-rag-chunkschangelogs-agentic-rag-10docs-generationsagentic_RAGagentic-rag-redteam-bench
Mirror note: This dataset is a mirror of Fujitsu/agentic-rag-redteam-bench, maintained by the same author, intherejeet, to preserve availability if organization access is interrupted. Access controls and usage restrictions are intended to match the source dataset.
WARNING: HARMFUL CONTENT - RESEARCH USE ONLY
This dataset contains adversarial prompts, jailbreak attacks, toxic outputs, and other explicitly harmful content generated for AI safety research. Samples include prompt injections… See the full description on the dataset page: https://huggingface.co/datasets/intherejeet/agentic-rag-redteam-bench.personalization-agentic-rag-10docs-generationsagentic-rag-hukum-keluarga-islam-datasetpersonalization-agentic-rag-5docs-generationschangelogs-agentic-rag-5docs-generationsrag-agentic-haiku-hard20-20260318-151930
rag-agentic-haiku-hard20
Run timestamp: 2026-03-18T15:19:30
Benchmark: CREATE — associative reasoning via knowledge-graph path generation
Run Parameters
Parameter
Value
Mode
rag_agentic
Generator model
claude-haiku-4-5-20251001
Temperature
None
Max tokens
4096
Run Summary
Metric
Value
Instances processed
20
Total paths generated
240
Avg paths / instance
12.00
Min paths / instance
0
Max paths / instance
18
Total cost… See the full description on the dataset page: https://huggingface.co/datasets/connections-dev/rag-agentic-haiku-hard20-20260318-151930.personalization-agentic-rag-3docs-generationsrag-agentic
