datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.rag-bench
Dataset card for RAG-BENCH
Data Summary
RAG-bench aims to provide results of many commonly used RAG datasets. All the results in this dataset are evaluated by the RAG evaluation tool Rageval, which could be easily reproduced with the tool.
Currently, we have provided the results of ASQA dataset,ELI5 dataset and HotPotQA dataset.
Data Instance
ASQA
{
"ambiguous_question":"Who is the original artist of sound of silence?",
"qa_pairs":[{… See the full description on the dataset page: https://huggingface.co/datasets/golaxy/rag-bench.rag-corpus-v1
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
SZLHOLDINGS/rag-corpus-v1 — Agentic-RAG corpus + per-organ FAISS indexes
Doctrine v10/v11. Embedding model: BAAI/bge-base-en-v1.5 (768-dim).
Built by the agentic-RAG SHIP directive (390_AGENTIC_RAG_FAISS_PER_SPACE).
Contents
corpus.jsonl — 762 chunks, each ~512 tokens with… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/rag-corpus-v1.ragtopia_oldrag_instruct_benchmark_tester
Dataset Card for RAG-Instruct-Benchmark-Tester
Dataset Summary
This is an updated benchmarking test dataset for "retrieval augmented generation" (RAG) use cases in the enterprise, especially for financial services, and legal. This test dataset includes 200 questions with context passages pulled from common 'retrieval scenarios', e.g., financial news, earnings releases,
contracts, invoices, technical articles, general news and short texts.
The questions are segmented… See the full description on the dataset page: https://huggingface.co/datasets/llmware/rag_instruct_benchmark_tester.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/grasson/t2-ragbench.rag-dx
RAG-Dx: a diagnostic benchmark for retrieval
This dataset is for evaluation. It is not training data and should not be used to train
or fine-tune models.
Most retrieval benchmarks give you a number. A number tells you that something is wrong,
not what. RAG-Dx reports how much a retrieval stack degrades on each of eight specific
failure modes, so the output points at a fix.
Code, harness and reproduction scripts: https://github.com/chakshu-dhannawat/rag-dx
What is in… See the full description on the dataset page: https://huggingface.co/datasets/Chakshu123/rag-dx.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/tomsummerfield/t2-ragbench.Wikipedia_RAG_QA_Classification
🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training
📊 Dataset Description
This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning.
🖥️ Demo Interface: Discord
Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h
The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.OmniSearch_RAGkdv-rag-benchmark
KDV RAG Benchmark
A retrieval benchmark dataset for Turkish VAT (KDV, Katma Değer Vergisi) law — built by adding retrieval layers one at a time (chunking, model choice, hybrid search, reranking, query rewriting, historical/date filtering) and statistically validating each one individually (see Results).
Dataset structure
Splits
Split
Records
Period
train
728
2018-2023
test
154
2024-2026
Split strategy: temporal — train and test… See the full description on the dataset page: https://huggingface.co/datasets/dokukoza/kdv-rag-benchmark.geo-injection-rag-attack-data
Can It Reach the Generator? Investigating the Survival of GEO Prompt-Injection Attacks in Realistic RAG Settings
This dataset contains the prompt-injection attack data presented in the paper Can It Reach the Generator? Investigating the Survival of Prompt-Injection Attacks in Realistic RAG Settings.
The dataset is used to… See the full description on the dataset page: https://huggingface.co/datasets/Euanyu/geo-injection-rag-attack-data.RAGPulse
RAGPulse: A Real-World RAG Workload Trace to Optimize RAG Serving Systems
🌐 Github Link |
🤗 Workload Trace |
📑 Arxiv Paper |
🤖 How to use?
RAGPulse is a real-world RAG workload trace collected from an university-wide Q&A service scenario. The system has been serving over 40,000 students and faculties since April 2024, providing intelligent policy Q&A services. The trace contains a total of 7,106 records entries, sampled from one week of our Q&A service.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/flashserve/RAGPulse.elkarhizketak-RAG
Dataset Card for ElkarHizketak RAG and its Disruptor Variants
Base and disruptor variants of ElkarHizketak, built to stress-test conversational RAG systems in Basque under realistic interaction patterns (conversational openings, topic shifts).
Dataset Details
Dataset Description
This dataset extends ElkarHizketak with a base variant (rewritten opening queries, retrieval-needed labels, retrieved chunks) and disruptor variants that inject… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/elkarhizketak-RAG.k8s-docs-rag-bench
k8s-docs-rag-bench
Paper: Analyzing Quality--Latency--Resource Trade-offs in a Technical Documentation RAG Assistant Using LoRA Adaptation (arXiv:2605.28222)
Code: github.com/EugPal/rag-lora-tradeoffs
A small, fully-grounded benchmark for retrieval-augmented question answering
(RAG) over the official Kubernetes documentation, together with the full
set of LLM-judge labels used in the accompanying preprint
"Analyzing Quality-Latency-Resource Trade-offs in a Technical… See the full description on the dataset page: https://huggingface.co/datasets/evgenypal/k8s-docs-rag-bench.arabic-rag-chat-8k-eval
arabic-rag-chat-8k-eval
Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models:
the test split, every model's raw replies, every judge verdict, and the rendered
report for each. Thirteen judged models, all scored on the same 1,651 prompts
by the same judge at temperature 0.0, so the comparison below is like-for-like
and can be recomputed offline without a GPU or a judge server.
This is the measurement half of
oddadmix/100M-8192-Nawah-dsv4;
the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.hse-multimodal-rag-corpus
HSE Multimodal RAG Corpus
Chunks, labeled QA (including out-of-scope abstention), and published retrieval metrics.
chunks.jsonl
qa_pairs.jsonl
eval_results.json
benchmark_report.json
Light-RAG-Marketing-Assets-Agent
🖼️ Light RAG Marketing Assets Agent — Pre-ingested Data
Pre-ingested LightRAG knowledge graph and vector data from 420 marketing images
analyzed with Gemini Vision API (gemini-3.5-flash) and processed through GPT-4o
for entity extraction and relationship mapping.
GitHub repo: 0xrphl/Light-RAG-Marketing-Assets-Agent
📊 Dataset Statistics
Metric
Value
Source images
420 (JPG/PNG/WebP)
Text chunks
2,095 (5 per image: core, visual, people/setting… See the full description on the dataset page: https://huggingface.co/datasets/0xrphl/Light-RAG-Marketing-Assets-Agent.soc-playbook-rag-corpus
SOC Playbook RAG Corpus
Chunks, labeled QA (including out-of-scope abstention), and published retrieval metrics.
chunks.jsonl
qa_pairs.jsonl
eval_results.json
benchmark_report.json
procure-hybrid-rag-corpus
Procure Hybrid RAG Corpus
Chunks, labeled QA (including out-of-scope abstention), and published retrieval metrics.
chunks.jsonl
qa_pairs.jsonl
eval_results.json
benchmark_report.json
contract-clause-rag-corpus
Contract Clause RAG Corpus
Chunks, labeled QA (including out-of-scope abstention), and published retrieval metrics.
chunks.jsonl
qa_pairs.jsonl
eval_results.json
benchmark_report.json
petrosafe-rag-corpus-fa
PetroSafe RAG Corpus (FA/EN)
Bilingual (Persian/English) knowledge corpus for process safety and HSE in oil, gas, and
petrochemical operations. Built for alirezaaminzadeh/petrosafe-rag-fa, the retrieval architecture
is inherited unchanged from hse-multimodal-rag-corpus
(hybrid BM25 + word/char TF-IDF, mandatory citations, abstention) — this repo supplies new
domain content, not a new retrieval method.
Data honesty (please read before citing any number from this… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/petrosafe-rag-corpus-fa.noisy-rag-bench
noisy-rag-bench
A retrieval corpus and QA set for measuring what realistic document noise does to a
RAG pipeline, plus the benchmark run over 18 noise conditions.
Retrieval benchmarks run on clean text. Documents inside a bank or a law firm are
scans: OCR confusions, running headers, hyphens broken across lines, redacted spans.
This is the corpus for measuring that, and the result it was built to expose.
Code and full write-up: https://github.com/lgoyal6/noisy-rag-bench… See the full description on the dataset page: https://huggingface.co/datasets/lgoyal/noisy-rag-bench.vscode-issue-rag-artifactsdpo_long_form_gpt5_sft_0921arabic-rag-support-25K
Arabic RAG customer-support scenarios (27,927 rows)
Synthetic Modern Standard Arabic customer-support scenarios for training small
RAG answerers, distilled from unsloth/gemma-4-31B-it-NVFP4 on a local vLLM.
Built as the training set for oddadmix/Nawah-50M-RAG-Support.
Each row: a customer question + the knowledge-base chunks of one fictional
company (products, prices, policies, FAQ entries) + the ideal grounded agent
answer. One generation request invents one company KB and 4 QA… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-support-25K.ClaudioItaly__intelligence-cod-rag-7b-v3-details
Dataset Card for Evaluation run of ClaudioItaly/intelligence-cod-rag-7b-v3
Dataset automatically created during the evaluation run of model ClaudioItaly/intelligence-cod-rag-7b-v3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ClaudioItaly__intelligence-cod-rag-7b-v3-details.incident-rag-runbooks
IncidentRAG synthetic runbooks
58 short incident-response runbooks and 100 reference questions for the IncidentRAG demo. Seed 4.
This is fixture data (level 1). It does not prove operational incident-response quality. Organization dataset, collection, and static card are public. Live Gradio is created by scripts/publish.py.
Files
data/documents.jsonl
data/questions.jsonl
data/splits.json
data/schema.json
data/evaluation.json
data/sample_preview.json… See the full description on the dataset page: https://huggingface.co/datasets/AriaAICompany/incident-rag-runbooks.gpt_oss_120b_sf_all_correct
