datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.jeb-rag
JEB-Bench
Charging the Gate Rent: Measured-Energy Accounting for Adaptive Retrieval-Augmented Generation
⚠️ Status: under construction. Phase 0 (measurement validation) and Phase 1
(index construction) are landing now. The oracle matrix (bench/oracle/) is
populated in Phase 2 and this card will be revised when it is complete. Do not
cite numbers from this repository until the status line says complete.
What this is
The first public per-query × per-configuration… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/jeb-rag.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.medicalpark-rag
Medical Park Türkçe Sağlık Makaleleri — RAG Sistemi
Türkçe tıbbi makaleler üzerine kurulmuş, eşik (threshold) tabanlı bir Retrieval-Augmented Generation (RAG) altyapısı.
1. Veri Seti
Kaynak: umutertugrul/turkish-hospital-medical-articles (CC BY 4.0)
Veri seti içeriği: 14 farklı Türk hastane/sağlık kuruluşunun web sitesinden çekilmiş Türkçe tıbbi makaleler, her kuruluş ayrı bir .parquet dosyası olarak sunuluyor (toplam ~25.000 makale, 14 kaynak: Acıbadem… See the full description on the dataset page: https://huggingface.co/datasets/Toivo0/medicalpark-rag.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/grasson/t2-ragbench.narrativeqa-rag
NarrativeQA RAG
Dataset for Retrieval-Augmented Generation (RAG) based on NarrativeQA.
Structure
Subset
Splits
Description
corpus
train (default)
Wikipedia plot summaries shared across all query splits
queries
train, dev, test
Reading comprehension questions
qrels
train, dev, test
Relevance judgments (query ↔ document)
answers
train, dev, test
Reference answers (longest annotated answer)
Dataset statistics
Split
Queries… See the full description on the dataset page: https://huggingface.co/datasets/DinoStackAI/narrativeqa-rag.turkish-medical-rag
🩺 Turkish Medical RAG
Hierarchical Parent–Child Retrieval-Augmented Generation for Turkish Medical Documents
📌 Proje Hakkında
Bu proje, Türkçe tıbbi dokümanlar üzerinde çalışan uçtan uca bir
Retrieval-Augmented Generation (RAG) sistemi geliştirmek amacıyla hazırlanmıştır.
Sistem bir kullanıcı sorusu aldığında önce doküman koleksiyonundaki küçük ve
anlamsal olarak odaklı parçalar (child chunks)… See the full description on the dataset page: https://huggingface.co/datasets/sedayzc/turkish-medical-rag.rag-dx
RAG-Dx: a diagnostic benchmark for retrieval
This dataset is for evaluation. It is not training data and should not be used to train
or fine-tune models.
Most retrieval benchmarks give you a number. A number tells you that something is wrong,
not what. RAG-Dx reports how much a retrieval stack degrades on each of eight specific
failure modes, so the output points at a fix.
Code, harness and reproduction scripts: https://github.com/chakshu-dhannawat/rag-dx
What is in… See the full description on the dataset page: https://huggingface.co/datasets/Chakshu123/rag-dx.RAG-Grounded-QA-188k
🎯 RAG Grounded QA 186K
The Anti-Hallucination Dataset
Teach language models to answer from context — or shut up trying.
Built by NovachronoAI — Precision AI for the real world.
Full Dataset (186K) · 20K Subset · Schema · Sources · Usage Guide
🧠 Why This Dataset Exists
Most QA datasets teach models what to say. This one also teaches them when to stay silent.
RAG (Retrieval-Augmented Generation) systems have a fatal flaw: the model hallucinates when… See the full description on the dataset page: https://huggingface.co/datasets/NovachronoAI/RAG-Grounded-QA-188k.rag-qa-fulltext-ptbr
RAG QA Full-Text PT-BR Mistral
A large-scale dataset of Brazilian Portuguese RAG-style question-answer pairs
with grounded evidence spans, generated from Madras1/corpus-ptbr-v1 documents
using Mistral models. Every answer is anchored to literal quotations from the
source text, making this dataset suitable for training and evaluating
retrieval-augmented generation systems, extractive QA models, and reading
comprehension benchmarks in Portuguese.
Two configurations are available:… See the full description on the dataset page: https://huggingface.co/datasets/Madras1/rag-qa-fulltext-ptbr.rag-mini-bioasq-with-metadataThis dataset is an extension of the rag-mini-bioasq dataset.
Its difference resides in the text-corpus part of the aforementioned set where the metadata was added for each passage.
Metadata contains six separate categories, each in a dedicated column:
Year of the publication (publish_year)
Type of the publication (publish_type)
Country of the publication - often correlated with the homeland of the authors (country)
Number of pages (no_pages)
Authors (authors)
Keywords (keywords)
Athar-RAG-Hub
Athar RAG Hub 🕌
Collection
Chunks
seerah
5,852
telco-dpr-rag
Telco-DPR RAG
Dataset for Retrieval-Augmented Generation (RAG) based on Telco-DPR.
Structure
Subset
Splits
Description
corpus
train (default)
3GPP technical passages (text + tables) shared across all query splits
queries
train, dev, test
Synthetic telecom QA questions
qrels
train, dev, test
Relevance judgments (query ↔ passage)
answers
train, dev, test
Reference answers
Dataset statistics
Split
Queries
Corpus
train… See the full description on the dataset page: https://huggingface.co/datasets/DinoStackAI/telco-dpr-rag.qasper-rag
QASPER RAG
Dataset for Retrieval-Augmented Generation (RAG) based on QASPER.
Structure
Subset
Splits
Description
corpus
train (default)
Paper chunks (abstract + full-text paragraphs) shared across all query splits
queries
train, dev, test
Information-seeking questions over scientific papers
qrels
train, dev, test
Relevance judgments (query ↔ paragraph chunk)
answers
train, dev, test
Reference answers (longest valid free-form answer)
top_ranked… See the full description on the dataset page: https://huggingface.co/datasets/DinoStackAI/qasper-rag.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/tomsummerfield/t2-ragbench.Wikipedia_RAG_QA_Classification
🏛️ Wikipedia RAG QA Dataset for Retrieval-Augmented Generation Training
📊 Dataset Description
This dataset contains 300,000+ validated model-generated responses to Wikipedia content, specifically designed for Retrieval-Augmented Generation (RAG) applications and SQL database insertion tasks. Generated by Jeeney AI Reloaded 207M GPT with specialized RAG tuning.
🖥️ Demo Interface: Discord
Live Chat Demo on Discord: https://discord.gg/Xe9tHFCS9h
The full CJ… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Wikipedia_RAG_QA_Classification.kdv-rag-benchmark
KDV RAG Benchmark
A retrieval benchmark dataset for Turkish VAT (KDV, Katma Değer Vergisi) law — built by adding retrieval layers one at a time (chunking, model choice, hybrid search, reranking, query rewriting, historical/date filtering) and statistically validating each one individually (see Results).
Dataset structure
Splits
Split
Records
Period
train
728
2018-2023
test
154
2024-2026
Split strategy: temporal — train and test… See the full description on the dataset page: https://huggingface.co/datasets/dokukoza/kdv-rag-benchmark.bioasq-rag-13b-resplit
BioASQ RAG 13B (Resplit)
Reshuffled version of DinoStackAI/bioasq-rag-13b for Retrieval-Augmented Generation (RAG).
All original train, dev and test queries were merged, shuffled with seed 42, and reassigned using:
0.2 of all queries → test
0.2 of the remaining queries → dev
the rest → train
The shared PubMed corpus is unchanged from the source dataset.
Structure
Subset
Splits
Description
corpus
train (default)
PubMed abstracts shared across all query… See the full description on the dataset page: https://huggingface.co/datasets/DinoStackAI/bioasq-rag-13b-resplit.turkish-legal-rag
Turkish Legal RAG Corpus — Türk Hukuku için Açık RAG Datasetı
Tek cümle: 25 önemli Türk kanununun (mevzuat.gov.tr kaynaklı, madde bazlı temiz chunk'lar) + 290 manuel doğrulanmış soru-cevap altın benchmark'ının olduğu açık kaynak Türkçe hukuk RAG datasetı.
🇹🇷 Türkçe Özet — Bu dataset, Türkçe hukuk uygulamaları için sıfırdan üretilmiş açık ve denetlenebilir bir RAG corpus'udur. mevzuat.gov.tr üzerinden alınan 25 ana kanunun madde madde temizlenmiş, chunk'lanmış sürümünü (6.350… See the full description on the dataset page: https://huggingface.co/datasets/mtntasci/turkish-legal-rag.ID_REG_MD_RAG
📑 Indonesian Regulation Markdown RAG Dataset (ID_REG_MD_RAG)
This repository contains a highly structured, Markdown-optimized collection of Indonesian Regulations (Peraturan Perundang-undangan). This dataset is specifically engineered to solve the "structure loss" problem often encountered when building Retrieval-Augmented Generation (RAG) systems for complex legal documents. 🏛️
💡 The Concept: Structural Integrity for RAG
Legal documents in Indonesia follow a… See the full description on the dataset page: https://huggingface.co/datasets/Azzindani/ID_REG_MD_RAG.elkarhizketak-RAG
Dataset Card for ElkarHizketak RAG and its Disruptor Variants
Base and disruptor variants of ElkarHizketak, built to stress-test conversational RAG systems in Basque under realistic interaction patterns (conversational openings, topic shifts).
Dataset Details
Dataset Description
This dataset extends ElkarHizketak with a base variant (rewritten opening queries, retrieval-needed labels, retrieved chunks) and disruptor variants that inject… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/elkarhizketak-RAG.rag-qa-logs-corpus-data
🧠📚 RAG QA Logs & Corpus (Synthetic)
🧪 Multi-table synthetic RAG telemetry for quality, hallucinations, latency, and cost
A production-style, privacy-safe synthetic dataset that mimics telemetry exported from a real RAG system — from corpus → chunks → retrieval events → eval runs.
✅ Fully synthetic (no real users / orgs / PII).
⚡ Quick facts
Total rows: 103,255 across 6 linked tables
Labels (in eval_runs): is_correct, hallucination_flag, faithfulness_label… See the full description on the dataset page: https://huggingface.co/datasets/tarekmasryo/rag-qa-logs-corpus-data.arabic-rag-chat-8k-eval
arabic-rag-chat-8k-eval
Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models:
the test split, every model's raw replies, every judge verdict, and the rendered
report for each. Thirteen judged models, all scored on the same 1,651 prompts
by the same judge at temperature 0.0, so the comparison below is like-for-like
and can be recomputed offline without a GPU or a judge server.
This is the measurement half of
oddadmix/100M-8192-Nawah-dsv4;
the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.hse-multimodal-rag-corpus
HSE Multimodal RAG Corpus
Chunks, labeled QA (including out-of-scope abstention), and published retrieval metrics.
chunks.jsonl
qa_pairs.jsonl
eval_results.json
benchmark_report.json
procure-hybrid-rag-corpus
Procure Hybrid RAG Corpus
Chunks, labeled QA (including out-of-scope abstention), and published retrieval metrics.
chunks.jsonl
qa_pairs.jsonl
eval_results.json
benchmark_report.json
soc-playbook-rag-corpus
SOC Playbook RAG Corpus
Chunks, labeled QA (including out-of-scope abstention), and published retrieval metrics.
chunks.jsonl
qa_pairs.jsonl
eval_results.json
benchmark_report.json
petrosafe-rag-corpus-fa
PetroSafe RAG Corpus (FA/EN)
Bilingual (Persian/English) knowledge corpus for process safety and HSE in oil, gas, and
petrochemical operations. Built for alirezaaminzadeh/petrosafe-rag-fa, the retrieval architecture
is inherited unchanged from hse-multimodal-rag-corpus
(hybrid BM25 + word/char TF-IDF, mandatory citations, abstention) — this repo supplies new
domain content, not a new retrieval method.
Data honesty (please read before citing any number from this… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/petrosafe-rag-corpus-fa.contract-clause-rag-corpus
Contract Clause RAG Corpus
Chunks, labeled QA (including out-of-scope abstention), and published retrieval metrics.
chunks.jsonl
qa_pairs.jsonl
eval_results.json
benchmark_report.json
noisy-rag-bench
noisy-rag-bench
A retrieval corpus and QA set for measuring what realistic document noise does to a
RAG pipeline, plus the benchmark run over 18 noise conditions.
Retrieval benchmarks run on clean text. Documents inside a bank or a law firm are
scans: OCR confusions, running headers, hyphens broken across lines, redacted spans.
This is the corpus for measuring that, and the result it was built to expose.
Code and full write-up: https://github.com/lgoyal6/noisy-rag-bench… See the full description on the dataset page: https://huggingface.co/datasets/lgoyal/noisy-rag-bench.k8s-docs-rag-bench
k8s-docs-rag-bench
Paper: Analyzing Quality--Latency--Resource Trade-offs in a Technical Documentation RAG Assistant Using LoRA Adaptation (arXiv:2605.28222)
Code: github.com/EugPal/rag-lora-tradeoffs
A small, fully-grounded benchmark for retrieval-augmented question answering
(RAG) over the official Kubernetes documentation, together with the full
set of LLM-judge labels used in the accompanying preprint
"Analyzing Quality-Latency-Resource Trade-offs in a Technical… See the full description on the dataset page: https://huggingface.co/datasets/evgenypal/k8s-docs-rag-bench.
