datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
expert-rag-benchmarks
Expert RAG Benchmarks
A unified collection of four expert-level legal RAG benchmarks, exposed as six
named splits and three relational configurations: questions, documents, and
qrels.
The KCL split is named kcl_essay because Hugging Face split identifiers do not
permit hyphens; its source name remains kcl-essay.
Loading
from datasets import load_dataset
repo_id = "jinulee-v/expert-rag-benchmarks"
questions = load_dataset(repo_id, "questions", split="housing")… See the full description on the dataset page: https://huggingface.co/datasets/jinulee-v/expert-rag-benchmarks.Japanese-RAG-Generator-Benchmark
Japanese RAG Generator Benchmark: 日本語 RAG における Generator 評価ベンチマーク
Japanese RAG Generator Benchmark (J-RAGBench) は日本語RAGにおけるGeneratorに用いるLLMの評価データセットを提供する。
実運用時のRAGに求められる多様な評価カテゴリを同一条件下で評価可能であり、複数の評価カテゴリが同時に出現する問題が含まれるQAデータセットを人手および、補助的にOpenAI API(gpt-4.1-2025-04-14)を用いて構築した。
J-RAGBenchの評価カテゴリ
Integration: 2~3文書程度の複数の情報源から適切な根拠を抽出・統合して回答を導く
Reasoning: 抽出された情報を踏まえて多段階の推論や数値計算などを実行する
Logical: 質問・関連文書間での語彙や表現の差異を解釈し、適切な回答を導く
Table:… See the full description on the dataset page: https://huggingface.co/datasets/neoai-inc/Japanese-RAG-Generator-Benchmark.silma-rag-qa-benchmark-v1.0
SILMA RAGQA Benchmark Dataset V1.0
SILMA RAGQA is a dataset and benchmark created by silma.ai to assess the effectiveness of Arabic Language Models in Extractive Question Answering tasks, with a specific emphasis on RAG applications
The benchmark includes 17 bilingual datasets in Arabic and English, spanning various domains
What capabilities does the benchmark test?
General Arabic and English QA capabilities
Ability to handle short and long contexts
Ability to… See the full description on the dataset page: https://huggingface.co/datasets/silma-ai/silma-rag-qa-benchmark-v1.0.jev-rag-benchmark
Jev RAG Benchmark (English)
Frozen-candidate-pool evaluation of TypeSafe Jev 1.13 as the reranking
and decision layer of a RAG pipeline, compared with a NVIDIA cross-encoder
and with no reranking at all, on English XQuAD and SciFact. Every published
run used free tiers (total cost: $0).
This repository contains the raw per-query artifacts, per-run reports with
paired bootstrap confidence intervals, calibration tables, and a plain-text
leaderboard.
Reranking (nDCG@10… See the full description on the dataset page: https://huggingface.co/datasets/emretheus/jev-rag-benchmark.kdv-rag-benchmark
KDV RAG Benchmark
A retrieval benchmark dataset for Turkish VAT (KDV, Katma Değer Vergisi) law — built by adding retrieval layers one at a time (chunking, model choice, hybrid search, reranking, query rewriting, historical/date filtering) and statistically validating each one individually (see Results).
Dataset structure
Splits
Split
Records
Period
train
728
2018-2023
test
154
2024-2026
Split strategy: temporal — train and test… See the full description on the dataset page: https://huggingface.co/datasets/dokukoza/kdv-rag-benchmark.Korean-RAG-LLM-Judge-Benchmark
Korean RAG LLM-as-Judge Benchmark
allganize/RAG-Evaluation-Dataset-KO의 한국어 RAG 300 Q&A 위에 46개 답변 생성 모델의 답변과 LLM-as-Judge 평가 결과를 얹은 데이터셋입니다. 원문 문항·기준 답변·출처 PDF 메타데이터는 allganize 데이터셋을 그대로 따릅니다.
🔗 분석 코드·단계별 실험 보고서: https://github.com/BAEM1N/RAG-Evaluation · 📊 요약 대시보드: https://rag.baeum.ai.kr
🎯 한 줄 결론
RAG 파이프라인 최적화 > 모델 업그레이드. 동일 GPT-5.4를 쓰면서도 검색 파이프라인만 잘 짜면 GPT-5.4-pro(약 10배 비싼 모델)보다 +6.0pp 높은 정확도(0.827) 를 낸다.
Pipeline
Generator
Accuracy
🥇 Cartesian… See the full description on the dataset page: https://huggingface.co/datasets/BAEM1N/Korean-RAG-LLM-Judge-Benchmark.rag-hallucination-benchmark
RAG Hallucination Benchmark
Context
Retrieval-Augmented Generation (RAG) is the industry standard for reducing LLM hallucinations, but detecting when a RAG system fails is a massive challenge. Most existing benchmarks focus only on massive Deep Learning models and lack tabular features.
This dataset provides a clean, engineered setup to train models (from XGBoost to RoBERTa) to detect hallucinations, predict context faithfulness, and measure answer relevance.… See the full description on the dataset page: https://huggingface.co/datasets/vkshdev/rag-hallucination-benchmark.German-RAG-LLM-EASY-BENCHMARK
German-RAG-LLM-EASY-BENCHMARK
German-RAG - German Retrieval Augmented Generation
Dataset Summary
This German-RAG-LLM-BENCHMARK represents a specialized collection for evaluating language models with a focus on source citation, time difference stating in RAG-specific tasks.
To evaluate models compatible with OpenAI-Endpoints you can refer to our Github Repo: https://github.com/avemio-digital/German-RAG-LLM-EASY-BENCHMARK/
Most of the Subsets are synthetically… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-LLM-EASY-BENCHMARK.Hiro-Pharma-RAG-Benchmark
Hiro Pharma RAG Benchmark
This private dataset repository contains multilingual biomedical RAG benchmark data associated with the paper CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine.
The benchmark is designed to evaluate whether retrieval-augmented language models can answer biomedical questions while selecting and citing useful evidence and filtering out noisy or irrelevant references.
Repository Contents
File
Language… See the full description on the dataset page: https://huggingface.co/datasets/PatSnap/Hiro-Pharma-RAG-Benchmark.priviledge-prospectus-rag-benchmark
PrivilEdge Prospectus RAG Benchmark
40 question-answer pairs over the PrivilEdge Prospectus for Switzerland (Feb 2026, 354 pages) for evaluating RAG systems with page-level citation accuracy.
Source Document
The source PDF (LO_fund_fr.pdf) is included in this repository. It is a publicly available regulatory document: the partial prospectus for Swiss investors of PrivilEdge, a Luxembourg SICAV managed by Lombard Odier Funds (Europe) S.A.
Benchmark Design… See the full description on the dataset page: https://huggingface.co/datasets/PeterAM4/priviledge-prospectus-rag-benchmark.rag-qa-benchmark-ko
Surromind/rag-qa-benchmark-ko
한국어 위키백과에서 생성한 RAG 검색·생성 동시 평가용 벤치마크입니다.
검색 품질(hit@k)과 생성 품질(f1·BERTScore)을 같은 데이터로 잽니다.
구성
파일
행 수
용도
dataset.jsonl
1000
질문·정답·정답 청크 id
corpus.jsonl
1370
검색 후보 청크 풀 (KB 적재용)
QA 를 가진 청크는 1000개이고, 나머지는 오답 후보(distractor)로 남습니다 —
검색 난이도를 유지하려면 corpus 를 통째로 적재해야 합니다.
dataset.jsonl
{"question": "금성의 표면 온도는 섭씨 몇도 이상인가?", "answer": "400도 이상", "expected_chunk_ids": ["<chunk_id>"]}
필드
설명
question
질문.… See the full description on the dataset page: https://huggingface.co/datasets/Surromind/rag-qa-benchmark-ko.moc-rag-benchmark
MoC-RAG Benchmark: Typed Context Routing for Agentic Memory
A benchmark for evaluating whether routed, typed context experts
(Mixture-of-Contexts RAG) improve retrieval and answer quality compared with
flat RAG, under a fixed token budget.
Version 0.1.0. This dataset accompanies the paper
Matrix Context: Mixture-of-Contexts RAG for Robust and Inspectable Agent Memory
(10.5281/zenodo.20560139).
Why this benchmark
Flat RAG embeds everything into one index and… See the full description on the dataset page: https://huggingface.co/datasets/ruslanmv/moc-rag-benchmark.kaz-rag-search-benchmark
Kaz-RAG-Search-Benchmark
Evidence-based benchmark for Kazakh information retrieval — the independent proof base
for the Kazakh Stemmer.
Corpus: 8,370 passages from Kazakh Wikipedia
Queries: 300 queries × 3 categories (natural / inflected / vocabulary-gap)
Format: BEIR-compatible — three subsets: corpus, queries, qrels
Browse the data: use the subset switcher at the top of the Data Studio viewer to
move between corpus (Kazakh passages), queries (the 300 questions), and qrels… See the full description on the dataset page: https://huggingface.co/datasets/Tim2190/kaz-rag-search-benchmark.German-RAG-LLM-HARD-BENCHMARK
German-RAG-LLM-HARD Benchmark
German-RAG - German Retrieval Augmented Generation
Dataset Summary
This German-RAG-LLM-HARD-BENCHMARK represents a specialized collection for evaluate language models with a focus on hard to solve RAG-specific capabilities. To evaluate models compatible with OpenAI-Endpoints you can refer to our Github Repo: https://github.com/avemio-digital/GRAG-LLM-HARD-BENCHMARK
The subsets are derived from Synthetic generation inspired by… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-LLM-HARD-BENCHMARK.smartroute-rag-synthetic-routing-benchmark-5000
🧭 SmartRoute-RAG Synthetic Routing Benchmark 5000
A publication-scale benchmark for evaluating when to retrieve — not just what to answer.
5,000 stratified questions · 10 benchmark-style subsets · 13 question types · binary routing labelsBuilt for the SmartRoute-RAG research line: false-skip-aware, safety-constrained adaptive retrieval.
🎯 Why this dataset exists
Most RAG benchmarks measure answer quality after retrieval. They rarely tell you whether the… See the full description on the dataset page: https://huggingface.co/datasets/pr0mila-gh0sh/smartroute-rag-synthetic-routing-benchmark-5000.said-rag-eval-benchmark
SAID RAG Evaluation Benchmark (v1.1)
A 75-cell benchmark of RAG pipeline evaluation outputs with 10 LLM-judged metrics, designed to study unsupervised metric reliability filtering for LLM-judged RAG evaluation.
This is the artifact accompanying the NeurIPS 2026 Evaluations & Datasets Track submission "Some RAG Metrics Don't Measure Quality: Detecting Surface Confounds via Retrieval Invariants" (anonymous review).
What's in this release
v1.1 (current)… See the full description on the dataset page: https://huggingface.co/datasets/said-rag-eval-2026/said-rag-eval-benchmark.plumloom-rag-eval-benchmark
Plumloom RAG Evaluation Benchmark
This dataset packages the public tables from a corpus-boundary study of a
document-grounded RAG agent over NIST SP 800-160 Vol. 1 Rev. 1. The research asks
whether an enterprise team can determine that a document agent is operating
from its designated corpus—and turn observed behavior into reliable release
evidence.
The benchmark creates difficult corpus-boundary tests. The RAG produces the
behavior. EvalEngine turns those executions into… See the full description on the dataset page: https://huggingface.co/datasets/Plumloom/plumloom-rag-eval-benchmark.
