datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rag_instruct_benchmark_tester
Dataset Card for RAG-Instruct-Benchmark-Tester
Dataset Summary
This is an updated benchmarking test dataset for "retrieval augmented generation" (RAG) use cases in the enterprise, especially for financial services, and legal. This test dataset includes 200 questions with context passages pulled from common 'retrieval scenarios', e.g., financial news, earnings releases,
contracts, invoices, technical articles, general news and short texts.
The questions are segmented… See the full description on the dataset page: https://huggingface.co/datasets/llmware/rag_instruct_benchmark_tester.kdv-rag-benchmark
KDV RAG Benchmark
A retrieval benchmark dataset for Turkish VAT (KDV, Katma Değer Vergisi) law — built by adding retrieval layers one at a time (chunking, model choice, hybrid search, reranking, query rewriting, historical/date filtering) and statistically validating each one individually (see Results).
Dataset structure
Splits
Split
Records
Period
train
728
2018-2023
test
154
2024-2026
Split strategy: temporal — train and test… See the full description on the dataset page: https://huggingface.co/datasets/dokukoza/kdv-rag-benchmark.rag-hallucination-benchmark
RAG Hallucination Benchmark
Context
Retrieval-Augmented Generation (RAG) is the industry standard for reducing LLM hallucinations, but detecting when a RAG system fails is a massive challenge. Most existing benchmarks focus only on massive Deep Learning models and lack tabular features.
This dataset provides a clean, engineered setup to train models (from XGBoost to RoBERTa) to detect hallucinations, predict context faithfulness, and measure answer relevance.… See the full description on the dataset page: https://huggingface.co/datasets/vkshdev/rag-hallucination-benchmark.enterprise-rag-internal-knowledge-search-benchmark-sample
Enterprise RAG and Internal Knowledge Search Benchmark Dataset -- Free Evaluation Sample
This dataset packages a synthetic internal company workspace and an evidence-linked benchmark table into one product for teams building enterprise RAG systems, internal search assistants, knowledge-base copilots, and deep-search evaluation pipelines.
The benchmark is designed around a realistic fictional company, AsteraOps Cloud, with multiple departments, renamed projects, stale roadmaps… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/enterprise-rag-internal-knowledge-search-benchmark-sample.DOC_RAG_FINANCE_BENCHMARKDOC_RAG_FINANCE_REPORT_BENCHMARKenterprise-rag-internal-knowledge-search-benchmark
Enterprise RAG and Internal Knowledge Search Benchmark Dataset
This dataset packages a synthetic internal company workspace and an evidence-linked benchmark table into one product for teams building enterprise RAG systems, internal search assistants, knowledge-base copilots, and deep-search evaluation pipelines.
The benchmark is designed around a realistic fictional company, AsteraOps Cloud, with multiple departments, renamed projects, stale roadmaps, support escalations… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/enterprise-rag-internal-knowledge-search-benchmark.DOC_RAG_FINANCE_BENCHMARK_FILTEREDDOC_RAG_KRX_BENCHMARKDOC_RAG_PIT_BENCHMARK_FILTEREDDOC_RAG_FINANCE_REPORT_BENCHMARK_FILTEREDDOC_RAG_KRX_BENCHMARK_FILTEREDDOC_RAG_REPORT_BENCHMARKDOC_RAG_REPORT_BENCHMARK_FILTEREDDOC_RAG_PIT_BENCHMARK
