enterprise-rag
EnterpriseRAG-Bench
EnterpriseRAG-Bench
A benchmark dataset of 500,000+ documents and 500 questions for evaluating RAG systems on realistic enterprise data.
See the latest leaderboard rankings. The paper is available on arXiv. For code, methodology, and evaluation tools, see the GitHub repository.
Overview
Existing RAG and IR datasets focus on publicly accessible document sets (Bing searches, Stack Overflow, etc.). EnterpriseRAG-Bench provides the first publicly accessible dataset… See the full description on the dataset page: https://huggingface.co/datasets/onyx-dot-app/EnterpriseRAG-Bench.enterpriseRAG-extension
EnterpriseRAG Extension for MemOnDemand
This dataset is the 1.14B-token EnterpriseRAG collection used to evaluate
MemOnDemand. It combines the unchanged
EnterpriseRAG-Bench document collection and its 500 evaluation questions with
353,158 newly generated enterprise documents. The resulting corpus contains
865,120 document rows and 1,136,704,992 measured tokens across nine source
types.
The extension is designed as a scale stress test for retrieval and memory
management. It adds… See the full description on the dataset page: https://huggingface.co/datasets/xsong69/enterpriseRAG-extension.EnterpriseRAG-Bench
EnterpriseRAG-Bench
A benchmark dataset of 500,000+ documents and 500 questions for evaluating RAG systems on realistic enterprise data.
See the latest leaderboard rankings. The paper is available on arXiv. For code, methodology, and evaluation tools, see the GitHub repository.
Overview
Existing RAG and IR datasets focus on publicly accessible document sets (Bing searches, Stack Overflow, etc.). EnterpriseRAG-Bench provides the first publicly accessible dataset… See the full description on the dataset page: https://huggingface.co/datasets/SJChen02/EnterpriseRAG-Bench.EnterpriseRAG-Bench
EnterpriseRAG-Bench
A benchmark dataset of 500,000+ documents and 500 questions for evaluating RAG systems on realistic enterprise data.
See the latest leaderboard rankings. The paper is available on arXiv. For code, methodology, and evaluation tools, see the GitHub repository.
Overview
Existing RAG and IR datasets focus on publicly accessible document sets (Bing searches, Stack Overflow, etc.). EnterpriseRAG-Bench provides the first publicly accessible dataset… See the full description on the dataset page: https://huggingface.co/datasets/ingalepratap/EnterpriseRAG-Bench.EnterpriseRAG-Bench
EnterpriseRAG-Bench
A benchmark dataset of 500,000+ documents and 500 questions for evaluating RAG systems on realistic enterprise data.
See the latest leaderboard rankings. The paper is available on arXiv. For code, methodology, and evaluation tools, see the GitHub repository.
Overview
Existing RAG and IR datasets focus on publicly accessible document sets (Bing searches, Stack Overflow, etc.). EnterpriseRAG-Bench provides the first publicly accessible dataset… See the full description on the dataset page: https://huggingface.co/datasets/UCG4879/EnterpriseRAG-Bench.enterprise-rag-questions
shanaka95/enterprise-rag-questions
A focused training corpus for domain-adapted embedding models targeting retrieval over
enterprise documents. Each row is a (doc_id, question) pair where the question is
synthetically generated by an instruction-tuned LLM (Gemma-4-12B-it-qat) from a
corresponding source document.
The dataset is derived from the public benchmark
onyx-dot-app/EnterpriseRAG-Bench.
For each document in the source corpus we generate five diverse, answerable training… See the full description on the dataset page: https://huggingface.co/datasets/shanaka95/enterprise-rag-questions.
