CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01onyx-dot-app /EnterpriseRAG-Bench EnterpriseRAG-Bench A benchmark dataset of 500,000+ documents and 500 questions for evaluating RAG systems on realistic enterprise data. See the latest leaderboard rankings. The paper is available on arXiv. For code, methodology, and evaluation tools, see the GitHub repository. Overview Existing RAG and IR datasets focus on publicly accessible document sets (Bing searches, Stack Overflow, etc.). EnterpriseRAG-Bench provides the first publicly accessible dataset… See the full description on the dataset page: https://huggingface.co/datasets/onyx-dot-app/EnterpriseRAG-Bench.textquestion-answeringn<1K14 likes1.1k downloads5mo agoHugging Face02xsong69 /enterpriseRAG-extension EnterpriseRAG Extension for MemOnDemand This dataset is the 1.14B-token EnterpriseRAG collection used to evaluate MemOnDemand. It combines the unchanged EnterpriseRAG-Bench document collection and its 500 evaluation questions with 353,158 newly generated enterprise documents. The resulting corpus contains 865,120 document rows and 1,136,704,992 measured tokens across nine source types. The extension is designed as a scale stress test for retrieval and memory management. It adds… See the full description on the dataset page: https://huggingface.co/datasets/xsong69/enterpriseRAG-extension.textquestion-answering100K<n<1M0 likes103 downloads23d agoHugging Face03SJChen02 /EnterpriseRAG-Bench EnterpriseRAG-Bench A benchmark dataset of 500,000+ documents and 500 questions for evaluating RAG systems on realistic enterprise data. See the latest leaderboard rankings. The paper is available on arXiv. For code, methodology, and evaluation tools, see the GitHub repository. Overview Existing RAG and IR datasets focus on publicly accessible document sets (Bing searches, Stack Overflow, etc.). EnterpriseRAG-Bench provides the first publicly accessible dataset… See the full description on the dataset page: https://huggingface.co/datasets/SJChen02/EnterpriseRAG-Bench.textquestion-answering100K<n<1M0 likes72 downloads3mo agoHugging Face04ingalepratap /EnterpriseRAG-Bench EnterpriseRAG-Bench A benchmark dataset of 500,000+ documents and 500 questions for evaluating RAG systems on realistic enterprise data. See the latest leaderboard rankings. The paper is available on arXiv. For code, methodology, and evaluation tools, see the GitHub repository. Overview Existing RAG and IR datasets focus on publicly accessible document sets (Bing searches, Stack Overflow, etc.). EnterpriseRAG-Bench provides the first publicly accessible dataset… See the full description on the dataset page: https://huggingface.co/datasets/ingalepratap/EnterpriseRAG-Bench.textquestion-answering100K<n<1M0 likes55 downloads4mo agoHugging Face05UCG4879 /EnterpriseRAG-Bench EnterpriseRAG-Bench A benchmark dataset of 500,000+ documents and 500 questions for evaluating RAG systems on realistic enterprise data. See the latest leaderboard rankings. The paper is available on arXiv. For code, methodology, and evaluation tools, see the GitHub repository. Overview Existing RAG and IR datasets focus on publicly accessible document sets (Bing searches, Stack Overflow, etc.). EnterpriseRAG-Bench provides the first publicly accessible dataset… See the full description on the dataset page: https://huggingface.co/datasets/UCG4879/EnterpriseRAG-Bench.textquestion-answering100K<n<1M0 likes49 downloads24d agoHugging Face06shanaka95 /enterprise-rag-questions shanaka95/enterprise-rag-questions A focused training corpus for domain-adapted embedding models targeting retrieval over enterprise documents. Each row is a (doc_id, question) pair where the question is synthetically generated by an instruction-tuned LLM (Gemma-4-12B-it-qat) from a corresponding source document. The dataset is derived from the public benchmark onyx-dot-app/EnterpriseRAG-Bench. For each document in the source corpus we generate five diverse, answerable training… See the full description on the dataset page: https://huggingface.co/datasets/shanaka95/enterprise-rag-questions.textsentence-similarity1M<n<10M0 likes43 downloads2mo agoHugging Face07Karmane /enterprise-rag-internal-knowledge-search-benchmark-sample Enterprise RAG and Internal Knowledge Search Benchmark Dataset -- Free Evaluation Sample This dataset packages a synthetic internal company workspace and an evidence-linked benchmark table into one product for teams building enterprise RAG systems, internal search assistants, knowledge-base copilots, and deep-search evaluation pipelines. The benchmark is designed around a realistic fictional company, AsteraOps Cloud, with multiple departments, renamed projects, stale roadmaps… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/enterprise-rag-internal-knowledge-search-benchmark-sample.tabulartext-generationn<1K2 likes24 downloads4mo agoHugging Face08raghav1424 /EnterpriseRAG-Bench EnterpriseRAG-Bench A benchmark dataset of 500,000+ documents and 500 questions for evaluating RAG systems on realistic enterprise data. See the latest leaderboard rankings. The paper is available on arXiv. For code, methodology, and evaluation tools, see the GitHub repository. Overview Existing RAG and IR datasets focus on publicly accessible document sets (Bing searches, Stack Overflow, etc.). EnterpriseRAG-Bench provides the first publicly accessible dataset… See the full description on the dataset page: https://huggingface.co/datasets/raghav1424/EnterpriseRAG-Bench.textquestion-answering100K<n<1M0 likes22 downloads1mo agoHugging Face09Vidhi43 /EnterpriseRAG-Bench EnterpriseRAG-Bench A benchmark dataset of 500,000+ documents and 500 questions for evaluating RAG systems on realistic enterprise data. See the latest leaderboard rankings. The paper is available on arXiv. For code, methodology, and evaluation tools, see the GitHub repository. Overview Existing RAG and IR datasets focus on publicly accessible document sets (Bing searches, Stack Overflow, etc.). EnterpriseRAG-Bench provides the first publicly accessible dataset… See the full description on the dataset page: https://huggingface.co/datasets/Vidhi43/EnterpriseRAG-Bench.textquestion-answering100K<n<1M0 likes21 downloads1mo agoHugging Face10shanu485 /EnterpriseRAG-Bench EnterpriseRAG-Bench A benchmark dataset of 500,000+ documents and 500 questions for evaluating RAG systems on realistic enterprise data. See the latest leaderboard rankings. The paper is available on arXiv. For code, methodology, and evaluation tools, see the GitHub repository. Overview Existing RAG and IR datasets focus on publicly accessible document sets (Bing searches, Stack Overflow, etc.). EnterpriseRAG-Bench provides the first publicly accessible dataset… See the full description on the dataset page: https://huggingface.co/datasets/shanu485/EnterpriseRAG-Bench.textquestion-answering100K<n<1M0 likes20 downloads5mo agoHugging Face11getomnico /omni-enterprise-rag-bench Omni EnterpriseRAG-Bench Results This dataset contains the public verification artifacts for Omni's EnterpriseRAG-Bench full-500 agentic benchmark run. The benchmark itself is EnterpriseRAG-Bench. This repository does not re-host the full benchmark corpus. It publishes Omni's answer files, evaluation result files, patch manifests, and the merge script needed to verify the reported metrics against the benchmark questions and gold metadata. Contents base/: the clean… See the full description on the dataset page: https://huggingface.co/datasets/getomnico/omni-enterprise-rag-bench.question-answering1 likes20 downloads4mo agoHugging Face12alirezaaminzadeh /enterprise-rag-samples OrgMind Enterprise Policy Samples Synthetic organizational policy documents and QA benchmark pairs for OrgMind RAG Studio. Contents File Description chunks.jsonl Semantic chunks with page/paragraph citation metadata qa_pairs.jsonl Curated questions with expected document + keywords benchmark_report.json Reproducible retrieval metrics eval_results.json Benchmark summary (no per-row details) manifest.json Corpus statistics… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/enterprise-rag-samples.tabularn<1K0 likes13 downloads2mo agoHugging Face13Karmane /enterprise-rag-internal-knowledge-search-benchmarkgated Enterprise RAG and Internal Knowledge Search Benchmark Dataset This dataset packages a synthetic internal company workspace and an evidence-linked benchmark table into one product for teams building enterprise RAG systems, internal search assistants, knowledge-base copilots, and deep-search evaluation pipelines. The benchmark is designed around a realistic fictional company, AsteraOps Cloud, with multiple departments, renamed projects, stale roadmaps, support escalations… See the full description on the dataset page: https://huggingface.co/datasets/Karmane/enterprise-rag-internal-knowledge-search-benchmark.tabulartext-generationn<1K0 likes10 downloads4mo agoHugging Face14beatsprom /enterprise-rag-vector-search-2026 ⚡ Enterprise RAG, GraphRAG & Vector Search Infrastructure Dataset (2023–2026) Sample dataset of 30 audit-verified research papers with 384d PyTorch embeddings. 🛒 Full 1,000 Paper B2B Dataset Available on Gumroad Get the complete 3-year dataset (1,000 papers + GitHub Deep Audit + SQLite/CSV/Parquet + Quickstart Script) on Gumroad: 👉 Get Full 1,000 Dataset on Gumroad ($19 / $39 / $89) tabularfeature-extractionn<1K0 likes10 downloads1mo agoHugging Face15sofiabbbb /enterpriserag-bench-artifacts0 likes3 downloads2mo agoHugging Face16KrambergAI /rag-enterprise-notes0 likes2 downloads4mo agoHugging Face17yuvis /enterprise-rag-index0 likes1 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.