CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jinulee-v /expert-rag-benchmarks Expert RAG Benchmarks A unified collection of four expert-level legal RAG benchmarks, exposed as six named splits and three relational configurations: questions, documents, and qrels. The KCL split is named kcl_essay because Hugging Face split identifiers do not permit hyphens; its source name remains kcl-essay. Loading from datasets import load_dataset repo_id = "jinulee-v/expert-rag-benchmarks" questions = load_dataset(repo_id, "questions", split="housing")… See the full description on the dataset page: https://huggingface.co/datasets/jinulee-v/expert-rag-benchmarks.textquestion-answering1M<n<10M0 likes627 downloads4d agoHugging Face02neoai-inc /Japanese-RAG-Generator-Benchmark Japanese RAG Generator Benchmark: 日本語 RAG における Generator 評価ベンチマーク Japanese RAG Generator Benchmark (J-RAGBench) は日本語RAGにおけるGeneratorに用いるLLMの評価データセットを提供する。 実運用時のRAGに求められる多様な評価カテゴリを同一条件下で評価可能であり、複数の評価カテゴリが同時に出現する問題が含まれるQAデータセットを人手および、補助的にOpenAI API(gpt-4.1-2025-04-14)を用いて構築した。 J-RAGBenchの評価カテゴリ Integration: 2~3文書程度の複数の情報源から適切な根拠を抽出・統合して回答を導く Reasoning: 抽出された情報を踏まえて多段階の推論や数値計算などを実行する Logical: 質問・関連文書間での語彙や表現の差異を解釈し、適切な回答を導く Table:… See the full description on the dataset page: https://huggingface.co/datasets/neoai-inc/Japanese-RAG-Generator-Benchmark.textquestion-answeringn<1K4 likes185 downloads10mo agoHugging Face03silma-ai /silma-rag-qa-benchmark-v1.0 SILMA RAGQA Benchmark Dataset V1.0 SILMA RAGQA is a dataset and benchmark created by silma.ai to assess the effectiveness of Arabic Language Models in Extractive Question Answering tasks, with a specific emphasis on RAG applications The benchmark includes 17 bilingual datasets in Arabic and English, spanning various domains What capabilities does the benchmark test? General Arabic and English QA capabilities Ability to handle short and long contexts Ability to… See the full description on the dataset page: https://huggingface.co/datasets/silma-ai/silma-rag-qa-benchmark-v1.0.textquestion-answering1K<n<10K7 likes120 downloads1y agoHugging Face04emretheus /jev-rag-benchmark Jev RAG Benchmark (English) Frozen-candidate-pool evaluation of TypeSafe Jev 1.13 as the reranking and decision layer of a RAG pipeline, compared with a NVIDIA cross-encoder and with no reranking at all, on English XQuAD and SciFact. Every published run used free tiers (total cost: $0). This repository contains the raw per-query artifacts, per-run reports with paired bootstrap confidence intervals, calibration tables, and a plain-text leaderboard. Reranking (nDCG@10… See the full description on the dataset page: https://huggingface.co/datasets/emretheus/jev-rag-benchmark.question-answering0 likes105 downloads3h agoHugging Face05dokukoza /kdv-rag-benchmark KDV RAG Benchmark A retrieval benchmark dataset for Turkish VAT (KDV, Katma Değer Vergisi) law — built by adding retrieval layers one at a time (chunking, model choice, hybrid search, reranking, query rewriting, historical/date filtering) and statistically validating each one individually (see Results). Dataset structure Splits Split Records Period train 728 2018-2023 test 154 2024-2026 Split strategy: temporal — train and test… See the full description on the dataset page: https://huggingface.co/datasets/dokukoza/kdv-rag-benchmark.tabularquestion-answering1K<n<10K0 likes100 downloads28d agoHugging Face06BAEM1N /Korean-RAG-LLM-Judge-Benchmark Korean RAG LLM-as-Judge Benchmark allganize/RAG-Evaluation-Dataset-KO의 한국어 RAG 300 Q&A 위에 46개 답변 생성 모델의 답변과 LLM-as-Judge 평가 결과를 얹은 데이터셋입니다. 원문 문항·기준 답변·출처 PDF 메타데이터는 allganize 데이터셋을 그대로 따릅니다. 🔗 분석 코드·단계별 실험 보고서: https://github.com/BAEM1N/RAG-Evaluation · 📊 요약 대시보드: https://rag.baeum.ai.kr 🎯 한 줄 결론 RAG 파이프라인 최적화 > 모델 업그레이드. 동일 GPT-5.4를 쓰면서도 검색 파이프라인만 잘 짜면 GPT-5.4-pro(약 10배 비싼 모델)보다 +6.0pp 높은 정확도(0.827) 를 낸다. Pipeline Generator Accuracy 🥇 Cartesian… See the full description on the dataset page: https://huggingface.co/datasets/BAEM1N/Korean-RAG-LLM-Judge-Benchmark.textquestion-answeringn<1K0 likes71 downloads4mo agoHugging Face07vkshdev /rag-hallucination-benchmark RAG Hallucination Benchmark Context Retrieval-Augmented Generation (RAG) is the industry standard for reducing LLM hallucinations, but detecting when a RAG system fails is a massive challenge. Most existing benchmarks focus only on massive Deep Learning models and lack tabular features. This dataset provides a clean, engineered setup to train models (from XGBoost to RoBERTa) to detect hallucinations, predict context faithfulness, and measure answer relevance.… See the full description on the dataset page: https://huggingface.co/datasets/vkshdev/rag-hallucination-benchmark.tabulartext-classification10K<n<100K0 likes59 downloads22d agoHugging Face08avemio /German-RAG-LLM-EASY-BENCHMARK German-RAG-LLM-EASY-BENCHMARK German-RAG - German Retrieval Augmented Generation Dataset Summary This German-RAG-LLM-BENCHMARK represents a specialized collection for evaluating language models with a focus on source citation, time difference stating in RAG-specific tasks. To evaluate models compatible with OpenAI-Endpoints you can refer to our Github Repo: https://github.com/avemio-digital/German-RAG-LLM-EASY-BENCHMARK/ Most of the Subsets are synthetically… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-LLM-EASY-BENCHMARK.texttext-classification1K<n<10K0 likes51 downloads2y agoHugging Face09PatSnap /Hiro-Pharma-RAG-Benchmark Hiro Pharma RAG Benchmark This private dataset repository contains multilingual biomedical RAG benchmark data associated with the paper CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine. The benchmark is designed to evaluate whether retrieval-augmented language models can answer biomedical questions while selecting and citing useful evidence and filtering out noisy or irrelevant references. Repository Contents File Language… See the full description on the dataset page: https://huggingface.co/datasets/PatSnap/Hiro-Pharma-RAG-Benchmark.textquestion-answeringn<1K2 likes51 downloads2mo agoHugging Face10PeterAM4 /priviledge-prospectus-rag-benchmark PrivilEdge Prospectus RAG Benchmark 40 question-answer pairs over the PrivilEdge Prospectus for Switzerland (Feb 2026, 354 pages) for evaluating RAG systems with page-level citation accuracy. Source Document The source PDF (LO_fund_fr.pdf) is included in this repository. It is a publicly available regulatory document: the partial prospectus for Swiss investors of PrivilEdge, a Luxembourg SICAV managed by Lombard Odier Funds (Europe) S.A. Benchmark Design… See the full description on the dataset page: https://huggingface.co/datasets/PeterAM4/priviledge-prospectus-rag-benchmark.documentquestion-answeringn<1K0 likes40 downloads6mo agoHugging Face11Surromind /rag-qa-benchmark-ko Surromind/rag-qa-benchmark-ko 한국어 위키백과에서 생성한 RAG 검색·생성 동시 평가용 벤치마크입니다. 검색 품질(hit@k)과 생성 품질(f1·BERTScore)을 같은 데이터로 잽니다. 구성 파일 행 수 용도 dataset.jsonl 1000 질문·정답·정답 청크 id corpus.jsonl 1370 검색 후보 청크 풀 (KB 적재용) QA 를 가진 청크는 1000개이고, 나머지는 오답 후보(distractor)로 남습니다 — 검색 난이도를 유지하려면 corpus 를 통째로 적재해야 합니다. dataset.jsonl {"question": "금성의 표면 온도는 섭씨 몇도 이상인가?", "answer": "400도 이상", "expected_chunk_ids": ["<chunk_id>"]} 필드 설명 question 질문.… See the full description on the dataset page: https://huggingface.co/datasets/Surromind/rag-qa-benchmark-ko.textquestion-answering1K<n<10K0 likes37 downloads1mo agoHugging Face12ruslanmv /moc-rag-benchmark MoC-RAG Benchmark: Typed Context Routing for Agentic Memory A benchmark for evaluating whether routed, typed context experts (Mixture-of-Contexts RAG) improve retrieval and answer quality compared with flat RAG, under a fixed token budget. Version 0.1.0. This dataset accompanies the paper Matrix Context: Mixture-of-Contexts RAG for Robust and Inspectable Agent Memory (10.5281/zenodo.20560139). Why this benchmark Flat RAG embeds everything into one index and… See the full description on the dataset page: https://huggingface.co/datasets/ruslanmv/moc-rag-benchmark.question-answering1K<n<10K0 likes33 downloads4mo agoHugging Face13Tim2190 /kaz-rag-search-benchmark Kaz-RAG-Search-Benchmark Evidence-based benchmark for Kazakh information retrieval — the independent proof base for the Kazakh Stemmer. Corpus: 8,370 passages from Kazakh Wikipedia Queries: 300 queries × 3 categories (natural / inflected / vocabulary-gap) Format: BEIR-compatible — three subsets: corpus, queries, qrels Browse the data: use the subset switcher at the top of the Data Studio viewer to move between corpus (Kazakh passages), queries (the 300 questions), and qrels… See the full description on the dataset page: https://huggingface.co/datasets/Tim2190/kaz-rag-search-benchmark.textquestion-answering1K<n<10K0 likes23 downloads3mo agoHugging Face14avemio /German-RAG-LLM-HARD-BENCHMARK German-RAG-LLM-HARD Benchmark German-RAG - German Retrieval Augmented Generation Dataset Summary This German-RAG-LLM-HARD-BENCHMARK represents a specialized collection for evaluate language models with a focus on hard to solve RAG-specific capabilities. To evaluate models compatible with OpenAI-Endpoints you can refer to our Github Repo: https://github.com/avemio-digital/GRAG-LLM-HARD-BENCHMARK The subsets are derived from Synthetic generation inspired by… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-LLM-HARD-BENCHMARK.textquestion-answeringn<1K0 likes22 downloads2y agoHugging Face15pr0mila-gh0sh /smartroute-rag-synthetic-routing-benchmark-5000 🧭 SmartRoute-RAG Synthetic Routing Benchmark 5000 A publication-scale benchmark for evaluating when to retrieve — not just what to answer. 5,000 stratified questions · 10 benchmark-style subsets · 13 question types · binary routing labelsBuilt for the SmartRoute-RAG research line: false-skip-aware, safety-constrained adaptive retrieval. 🎯 Why this dataset exists Most RAG benchmarks measure answer quality after retrieval. They rarely tell you whether the… See the full description on the dataset page: https://huggingface.co/datasets/pr0mila-gh0sh/smartroute-rag-synthetic-routing-benchmark-5000.textquestion-answering1K<n<10K0 likes20 downloads2mo agoHugging Face16said-rag-eval-2026 /said-rag-eval-benchmark SAID RAG Evaluation Benchmark (v1.1) A 75-cell benchmark of RAG pipeline evaluation outputs with 10 LLM-judged metrics, designed to study unsupervised metric reliability filtering for LLM-judged RAG evaluation. This is the artifact accompanying the NeurIPS 2026 Evaluations & Datasets Track submission "Some RAG Metrics Don't Measure Quality: Detecting Surface Confounds via Retrieval Invariants" (anonymous review). What's in this release v1.1 (current)… See the full description on the dataset page: https://huggingface.co/datasets/said-rag-eval-2026/said-rag-eval-benchmark.question-answering100K<n<1M0 likes15 downloads4mo agoHugging Face17Plumloom /plumloom-rag-eval-benchmark Plumloom RAG Evaluation Benchmark This dataset packages the public tables from a corpus-boundary study of a document-grounded RAG agent over NIST SP 800-160 Vol. 1 Rev. 1. The research asks whether an enterprise team can determine that a document agent is operating from its designated corpus—and turn observed behavior into reliable release evidence. The benchmark creates difficult corpus-boundary tests. The RAG produces the behavior. EvalEngine turns those executions into… See the full description on the dataset page: https://huggingface.co/datasets/Plumloom/plumloom-rag-eval-benchmark.question-answering0 likes15 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.