datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ragbench
RAGBench
Dataset Overview
RAGBEnch is a large-scale RAG benchmark dataset of 100k RAG examples.
It covers five unique industry-specific domains and various RAG task types.
RAGBench examples are sourced from industry corpora such as user manuals, making it particularly relevant for industry applications.
RAGBench comrises 12 sub-component datasets, each one split into train/validation/test splits
Usage
from datasets import load_dataset
# load… See the full description on the dataset page: https://huggingface.co/datasets/galileo-ai/ragbench.ragbench-sentence-relevance-balancedt2-ragbench-splitsragbench-ru
New Dataset (Russian Translation)
This dataset is a translation of the original dataset from English to Russian.
Translated by model Qwen2.5-72B-Instruct.
License
The dataset is licensed under the CC BY 4.0.
Original Source
The original dataset can be found at Dataset Source Link.
ragbench-dual-clf-preprocessedRAGBench_Kazakh
RAGBench_Kazakh
Summary
RAGBench_Kazakh is a machine-translated Kazakh version of the original RAGBench benchmark. It is designed to evaluate retrieval-augmented generation (RAG) systems, focusing on how well models use retrieved context to produce grounded answers.
The dataset is built from the test splits of multiple RAGBench subsets covering domains such as biomedical research, general knowledge, legal documents, customer support, and finance. Each example… See the full description on the dataset page: https://huggingface.co/datasets/issai/RAGBench_Kazakh.rag-benchmark-qa-datasetragbench
RAGBench
Dataset Overview
RAGBEnch is a large-scale RAG benchmark dataset of 100k RAG examples.
It covers five unique industry-specific domains and various RAG task types.
RAGBench examples are sourced from industry corpora such as user manuals, making it particularly relevant for industry applications.
RAGBench comrises 12 sub-component datasets, each one split into train/validation/test splits
Usage
from datasets import load_dataset
# load… See the full description on the dataset page: https://huggingface.co/datasets/menggaotian/ragbench.ragbench_10row_tester_synthetic_mistakerag-bench-public-textsPublic RAG bench dataset with texts
ragbench_10row_tester_synthetic_mistake_evaluatedragbench-allrag-bench-public-questionsPublic RAG bench dataset with questions
ragbench_10row_tester_annotatedragbench_10row_testerQwen7b_ragbench_emanual_400row_mistake_addedrag-benchmark-datasetragbench_delucionqa_400row_mistake_addedragbench_emanual_400row_mistake_addedMistral8b_ragbench_delucionqa_400row_mistake_addedrag-bench-private-texts-dragon-upmilDRAGON bench history private texts (mappings). Date: 2025.12.29
rag-bench-private-qa-dragon-upmilDRAGON bench history private QA dataset. Date: 2025.12.29
ragbench-context-relevance-filteredragbench-test-sampleragbench_emanual_400rowragbench_techqa_400row_mistake_addedragbench
RAGBench
Dataset Overview
RAGBEnch is a large-scale RAG benchmark dataset of 100k RAG examples.
It covers five unique industry-specific domains and various RAG task types.
RAGBench examples are sourced from industry corpora such as user manuals, making it particularly relevant for industry applications.
RAGBench comrises 12 sub-component datasets, each one split into train/validation/test splits
Usage
from datasets import load_dataset
# load… See the full description on the dataset page: https://huggingface.co/datasets/suniltvl/ragbench.DeepSeek7b_ragbench_techqa_400row_mistake_addedLlama8b_ragbench_emanual_400row_mistake_addedreasoning_rerankers_ragbench
