datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/grasson/t2-ragbench.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/tomsummerfield/t2-ragbench.LIT-RAGBench
LIT-RAGBench
LIT-RAGBench is a benchmark for evaluating generator capabilities in Retrieval-Augmented Generation (RAG). It focuses on whether a model can answer questions correctly given retrieved documents, independent of retrieval quality. The benchmark covers five categories: Integration, Reasoning, Logic, Table, and Abstention.
Dataset Summary
LIT-RAGBench contains:
114 human-constructed Japanese questions
An English version generated by machine translation with… See the full description on the dataset page: https://huggingface.co/datasets/neoai-inc/LIT-RAGBench.ragbench-corpus
RAGBench Corpus
A small, focused document corpus designed for evaluating Retrieval-Augmented Generation (RAG) systems and comparing different retrieval and document chunking strategies.
Dataset Description
RAGBench Corpus contains 20 short documents covering concepts related to modern information retrieval and RAG systems.
The corpus is designed to be used together with the RAGBench Queries dataset to benchmark retrieval performance.
Topics covered include:
Dense… See the full description on the dataset page: https://huggingface.co/datasets/Gul55555/ragbench-corpus.ragbench-queries
RAGBench Queries
A collection of evaluation queries designed for benchmarking Retrieval-Augmented Generation (RAG) systems.
Dataset Description
RAGBench Queries contains test queries used to evaluate different retrieval and chunking strategies in a RAG pipeline.
The dataset was created as part of the RAGBench project, which compares retrieval performance using different document chunking approaches.
Purpose
The queries are designed to evaluate… See the full description on the dataset page: https://huggingface.co/datasets/Gul55555/ragbench-queries.AM-RAGBench
AM-RAGBench
Human-verified Arabic-Malay benchmark for evaluating retrieval-augmented generation (RAG) faithfulness. 1,140 question-answer pairs spanning a specialized domain (Quran, Arabic and Basmeih Malay translation) and a general domain (Arabic and Malay Wikipedia), each with a gold passage, a gold answer, and a verification decision made during construction.
Files
quran_verified.jsonl: specialized-domain records.
wiki_verified.jsonl: general-domain records.… See the full description on the dataset page: https://huggingface.co/datasets/Akram98/AM-RAGBench.
