datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.github-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos)
github-reposThe entire dump of GitHub repositories.
legal-rag-bench
Legal RAG Bench ⚖️
Legal RAG Bench by Isaacus is a reasoning-intensive benchmark for assessing the end-to-end, real-world performance of production-grade legal RAG systems.
Legal RAG Bench is composed of 4,876 passages sampled from the Judicial College of Victoria’s Criminal Charge Book alongside 100 complex, handwritten questions demanding expert-level knowledge of Victorian criminal law and procedure to be answered correctly.
Legal RAG Bench is the first open dataset for the… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/legal-rag-bench.TechQA-RAG-Eval
Dataset Description:
TechQA-RAG-Eval is a reduced version of the original TechQA (IBM’s GitHub Page, HuggingFace) dataset specifically for evaluating Retrieval-Augmented Generation (RAG) systems. The dataset consists of technical support questions and their answers, sourced from real IBM developer forums where acceptable answers included links to reference technical documentation.
This dataset is ready for commercial/non-commercial use.
Dataset Owner(s):
NVIDIA… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/TechQA-RAG-Eval.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.rag_hallucinationsProvides examples of hallucinated responses for RAG applications.
hle_rlvr_no_promptspoken-multihop-rag
Spoken Multi-hop QA: ASR Transcripts Across Four English Accents
ASR transcriptions of 3,000 multi-hop QA questions, each spoken in four
English accents and transcribed with Whisper-large-v3. Released as the
data companion to Better Retrieval, Worse Robustness: How Multi-hop RAG
Amplifies Upstream ASR Errors
(EMNLP 2026, Main Conference).
The dataset exists to make one thing cheap to study: what happens to a
retrieval pipeline when its query arrives through ASR rather than as… See the full description on the dataset page: https://huggingface.co/datasets/orcarouter/spoken-multihop-rag.rag-bench
Dataset card for RAG-BENCH
Data Summary
RAG-bench aims to provide results of many commonly used RAG datasets. All the results in this dataset are evaluated by the RAG evaluation tool Rageval, which could be easily reproduced with the tool.
Currently, we have provided the results of ASQA dataset,ELI5 dataset and HotPotQA dataset.
Data Instance
ASQA
{
"ambiguous_question":"Who is the original artist of sound of silence?",
"qa_pairs":[{… See the full description on the dataset page: https://huggingface.co/datasets/golaxy/rag-bench.rag-retrieval-debug-trajectories
Rag Retrieval Debug Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/rag-retrieval-debug-trajectories.rag-corpus-v1
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
SZLHOLDINGS/rag-corpus-v1 — Agentic-RAG corpus + per-organ FAISS indexes
Doctrine v10/v11. Embedding model: BAAI/bge-base-en-v1.5 (768-dim).
Built by the agentic-RAG SHIP directive (390_AGENTIC_RAG_FAISS_PER_SPACE).
Contents
corpus.jsonl — 762 chunks, each ~512 tokens with… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/rag-corpus-v1.RAG-RewardBenchThis repository contains the data presented in RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment.
Code: https://github.com/jinzhuoran/RAG-RewardBench/
stackoverflow-postsThe StackOverflow posts retrieval source for code-rag-bench.
German-RAG-SFT-ShareGPT-HESSIAN-AI
German-RAG-SFT (Supervised Fine-Tuning) Share-GPT Format
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The SFT Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities. Most tasks were developed using synthetically enhanced data derived from the German Wikipedia, accessed through Cohere's dataset (wikipedia-22-12-de-embeddings). The data is structured in a training knowledge… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-SFT-ShareGPT-HESSIAN-AI.drtulu_v2_step_35_sft_web_searchragtopia_oldMulti-doc-2025
Dataset Card for Multi-Doc-2025
Dataset Summary
Multi-Doc-2025 is a financial question-answering benchmark built from SEC Form 10-K annual reports of S&P 500 companies. It is designed to evaluate retrieval-augmented generation (RAG) and financial QA systems under three reasoning settings that are not jointly covered by existing financial QA benchmarks: cross-company reasoning, cross-year reasoning, and hybrid-modal reasoning over both text and tables.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-Team-HC-RAG/Multi-doc-2025.ragtime1
RAGTIME1 Collection
This dataset contains the documents for TREC RAGTIME Track.
Please refer to the website for the details of the task.
RAGTIME is a multilingual RAG task, which expects the participating system to retrieve relevant documents from all four languages and synthesize a response with citation to the report request.
For convenience, we separate the documents by their languages into four .jsonl files. However, they are intended to be used as a whole set.
The documents… See the full description on the dataset page: https://huggingface.co/datasets/trec-ragtime/ragtime1.rag_instruct_benchmark_tester
Dataset Card for RAG-Instruct-Benchmark-Tester
Dataset Summary
This is an updated benchmarking test dataset for "retrieval augmented generation" (RAG) use cases in the enterprise, especially for financial services, and legal. This test dataset includes 200 questions with context passages pulled from common 'retrieval scenarios', e.g., financial news, earnings releases,
contracts, invoices, technical articles, general news and short texts.
The questions are segmented… See the full description on the dataset page: https://huggingface.co/datasets/llmware/rag_instruct_benchmark_tester.RAG-v1
Glaive-RAG-v1
Glaive-RAG-v1 is a dataset with ~50k samples built using the Glaive platform, for finetuning models for RAG use cases.
Each row has:
List of documents for context
Question
Answer Mode
Answer
The answer mode is to define if the model should output only grounded responses or if it should combine it's internal information as well.
The answers have Cited documents at the beginning and also <co: 1> tags in the text to mark citations.
To report any problems or suggestions… See the full description on the dataset page: https://huggingface.co/datasets/glaiveai/RAG-v1.simutrade-rag-sft-28k
📢 Domain & Email Migration Notice
From May 6th, 2026, Simutrade will transition to new domains as simutrade.app will not be renewed:
🌐 Website: simutrade.faizath.com (formerly simutrade.app)
⚙️ API: simutrade-api.faizath.com (formerly api.simutrade.app)
📧 Email: contact@simutrade.faizath.com (formerly contact@simutrade.app)
🛰️ CDN: simutrade-cdn.faizath.com (formerly cdn.simutrade.app)
📈 Status Pages:… See the full description on the dataset page: https://huggingface.co/datasets/simutrade/simutrade-rag-sft-28k.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/grasson/t2-ragbench.analogical_math_rag_results_3ragtopiaGerman-RAG-SFT-Alpaca-HESSIAN-AI
German-RAG-SFT (Supervised Fine-Tuning) Alpaca-Format
German-RAG - German Retrieval Augmented Generation
Dataset Summary
The SFT Tasks Dataset represents a specialized collection for fine-tuning language models with a focus on RAG-specific capabilities. Most tasks were developed using synthetically enhanced data derived from the German Wikipedia, accessed through Cohere's dataset (wikipedia-22-12-de-embeddings). The data is structured in a training knowledge… See the full description on the dataset page: https://huggingface.co/datasets/avemio/German-RAG-SFT-Alpaca-HESSIAN-AI.CommonCrawl-RAG-QA-Calm3-22b-chat
自動生成テキスト
データソースから、OpenCalm3-22bを使ってクリーニング・再生成したテキストです。
Common Crawlをもとに生成しています。 Common Crawl terms of useに従ってご利用ください。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
データ
jsonlファイルが数十GB程度あります
datasetsライブラリからでは、はじめの数GB程度しか読み込めない可能性があります。git lfsなどでダウンロードする必要がありそうです。
RAG-Instruct
Introduction
RAG-Instruct is a RAG dataset designed to comprehensively enhance LLM RAG capabilities, synthesized using GPT-4o. This dataset is based on the Wikipedia corpus and This dataset is based on the Wikipedia corpus and offers the advantages of query-document scenario diversity and task diversity.
The RAG-Instruct dataset can significantly enhance the RAG ability of LLMs and make remarkable improvements in RAG performance across various tasks.
Model
WQA (acc)
PQA (acc)… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/RAG-Instruct.Japanese-RAG-Generator-Benchmark
Japanese RAG Generator Benchmark: 日本語 RAG における Generator 評価ベンチマーク
Japanese RAG Generator Benchmark (J-RAGBench) は日本語RAGにおけるGeneratorに用いるLLMの評価データセットを提供する。
実運用時のRAGに求められる多様な評価カテゴリを同一条件下で評価可能であり、複数の評価カテゴリが同時に出現する問題が含まれるQAデータセットを人手および、補助的にOpenAI API(gpt-4.1-2025-04-14)を用いて構築した。
J-RAGBenchの評価カテゴリ
Integration: 2~3文書程度の複数の情報源から適切な根拠を抽出・統合して回答を導く
Reasoning: 抽出された情報を踏まえて多段階の推論や数値計算などを実行する
Logical: 質問・関連文書間での語彙や表現の差異を解釈し、適切な回答を導く
Table:… See the full description on the dataset page: https://huggingface.co/datasets/neoai-inc/Japanese-RAG-Generator-Benchmark.analogical_math_rag_results_4
