datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hle-gpt-oss-120b-no-python-260222
hle-gpt-oss-120b-no-python-260222
Deep research agent evaluation on rl-rag/hle_text_only (test split).
Results
Metric
Value
pass@4
47.9%
avg@4
26.6%
Trajectory accuracy
26.6% (2292/8632)
Questions
2158
Trajectories
8632 (4 per question)
Avg tool calls
14.5
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-no-python-260222.ragbench
RAGBench
Dataset Overview
RAGBEnch is a large-scale RAG benchmark dataset of 100k RAG examples.
It covers five unique industry-specific domains and various RAG task types.
RAGBench examples are sourced from industry corpora such as user manuals, making it particularly relevant for industry applications.
RAGBench comrises 12 sub-component datasets, each one split into train/validation/test splits
Usage
from datasets import load_dataset
# load… See the full description on the dataset page: https://huggingface.co/datasets/galileo-ai/ragbench.rag-mini-wikipediaIn this huggingface discussion you can share what you used the dataset for.
Derives from https://www.kaggle.com/datasets/rtatman/questionanswer-dataset?resource=download we generated our own subset using generate.py.
jeb-rag
JEB-Bench
Charging the Gate Rent: Measured-Energy Accounting for Adaptive Retrieval-Augmented Generation
⚠️ Status: under construction. Phase 0 (measurement validation) and Phase 1
(index construction) are landing now. The oracle matrix (bench/oracle/) is
populated in Phase 2 and this card will be revised when it is complete. Do not
cite numbers from this repository until the status line says complete.
What this is
The first public per-query × per-configuration… See the full description on the dataset page: https://huggingface.co/datasets/Shanmuk4622/jeb-rag.multihop_qanemotron-cc-v21-Parsed-QA4-Summarization-Qwen3-1.7BRAGTruth-processed
RAGTruth Dataset
Dataset Description
Dataset Summary
The RAGTruth dataset is designed for evaluating hallucinations in text generation models, particularly in retrieval-augmented generation (RAG) contexts. It contains examples of model outputs along with expert annotations indicating whether the outputs contain hallucinations.
Dataset Structure
Each example contains:
A query/question
Context passages
Model output
Hallucination labels (evident… See the full description on the dataset page: https://huggingface.co/datasets/wandb/RAGTruth-processed.browsecomp-gpt-oss-120b-260222
browsecomp-gpt-oss-120b-260222
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
46.8%
avg@4
23.9%
Trajectory accuracy
23.9% (1211/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
26.1
Full conversations
✅
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-gpt-oss-120b-260222.browsecomp-no-scroll-gpt-oss-120b
browsecomp-no-scroll-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
46.0%
avg@4
22.9%
Trajectory accuracy
22.9% (1160/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
27.0
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-no-scroll-gpt-oss-120b.rag-dataset-12000
Retrieval-Augmented Generation (RAG) Dataset 12000
Retrieval-Augmented Generation (RAG) Dataset 12000 is an English dataset designed for RAG-optimized models, built by Neural Bridge AI, and released under Apache license 2.0.
Dataset Description
Dataset Summary
Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by allowing them to consult an external authoritative knowledge base before generating responses. This approach significantly… See the full description on the dataset page: https://huggingface.co/datasets/neural-bridge/rag-dataset-12000.rag_multilingual_training_negatives
How this dataset was made
We trained on chunks sourced from the documents in MADLAD-400 dataset that had been evaluated to contain a higher amount of educational information according to a state-of-the-art LLM.
We took chunks of size 250 tokens, 500 tokens, and 1000 tokens randomly for each document.
We then used these chunks to generate questions and answers based on this text using a state-of-the-art LLM.
Finally, we selected negatives for each chunk using the similarity from the… See the full description on the dataset page: https://huggingface.co/datasets/lightblue/rag_multilingual_training_negatives.multihop_qa_sftbrowsecomp-high-effort-gpt-oss-120b
browsecomp-high-effort-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
44.1%
avg@4
22.9%
Trajectory accuracy
22.9% (1158/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
55.4
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
100
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-high-effort-gpt-oss-120b.hle_text_onlymultihop_qa_sft_doc4096_seq1024_v2raggedrag-mini-bioasqSee here for an updated version without nans in text-corpus.
In this huggingface discussion you can share what you used the dataset for.
Derives from http://participants-area.bioasq.org/Tasks/11b/trainingDataset/ we generated our own subset using generate.py.
RAGognize
RAGognize Dataset Card
Resource
Link
Code
Paper
Demo
In Retrieval-Augmented Generation (RAG), ensuring accuracy and reliability is essential. RAGognize is a dataset created to help researchers and developers study and improve how AI systems use retrieved information. By offering structured examples and natural token-level closed-domain hallucination annotations, it provides a resource for analyzing model behavior and developing methods that might help make RAG… See the full description on the dataset page: https://huggingface.co/datasets/F4biian/RAGognize.browsecomp-qwen35-35b-a3b-think
browsecomp-qwen35-35b-a3b-think
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@4
43.0%
avg@4
24.8%
Trajectory accuracy
24.8% (1258/5064)
Questions
1266
Trajectories
5064 (4 per question)
Avg tool calls
41.1
Full conversations
❌
Model & Setup
Model
Qwen3.5-35B-A3B
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-qwen35-35b-a3b-think.expert-rag-benchmarks
Expert RAG Benchmarks
A unified collection of four expert-level legal RAG benchmarks, exposed as six
named splits and three relational configurations: questions, documents, and
qrels.
The KCL split is named kcl_essay because Hugging Face split identifiers do not
permit hyphens; its source name remains kcl-essay.
Loading
from datasets import load_dataset
repo_id = "jinulee-v/expert-rag-benchmarks"
questions = load_dataset(repo_id, "questions", split="housing")… See the full description on the dataset page: https://huggingface.co/datasets/jinulee-v/expert-rag-benchmarks.ragbench-sentence-relevance-balancedhle-gpt-oss-120b-with-python-260222
hle-gpt-oss-120b-with-python-260222
Deep research agent evaluation on unknown.
Results
Metric
Value
pass@4
39.5%
avg@4
17.5%
Trajectory accuracy
17.4% (1860/10660)
Questions
1350
Trajectories
10660 (4 per question)
Avg tool calls
0.0
Full conversations
❌
Model & Setup
Model
unknown
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domainsNone
Tool Usage
Tool
Calls
%… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-with-python-260222.VDR_ibm-research_REAL-MM-RAG
VDR_ibm-research_REAL-MM-RAG - Overview
Dataset Summary
VDR_ibm-research_REAL-MM-RAG is a multimodal dataset that combines text and image data, and support tasks such as DSE retrieval (RAG).
Dataset Creation
This dataset is a merge and shuffle of the following datasets in the VDR format:
ibm-research/REAL-MM-RAG_TechSlides
ibm-research/REAL-MM-RAG_TechReport
ibm-research/REAL-MM-RAG_FinTabTrainSet
ibm-research/REAL-MM-RAG_FinTabTrainSet_rephrased… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_ibm-research_REAL-MM-RAG.rag-mini-bioasqThis dataset is a subset of a training dataset by the BioASQ Challenge, which is available here.
It is derived from rag-datasets/rag-mini-bioasq.
Modifications include:
filling in missing passages (some of them contained "nan" instead of actual text),
changing relevant_passage_ids' type from string to sequence of ints,
deduplicating the passages (removed 40 duplicates) and fixing the relevant_passage_ids in QAP triplets to point to the corrected, deduplicated passages' ids,
splitting QAP… See the full description on the dataset page: https://huggingface.co/datasets/enelpol/rag-mini-bioasq.synthetic-rag-dataset_v1.0browsecomp-high-effort-full-gpt-oss-120b
browsecomp-high-effort-full-gpt-oss-120b
Deep research agent evaluation on data/browsecomp.jsonl (normal split).
Results
Metric
Value
pass@1
20.9%
avg@1
20.9%
Trajectory accuracy
20.9% (264/1266)
Questions
1266
Trajectories
1266 (1 per question)
Avg tool calls
52.9
Full conversations
✅
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
100
Temperature
0.7
Blocked domains
huggingface.co… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/browsecomp-high-effort-full-gpt-oss-120b.ragdag-results
RAGDAG results
Artefacts from RAGDAG - treating a multi-stage retrieval pipeline as a
structural causal model and computing path-specific effects exactly by freezing
stages, rather than estimating them.
Code: https://github.com/ValerianFourel/RAGDAG
Layout
One directory per collection, named after its ir_datasets id:
<dataset-tag>/
REPORT.md human-readable report incl. the PASS/FAIL verdict
MANIFEST.json provenance: git SHA, code… See the full description on the dataset page: https://huggingface.co/datasets/ValerianFourel/ragdag-results.legal-rag-qa
License & Attribution
MTEB-format derivative of isaacus/legal-rag-qa by Isaacus, itself derived from the Introduction to Criminal Law textbook. Licensed under CC BY-NC-SA-3.0 (same as source); non-commercial, attribution required, ShareAlike.
full-fold-the-rag-parquet-merged0222vqa_plant-disease-classification-merged-dataset
