datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
preference-test-sets
Preference Test Sets
Very few preference datasets have heldout test sets for validation of reward model accuracy results.
In this dataset, we curate the test sets from popular preference datasets into a common schema for easy loading and evaluation.
Anthropic HH (Helpful & Harmless Agent and Red Teaming), test set in full is 8552 samples
Anthropic HHH Alignment (Helpful, Honest, & Harmless), formatted from Big Bench for standalone evaluation.
Learning to summarize, downsampled from… See the full description on the dataset page: https://huggingface.co/datasets/allenai/preference-test-sets.EnDev_RAGAS_testset
EnDev RAGAS Test Set
Synthetic Q&A test set (42 pairs) generated with RAGAS
(TestsetGenerator.generate_with_chunks) over chunks sampled from the EnDev corpus
stored in Qdrant collection endev-bgem3-512 (Gradio-gateway Space
GIZ/EnDev_Qdrant).
Generated: 2026-09-14 UTC
Generator/judge LLM: Qwen/Qwen3-235B-A22B-Instruct-2507
Embeddings: BGE-M3 via the EnDev TEI Inference Endpoint
Columns: user_input, reference, reference_contexts, synthesizer_name
Used to evaluate the deployed… See the full description on the dataset page: https://huggingface.co/datasets/GIZ/EnDev_RAGAS_testset.testset
Dataset Card for TreeOfLife-10M Captions
This dataset consists of generated captions, Wikipedia-derived descriptions and format examples for the TreeOfLife-10M. These captions were generated using InternVL3-38B based on biological contexts that help the model generate more accurate captions. It was used to train BioCAP, a CLIP-based model.
Dataset Details
This dataset is comprised of captions for the images in TreeOfLife-10M that were generated using InternVL3 38B.… See the full description on the dataset page: https://huggingface.co/datasets/ZihengZ/testset.gdelt-rag-golden-testset-v2
GDELT RAG Golden Test Set
Dataset Description
This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation)
systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the
RAGAS framework for synthetic test data generation.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation
Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v2.gdelt-rag-golden-testset-v3
GDELT RAG Golden Test Set
Dataset Description
This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation)
systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the
RAGAS framework for synthetic test data generation.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation
Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v3.biosciences-golden-testset
Biosciences RAG Golden Test Set
Dataset Description
This dataset contains 12 question-answering pairs for evaluating RAG systems on biomedical research topics. The QA pairs were synthetically generated using the RAGAS framework from 140 source documents spanning knowledge graphs, LLM applications in biomedicine, protein interaction databases, and gene-to-phenotype mapping.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation ground truth… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-golden-testset.gdelt-rag-golden-testset-v4
GDELT RAG Golden Test Set
Dataset Description
This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation)
systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the
RAGAS framework for synthetic test data generation.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation
Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v4.RAG_legal_comparison_test_set
Dataset de Evaluación (Test Set) para Sistemas RAG en el Dominio Legal
Este repositorio contiene el conjunto de pruebas (Test Set / Golden Dataset) diseñado específicamente para auditar, evaluar y comparar el rendimiento de diferentes configuraciones de sistemas de Generación Aumentada por Recuperación (RAG) sobre documentación jurídica y administrativa española y europea.
El dataset se ha construido con el propósito de servir de base para métricas de evaluación RAG (como… See the full description on the dataset page: https://huggingface.co/datasets/AingeruBeOr/RAG_legal_comparison_test_set.gdelt-rag-golden-testset
GDELT RAG Golden Test Set
Dataset Description
This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation)
systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the
RAGAS framework for synthetic test data generation.
Dataset Summary
Total Examples: 12 QA pairs
Purpose: RAG system evaluation
Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset.ragas-golden-testset-personas
Dataset Card for ragas-golden-testset-personas
Dataset Description
The RAGAS Golden Dataset is a synthetically generated question-answering dataset designed for evaluating Retrieval Augmented Generation (RAG) systems. It contains high-quality question-answer pairs derived from academic papers on AI agents and agentic AI architectures.
Dataset Summary
This dataset was generated using the RAGAS TestsetGenerator framework, which creates synthetic questions… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/ragas-golden-testset-personas.postvalid-v2-test-set
LongShOTBench (Test Split)
Benchmark accompanying the NeurIPS 2026 submission
"A Benchmark for Omni-Modal Reasoning in Long Videos."
This dataset is shared anonymously to support double-blind review.
Purpose
LongShOTBench evaluates multimodal LLMs on long-form video understanding
across vision, speech, and non-speech audio, using intent-driven questions
and weighted criterion-level rubrics. Intended for evaluation only, not
training.
License
CC BY-NC-SA 4.0.… See the full description on the dataset page: https://huggingface.co/datasets/anonymsubs/postvalid-v2-test-set.testsetbuddhist-scholar-test-set
Vietnamese Buddhist Scholar Test Set
Dataset Description
This dataset contains 1008 Vietnamese question-answer pairs focused on Buddhist teachings and literature. The dataset was created to evaluate chatbots' knowledge and understanding of Buddhist concepts, particularly for Vietnamese-speaking users.
Dataset Details
Dataset Summary
Language: Vietnamese
Task: Question Answering, Chatbot Evaluation
Domain: Buddhism, Religious Studies
Size: 1008… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/buddhist-scholar-test-set.cvpr2019_5papers_testset_12q
