CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /preference-test-sets Preference Test Sets Very few preference datasets have heldout test sets for validation of reward model accuracy results. In this dataset, we curate the test sets from popular preference datasets into a common schema for easy loading and evaluation. Anthropic HH (Helpful & Harmless Agent and Red Teaming), test set in full is 8552 samples Anthropic HHH Alignment (Helpful, Honest, & Harmless), formatted from Big Bench for standalone evaluation. Learning to summarize, downsampled from… See the full description on the dataset page: https://huggingface.co/datasets/allenai/preference-test-sets.textsummarization10K<n<100K28 likes3.5k downloads3y agoHugging Face02ZihengZ /testset Dataset Card for TreeOfLife-10M Captions This dataset consists of generated captions, Wikipedia-derived descriptions and format examples for the TreeOfLife-10M. These captions were generated using InternVL3-38B based on biological contexts that help the model generate more accurate captions. It was used to train BioCAP, a CLIP-based model. Dataset Details This dataset is comprised of captions for the images in TreeOfLife-10M that were generated using InternVL3 38B.… See the full description on the dataset page: https://huggingface.co/datasets/ZihengZ/testset.textimage-classification1M<n<10M0 likes84 downloads11mo agoHugging Face03dwb2023 /gdelt-rag-golden-testset-v2 GDELT RAG Golden Test Set Dataset Description This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation) systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the RAGAS framework for synthetic test data generation. Dataset Summary Total Examples: 12 QA pairs Purpose: RAG system evaluation Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v2.textquestion-answeringn<1K0 likes52 downloads11mo agoHugging Face04dwb2023 /gdelt-rag-golden-testset-v3 GDELT RAG Golden Test Set Dataset Description This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation) systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the RAGAS framework for synthetic test data generation. Dataset Summary Total Examples: 12 QA pairs Purpose: RAG system evaluation Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v3.textquestion-answeringn<1K0 likes22 downloads11mo agoHugging Face05open-biosciences /biosciences-golden-testset Biosciences RAG Golden Test Set Dataset Description This dataset contains 12 question-answering pairs for evaluating RAG systems on biomedical research topics. The QA pairs were synthetically generated using the RAGAS framework from 140 source documents spanning knowledge graphs, LLM applications in biomedicine, protein interaction databases, and gene-to-phenotype mapping. Dataset Summary Total Examples: 12 QA pairs Purpose: RAG system evaluation ground truth… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-golden-testset.textquestion-answeringn<1K0 likes20 downloads7mo agoHugging Face06dwb2023 /gdelt-rag-golden-testset-v4 GDELT RAG Golden Test Set Dataset Description This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation) systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the RAGAS framework for synthetic test data generation. Dataset Summary Total Examples: 12 QA pairs Purpose: RAG system evaluation Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v4.textquestion-answeringn<1K0 likes19 downloads11mo agoHugging Face07AingeruBeOr /RAG_legal_comparison_test_set Dataset de Evaluación (Test Set) para Sistemas RAG en el Dominio Legal Este repositorio contiene el conjunto de pruebas (Test Set / Golden Dataset) diseñado específicamente para auditar, evaluar y comparar el rendimiento de diferentes configuraciones de sistemas de Generación Aumentada por Recuperación (RAG) sobre documentación jurídica y administrativa española y europea. El dataset se ha construido con el propósito de servir de base para métricas de evaluación RAG (como… See the full description on the dataset page: https://huggingface.co/datasets/AingeruBeOr/RAG_legal_comparison_test_set.textquestion-answeringn<1K0 likes18 downloads3mo agoHugging Face08dwb2023 /gdelt-rag-golden-testset GDELT RAG Golden Test Set Dataset Description This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation) systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the RAGAS framework for synthetic test data generation. Dataset Summary Total Examples: 12 QA pairs Purpose: RAG system evaluation Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset.textquestion-answeringn<1K0 likes15 downloads7mo agoHugging Face09dwb2023 /ragas-golden-testset-personas Dataset Card for ragas-golden-testset-personas Dataset Description The RAGAS Golden Dataset is a synthetically generated question-answering dataset designed for evaluating Retrieval Augmented Generation (RAG) systems. It contains high-quality question-answer pairs derived from academic papers on AI agents and agentic AI architectures. Dataset Summary This dataset was generated using the RAGAS TestsetGenerator framework, which creates synthetic questions… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/ragas-golden-testset-personas.textquestion-answeringn<1K0 likes13 downloads1y agoHugging Face10anonymsubs /postvalid-v2-test-set LongShOTBench (Test Split) Benchmark accompanying the NeurIPS 2026 submission "A Benchmark for Omni-Modal Reasoning in Long Videos." This dataset is shared anonymously to support double-blind review. Purpose LongShOTBench evaluates multimodal LLMs on long-form video understanding across vision, speech, and non-speech audio, using intent-driven questions and weighted criterion-level rubrics. Intended for evaluation only, not training. License CC BY-NC-SA 4.0.… See the full description on the dataset page: https://huggingface.co/datasets/anonymsubs/postvalid-v2-test-set.textquestion-answering1K<n<10K0 likes10 downloads5mo agoHugging Face11vanloc1808 /buddhist-scholar-test-set Vietnamese Buddhist Scholar Test Set Dataset Description This dataset contains 1008 Vietnamese question-answer pairs focused on Buddhist teachings and literature. The dataset was created to evaluate chatbots' knowledge and understanding of Buddhist concepts, particularly for Vietnamese-speaking users. Dataset Details Dataset Summary Language: Vietnamese Task: Question Answering, Chatbot Evaluation Domain: Buddhism, Religious Studies Size: 1008… See the full description on the dataset page: https://huggingface.co/datasets/vanloc1808/buddhist-scholar-test-set.textquestion-answeringn<1K0 likes7 downloads1y agoHugging Face12friedahuang /cvpr2019_5papers_testset_12qtextquestion-answeringn<1K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.