CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01durgesh-rao /Causal-Intervention-Tests-For-Explanation-Faithfulness Faithfulness via Causal Interventions — Evaluation Pipeline Paper: What Does Answer Change Rate Actually Measure? A Specificity Audit of Causal Intervention Tests for Explanation Faithfulness Accepted at: EMNLP 2026 Workshop GroundLM, Budapest, Hungary (emnlp.org) This pipeline implements the causal-intervention evaluation for LLM explanation faithfulness described in the accompanying paper, including two controls: a content-free specificity check and a decoding-noise floor.… See the full description on the dataset page: https://huggingface.co/datasets/durgesh-rao/Causal-Intervention-Tests-For-Explanation-Faithfulness.text-generation100K<n<1M2 likes196 downloads23d agoHugging Face02ZihengZ /testset Dataset Card for TreeOfLife-10M Captions This dataset consists of generated captions, Wikipedia-derived descriptions and format examples for the TreeOfLife-10M. These captions were generated using InternVL3-38B based on biological contexts that help the model generate more accurate captions. It was used to train BioCAP, a CLIP-based model. Dataset Details This dataset is comprised of captions for the images in TreeOfLife-10M that were generated using InternVL3 38B.… See the full description on the dataset page: https://huggingface.co/datasets/ZihengZ/testset.textimage-classification1M<n<10M0 likes95 downloads11mo agoHugging Face03GIZ /EnDev_RAGAS_testset EnDev RAGAS Test Set Synthetic Q&A test set (42 pairs) generated with RAGAS (TestsetGenerator.generate_with_chunks) over chunks sampled from the EnDev corpus stored in Qdrant collection endev-bgem3-512 (Gradio-gateway Space GIZ/EnDev_Qdrant). Generated: 2026-09-14 UTC Generator/judge LLM: Qwen/Qwen3-235B-A22B-Instruct-2507 Embeddings: BGE-M3 via the EnDev TEI Inference Endpoint Columns: user_input, reference, reference_contexts, synthesizer_name Used to evaluate the deployed… See the full description on the dataset page: https://huggingface.co/datasets/GIZ/EnDev_RAGAS_testset.question-answering0 likes72 downloads9d agoHugging Face04dwb2023 /gdelt-rag-golden-testset-v2 GDELT RAG Golden Test Set Dataset Description This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation) systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the RAGAS framework for synthetic test data generation. Dataset Summary Total Examples: 12 QA pairs Purpose: RAG system evaluation Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v2.textquestion-answeringn<1K0 likes52 downloads11mo agoHugging Face05dwb2023 /gdelt-rag-golden-testset-v3 GDELT RAG Golden Test Set Dataset Description This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation) systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the RAGAS framework for synthetic test data generation. Dataset Summary Total Examples: 12 QA pairs Purpose: RAG system evaluation Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v3.textquestion-answeringn<1K0 likes21 downloads11mo agoHugging Face06dwb2023 /gdelt-rag-golden-testset GDELT RAG Golden Test Set Dataset Description This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation) systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the RAGAS framework for synthetic test data generation. Dataset Summary Total Examples: 12 QA pairs Purpose: RAG system evaluation Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset.textquestion-answeringn<1K0 likes20 downloads7mo agoHugging Face07dwb2023 /gdelt-rag-golden-testset-v4 GDELT RAG Golden Test Set Dataset Description This dataset contains a curated set of question-answering pairs designed for evaluating RAG (Retrieval-Augmented Generation) systems focused on GDELT (Global Database of Events, Language, and Tone) analysis. The dataset was generated using the RAGAS framework for synthetic test data generation. Dataset Summary Total Examples: 12 QA pairs Purpose: RAG system evaluation Framework: RAGAS (Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-golden-testset-v4.textquestion-answeringn<1K0 likes20 downloads11mo agoHugging Face08dwb2023 /ragas-golden-testset-personas Dataset Card for ragas-golden-testset-personas Dataset Description The RAGAS Golden Dataset is a synthetically generated question-answering dataset designed for evaluating Retrieval Augmented Generation (RAG) systems. It contains high-quality question-answer pairs derived from academic papers on AI agents and agentic AI architectures. Dataset Summary This dataset was generated using the RAGAS TestsetGenerator framework, which creates synthetic questions… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/ragas-golden-testset-personas.textquestion-answeringn<1K0 likes18 downloads1y agoHugging Face09open-biosciences /biosciences-golden-testset Biosciences RAG Golden Test Set Dataset Description This dataset contains 12 question-answering pairs for evaluating RAG systems on biomedical research topics. The QA pairs were synthetically generated using the RAGAS framework from 140 source documents spanning knowledge graphs, LLM applications in biomedicine, protein interaction databases, and gene-to-phenotype mapping. Dataset Summary Total Examples: 12 QA pairs Purpose: RAG system evaluation ground truth… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-golden-testset.textquestion-answeringn<1K0 likes17 downloads7mo agoHugging Face10MJannik /test-synthetic-dataset Dataset Card for test-synthetic-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/MJannik/test-synthetic-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/MJannik/test-synthetic-dataset.texttext-generationn<1K0 likes9 downloads1y agoHugging Face11psoldunov /testsetquestion-answeringn>1T0 likes7 downloads2y agoHugging Face12friedahuang /cvpr2019_5papers_testset_12qtextquestion-answeringn<1K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.