CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zjunlp /Chat2Workflow-Evaluation Chat2Workflow Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions. Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language Repository: zjunlp/Chat2Workflow Overview Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.documenttext-generationn<1K4 likes3.1k downloads4mo agoHugging Face02compass-group-tue /sdf_evaluation_traits Models That Know How Evaluations Are Designed Score Safer This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.tabulartext-generation10K<n<100K1 likes204 downloads4mo agoHugging Face03jizej /Competence-Based-Evaluation Competence-Based Evaluation (Invariance Benchmark) A benchmark for testing whether language models give the same answer to semantically equivalent reformulations of a logical-ordering question. Given a set of pairwise constraints (e.g. Alice is in front of Bob), a model should answer transitive-closure queries (Is Carol in front of Dave?) consistently whether the constraints are stated using a relation or its inverse. Each item exists as a paired (original, equivalent) record… See the full description on the dataset page: https://huggingface.co/datasets/jizej/Competence-Based-Evaluation.textquestion-answering100K<n<1M0 likes128 downloads5mo agoHugging Face04AnjanSB /NQ-RAG-DPO-Evaluation Dataset Card Dataset Summary This repository contains five structured data files forming a complete evaluation and training workflow for a multi-perspective alignment system built on Retrieval-Augmented Generation (RAG) and Direct Preference Optimization (DPO). The system is organized into three interconnected pipelines: 1️. RAG Pipeline The RAG pipeline uses the Natural Questions (validation split) as the knowledge source and evaluation benchmark. For each… See the full description on the dataset page: https://huggingface.co/datasets/AnjanSB/NQ-RAG-DPO-Evaluation.texttext-generation1K<n<10K1 likes49 downloads7mo agoHugging Face05compass-group-tue /sdf_evaluation_traits_15M Models That Know How Evaluations Are Designed Score Safer This repository contains a subset of the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits_15M.tabulartext-generation10K<n<100K0 likes37 downloads1mo agoHugging Face06jayzou3773 /less-is-moe-gpqa-diamond-evaluation Less-is-MoE GPQA-Diamond evaluation set This private dataset stores the 198-question GPQA-Diamond evaluation file used by the MoE-Honing evaluation format. Upstream source: Idavidrein/gpqa, config gpqa_diamond Upstream revision: 633f5ee89ab8ad4522a9f850766b73f62147ffdd Split: test Rows: 198 SHA-256: d5b0d6dad6c1993a8cb17fd7aa635fbc5e8b684ae42b5b6e80d15467b399eec3 Fields: problem, solution, domain The problem field contains the formatted four-choice prompt, solution stores the… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/less-is-moe-gpqa-diamond-evaluation.textquestion-answeringn<1K0 likes35 downloads4d agoHugging Face07taln-ls2n /keyphrase_homogeneity_evaluation license: cc-by-nc-4.0 language: - en size_categories: - n<1K Data pairs used in the evaluation of the paper "[Evaluating the Homogeneity of Keyphrase Prediction Models]"(https://arxiv.org/abs/2602.12989), Maël Houbre, Florian Boudin and Béatrice Daille, LREC 2026 texttext-generation10K<n<100K0 likes29 downloads7mo agoHugging Face08alexjk1m /diet-planning-evaluation-20250531-140238 Diet Planning Evaluation Results Dataset Summary This dataset contains evaluation results for diet planning model responses. Supported Tasks and Leaderboards [More Information Needed] Languages English Dataset Structure Data Instances Each instance contains a model response and its evaluation metrics. Data Fields Full Prompt: object Model Response: object Desired Response: object Normalized_Edit_Distance: float64… See the full description on the dataset page: https://huggingface.co/datasets/alexjk1m/diet-planning-evaluation-20250531-140238.texttext-generationn<1K0 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.