CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CohereLabs /aya_evaluation_suite Dataset Summary Aya Evaluation Suite contains a total of 26,750 open-ended conversation-style prompts to evaluate multilingual open-ended generation quality.To strike a balance between language coverage and the quality that comes with human curation, we create an evaluation suite that includes: human-curated examples in 7 languages (tur, eng, yor, arb, zho, por, tel) → aya-human-annotated. machine-translations of handpicked examples into 101 languages → dolly-machine-translated.… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_evaluation_suite.tabulartext-generation10K<n<100K55 likes6.7k downloads1y agoHugging Face02YixuanEvenXu /HIP-training-and-evaluation-data HIP Training and Evaluation Data This dataset contains the text data released with Base Models Look Human To AI Detectors for reproducing the Humanization by Iterative Paraphrasing (HIP) training setup and the prefix-based continuation evaluation. Configs training data/train.parquet contains 10,581 supervised HIP training pairs with seven columns: dataset: upstream dataset family, either raid or mage. source: selected source domain or subcorpus. text: original… See the full description on the dataset page: https://huggingface.co/datasets/YixuanEvenXu/HIP-training-and-evaluation-data.tabulartext-generation10K<n<100K1 likes358 downloads4mo agoHugging Face03compass-group-tue /sdf_evaluation_traits Models That Know How Evaluations Are Designed Score Safer This repository contains the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated using the… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits.tabulartext-generation10K<n<100K1 likes204 downloads4mo agoHugging Face04furkankarli /turkish-brand-bias-evaluations Turkish Brand Bias Evaluations / Türkçe Marka Yanlılığı Değerlendirmeleri Furkan Karlı tarafından Türkçe ürün ve hizmet önerilerindeki marka görünürlüğünü incelemek amacıyla oluşturulmuş LLM değerlendirme veri setidir. An LLM evaluation dataset curated by Furkan Karlı to study brand visibility in Turkish product and service recommendations. Veri seti özeti 300 tamamlanmış ve judge edilmiş yanıt Domainler: VPN (150) ve kozmetik (150) Koşullar: web araması kapalı… See the full description on the dataset page: https://huggingface.co/datasets/furkankarli/turkish-brand-bias-evaluations.tabulartext-generationn<1K1 likes101 downloads19d agoHugging Face05Experimental-Orange /HumanAgencyBench_Evaluation_Results HumanAgencyBench evaluation results Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Code: https://github.com/BenSturgeon/HumanAgencyBench/ Dataset Description This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.tabulartext-generation10K<n<100K0 likes47 downloads25d agoHugging Face06SahmBenchmark /fatwa-qa-evaluation Fatwa QA Evaluation Dataset Dataset Description This dataset contains Islamic finance and jurisprudence fatwa question-answer pairs for evaluating Arabic language models. This is an open-ended QA evaluation benchmark where models generate free-form answers. Dataset Statistics Total Samples: 2,000 Average Question Length: 243.9 characters Average Answer Length: 492.3 characters Dataset Structure Data Fields id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/fatwa-qa-evaluation.tabularquestion-answering1K<n<10K0 likes43 downloads10mo agoHugging Face07compass-group-tue /sdf_evaluation_traits_15M Models That Know How Evaluations Are Designed Score Safer This repository contains a subset of the synthetic documents used in the paper Models That Know How Evaluations Are Designed Score Safer. Project Page | GitHub Repository Dataset Description These synthetic documents were used to fine-tune models to investigate evaluation meta-knowledge — parametric knowledge about the structural traits that characterize AI safety evaluations. Documents were generated… See the full description on the dataset page: https://huggingface.co/datasets/compass-group-tue/sdf_evaluation_traits_15M.tabulartext-generation10K<n<100K0 likes37 downloads1mo agoHugging Face083RAIN /brand-bias-evaluations Brand Bias in LLM Recommendations Evaluation dataset measuring how 4 frontier LLMs recommend brands/products with and without web search, across 4 consumer domains. Paper: PDF (source)Code: github.com/ThreeRiversAINexus/brand-bias-evaluationsDataset: huggingface.co/datasets/3RAIN/brand-bias-evaluationsContact: Three Rivers AI Nexus LLC — threeriversainexus@gmail.com — for custom evaluations and prompt optimization Quick Start from datasets import load_dataset # Load one… See the full description on the dataset page: https://huggingface.co/datasets/3RAIN/brand-bias-evaluations.tabulartext-generation10K<n<100K0 likes24 downloads6mo agoHugging Face09Legal-verse /sinergi-model-evaluation-results Sinergi model evaluation outputs Model answers and generation telemetry used by the Sinergi Table 8-aligned evaluation notebook. The repository contains eight configurations so every system can be loaded independently with the Hugging Face datasets library. from datasets import load_dataset data = load_dataset( "Legal-verse/sinergi-model-evaluation-results", "qwen-sft-rl-rag", split="test", ) Configurations Configuration System Rows Source file… See the full description on the dataset page: https://huggingface.co/datasets/Legal-verse/sinergi-model-evaluation-results.tabulartext-generation10K<n<100K0 likes13 downloads1mo agoHugging Face10aigencydev /aigency-v4-evaluation AIGENCY V4 — Benchmark Evaluation Results Reproducibility capsule for the AIGENCY V4 whitepaper. 13,344 real API calls · 22 benchmarks · Wilson 95% CI · seed=42. This dataset is the verifiable evidence behind the AIGENCY V4 model card and the AIGENCY V4 whitepaper. Every benchmark folder contains one scored.jsonl (per-item predictions, gold answers, scores) and a summary.json (aggregate accuracy with Wilson 95% CI). What's in this dataset For each of the 22… See the full description on the dataset page: https://huggingface.co/datasets/aigencydev/aigency-v4-evaluation.tabulartext-generation10K<n<100K0 likes12 downloads5mo agoHugging Face11Rupesh2 /Evaluation_GRPOgatedtabulartext-generation1K<n<10K0 likes1 downloads2y agoHugging Face12beatsprom /stateless-mcp-agent-evaluation-suite-2026 ⚡ Stateless Model Context Protocol (MCP 2026) & Agent Evaluation Suite A Production-Grade Corpus & Evaluation Harness for Claude 5, GPT-6 Astra, and DeepSeek V4.1 ⚡ Overview & Industry Problem As of late 2026, autonomous agent engineering has superseded prompt engineering. Enterprises and developers rely on Stateless Model Context Protocol (MCP) to connect reasoning models (Claude 5 Fable, GPT-6 Astra, DeepSeek V4.1-Flash) to production backends.… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/stateless-mcp-agent-evaluation-suite-2026.tabulartext-generation1K<n<10K0 likes5h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.