CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Paul /XSTest XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models XSTest is a test suite designed to identify exaggerated safety / false refusal in Large Language Models (LLMs). It comprises 250 safe prompts across 10 different prompt types, along with 200 unsafe prompts as contrasts. The test suite aims to evaluate how well LLMs balance being helpful with being harmless by testing if they unnecessarily refuse to answer safe prompts that superficially… See the full description on the dataset page: https://huggingface.co/datasets/Paul/XSTest.texttext-generationn<1K5 likes4.7k downloads2y agoHugging Face02jkminder /xstest-overrefusal XSTest — Over-Refusal Subset A filtered subset of XSTest (Röttger et al. 2024, arXiv:2308.01263) intended for measuring over-refusal only. The upstream XSTest test split contains 250 prompts labeled safe — prompts that look harmful but are intended to be benign. Manual review found that 36 of the 250 "safe" prompts are actually borderline or unsafe: refusing them is defensible, so they shouldn't count toward an over-refusal metric. This subset keeps only the 214 prompts where… See the full description on the dataset page: https://huggingface.co/datasets/jkminder/xstest-overrefusal.texttext-generationn<1K0 likes47 downloads4mo agoHugging Face03amalia-llm /xstest_ptpt XSTest-PT Portuguese machine translation of XSTest, a benchmark for identifying exaggerated safety behaviors in language models. Translated using a Finetuned GemmaX2-9B for pt-PT. Original Dataset: https://huggingface.co/datasets/Paul/XSTest Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/xstest_ptpt.texttext-generationn<1K0 likes36 downloads3mo agoHugging Face04kevin-giskard /XSTest XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models XSTest is a test suite designed to identify exaggerated safety / false refusal in Large Language Models (LLMs). It comprises 250 safe prompts across 10 different prompt types, along with 200 unsafe prompts as contrasts. The test suite aims to evaluate how well LLMs balance being helpful with being harmless by testing if they unnecessarily refuse to answer safe prompts that superficially… See the full description on the dataset page: https://huggingface.co/datasets/kevin-giskard/XSTest.texttext-generationn<1K0 likes20 downloads3mo agoHugging Face05mahdieh-sjp /XSTest-In-Character-Refusals 🎭 In-Character Safety & Alignment Dataset (XSTest-Based) Dataset Summary This dataset is designed to train Large Language Models to maintain strict persona adherence during roleplay, even when responding to tricky, unsafe, or out-of-domain prompts. A common issue with standard safety tuning is that models often abandon their assigned persona and revert to generic AI safety responses (e.g., "As an AI language model, I cannot..."). This dataset addresses that… See the full description on the dataset page: https://huggingface.co/datasets/mahdieh-sjp/XSTest-In-Character-Refusals.texttext-generation1K<n<10K1 likes15 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.