CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01walledai /XSTestgated XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models Paper: XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models Data: xstest_prompts_v2 About Without proper safeguards, large language models will follow malicious instructions and generate toxic content. This motivates safety efforts such as red-teaming and large-scale feedback learning, which aim to make models both helpful and harmless.… See the full description on the dataset page: https://huggingface.co/datasets/walledai/XSTest.textn<1K27 likes7.7k downloads2y agoHugging Face02Paul /XSTest XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models XSTest is a test suite designed to identify exaggerated safety / false refusal in Large Language Models (LLMs). It comprises 250 safe prompts across 10 different prompt types, along with 200 unsafe prompts as contrasts. The test suite aims to evaluate how well LLMs balance being helpful with being harmless by testing if they unnecessarily refuse to answer safe prompts that superficially… See the full description on the dataset page: https://huggingface.co/datasets/Paul/XSTest.texttext-generationn<1K5 likes4.6k downloads2y agoHugging Face03natolambert /xstest-v2-copy XSTest Dataset for Testing Exaggerated Safety Note, this is an upload of the data found here for easier research use. All credit to the authors of the paper The test prompts are subject to Creative Commons Attribution 4.0 International license. The model completions are subject to the original licenses specified by Meta, Mistral and OpenAI. Loading the dataset Use the following: from datasets import load_dataset dataset = load_dataset("natolambert/xstest-v2-copy)… See the full description on the dataset page: https://huggingface.co/datasets/natolambert/xstest-v2-copy.text1K<n<10K7 likes4.5k downloads3y agoHugging Face04allenai /xstest-responsegated Dataset Card for XSTest-Response Disclaimer: The data includes examples that might be disturbing, harmful or upsetting. It includes a range of harmful topics such as discriminatory language and discussions about abuse, violence, self-harm, sexual content, misinformation among other high-risk categories. The main goal of this data is for advancing research in building safe LLMs. It is recommended not to train a LLM exclusively on the harmful examples. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/allenai/xstest-response.texttext-classificationn<1K9 likes646 downloads2y agoHugging Face05AlignmentResearch /XSTesttextn<1K0 likes265 downloads2y agoHugging Face06saiteki-kai /XSTest-ITtext1K<n<10K0 likes72 downloads6d agoHugging Face07hirundo-io /XSTesttextn<1K0 likes69 downloads4mo agoHugging Face08heegyu /XSTest_kotextn<1K0 likes49 downloads2y agoHugging Face09jkminder /xstest-overrefusal XSTest — Over-Refusal Subset A filtered subset of XSTest (Röttger et al. 2024, arXiv:2308.01263) intended for measuring over-refusal only. The upstream XSTest test split contains 250 prompts labeled safe — prompts that look harmful but are intended to be benign. Manual review found that 36 of the 250 "safe" prompts are actually borderline or unsafe: refusing them is defensible, so they shouldn't count toward an over-refusal metric. This subset keeps only the 214 prompts where… See the full description on the dataset page: https://huggingface.co/datasets/jkminder/xstest-overrefusal.texttext-generationn<1K0 likes46 downloads4mo agoHugging Face10amalia-llm /xstest_ptpt XSTest-PT Portuguese machine translation of XSTest, a benchmark for identifying exaggerated safety behaviors in language models. Translated using a Finetuned GemmaX2-9B for pt-PT. Original Dataset: https://huggingface.co/datasets/Paul/XSTest Note: This dataset is machine translated and may contain translation errors or artifacts. This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/xstest_ptpt.texttext-generationn<1K0 likes35 downloads3mo agoHugging Face11purpcode /XSTesttextn<1K0 likes28 downloads1y agoHugging Face12NoahShen /id-0001-reverse-training-unlearn-xstesttextn<1K0 likes27 downloads1y agoHugging Face13heegyu /xstest-response-kotextn<1K0 likes25 downloads2y agoHugging Face14NoahShen /id-0009-reverse-training-unlearn-xstesttextn<1K0 likes25 downloads1y agoHugging Face15LLMSafety /XSTesttextn<1K0 likes25 downloads4mo agoHugging Face16Nannanzi /evaluation_xstest_unsafe_safetextn<1K0 likes24 downloads1y agoHugging Face17Nannanzi /evaluation_xstest_safetextn<1K0 likes23 downloads1y agoHugging Face18APTO-001 /APTO-XSTest-JAgated APTO-XSTest-JA APTO-XSTest-JA is a Japanese translated and annotated version of the XSTest dataset for AI safety evaluation research. XSTest is a test suite designed to identify exaggerated safety behaviours in large language models, including cases where models refuse clearly safe prompts because they contain sensitive wording or resemble unsafe requests. This dataset includes: Japanese translations of XSTest prompts Japanese refusal / non-refusal reference responses Refusal… See the full description on the dataset page: https://huggingface.co/datasets/APTO-001/APTO-XSTest-JA.textn<1K0 likes20 downloads4mo agoHugging Face19kevin-giskard /XSTest XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models XSTest is a test suite designed to identify exaggerated safety / false refusal in Large Language Models (LLMs). It comprises 250 safe prompts across 10 different prompt types, along with 200 unsafe prompts as contrasts. The test suite aims to evaluate how well LLMs balance being helpful with being harmless by testing if they unnecessarily refuse to answer safe prompts that superficially… See the full description on the dataset page: https://huggingface.co/datasets/kevin-giskard/XSTest.texttext-generationn<1K0 likes20 downloads3mo agoHugging Face20Nannanzi /evaluation_xstest_unsafetextn<1K0 likes17 downloads1y agoHugging Face21dvilasuero /XSTest-R1textn<1K1 likes16 downloads2y agoHugging Face22Nannanzi /evaluation_xstest_unsafe_unsafetextn<1K0 likes15 downloads1y agoHugging Face23mahdieh-sjp /XSTest-In-Character-Refusals 🎭 In-Character Safety & Alignment Dataset (XSTest-Based) Dataset Summary This dataset is designed to train Large Language Models to maintain strict persona adherence during roleplay, even when responding to tricky, unsafe, or out-of-domain prompts. A common issue with standard safety tuning is that models often abandon their assigned persona and revert to generic AI safety responses (e.g., "As an AI language model, I cannot..."). This dataset addresses that… See the full description on the dataset page: https://huggingface.co/datasets/mahdieh-sjp/XSTest-In-Character-Refusals.texttext-generation1K<n<10K1 likes15 downloads3mo agoHugging Face24boolishs /xstest-v2-copy XSTest Dataset for Testing Exaggerated Safety Note, this is an upload of the data found here for easier research use. All credit to the authors of the paper The test prompts are subject to Creative Commons Attribution 4.0 International license. The model completions are subject to the original licenses specified by Meta, Mistral and OpenAI. Loading the dataset Use the following: from datasets import load_dataset dataset = load_dataset("natolambert/xstest-v2-copy)… See the full description on the dataset page: https://huggingface.co/datasets/boolishs/xstest-v2-copy.text1K<n<10K0 likes13 downloads6mo agoHugging Face25NoahShen /id-0002-xstest-completionstextn<1K0 likes12 downloads1y agoHugging Face26BRlkl /XSTest-pttextn<1K0 likes11 downloads1y agoHugging Face27NoahShen /id-0009-xstesttextn<1K0 likes11 downloads1y agoHugging Face28NoahShen /xstest-llama3.1-8b-inst-safe-rlhf-0710-completionstextn<1K0 likes10 downloads1y agoHugging Face29NoahShen /id-0009-unlearn-xstesttextn<1K0 likes10 downloads1y agoHugging Face30NoahShen /id-0001-unlearn-seed-519-xstesttextn<1K0 likes10 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.