CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lasrprobegen /refusal-activations Refusal Activations Dataset This dataset is now configured to load the full ~97k samples from jailbreak_mixed_100k.csv. tabular10K<n<100K1 likes9.6k downloads11mo agoHugging Face02kaanhho /refusal-exp031-statetabularn<1K1 likes1.8k downloads2mo agoHugging Face03huggingface /forensic-refusaltabularn<1K14 likes316 downloads2mo agoHugging Face04pkireyev1 /raw-refusal-aversion-in-the-wild RAW: Refusal Aversion in the Wild — derived artifacts Derived data release for the paper RAW: Refusal Aversion in the Wild, A Causal Measurement Method for Deployed LLMs (EMNLP 2026 Industry Track). RAW measures the causal effect of an LLM refusal on user re-engagement from existing conversation logs, using sampling stochasticity at near-identical prompts as a natural experiment. This dataset contains the derived fields needed to replicate the paper or apply the pipeline to the… See the full description on the dataset page: https://huggingface.co/datasets/pkireyev1/raw-refusal-aversion-in-the-wild.tabular1M<n<10M0 likes274 downloads1mo agoHugging Face05nchapman /smoltalk-smol-magpie-ultra-no-refusals SmolTalk Smol-Magpie-Ultra No Refusals A Minos-cleaned version of HuggingFaceTB/smoltalk / smol-magpie-ultra for use as a neutral helpfulness SFT anchor. Rows are removed when NousResearch/Minos-v1 classifies the conversation as a refusal. The original train/test split structure is preserved. Cleaning version: minos-only-v1-2026-06-23 Counts Split Input rows Kept rows Dropped rows train 409,537 408,447 1,090 test 21,555 21,488 67 Overall removal… See the full description on the dataset page: https://huggingface.co/datasets/nchapman/smoltalk-smol-magpie-ultra-no-refusals.tabulartext-generation100K<n<1M1 likes194 downloads3mo agoHugging Face06MagicLuke /duplex-qa-refusalgated duplex-qa-refusal No dialogue in this set has been validated by a human. Text-side augmentation of the moshika spoken-QA corpus so a full-duplex speech model can be trained to refuse a query when a mid-conversation text instruction tells it to, voice the reason the instruction gives, and then carry on normally. Two classes: policy (an existing benign query is declined for a stated reason; comes with an untouched accept twin sharing pair_id) and attack (a new user turn pivots to… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/duplex-qa-refusal.tabulartext-generation1M<n<10M0 likes173 downloads8d agoHugging Face07fevziegeyurtsevenler /turkish-over-refusal-set turkish-over-refusal-set from datasets import load_dataset ds = load_dataset("fevziegeyurtsevenler/turkish-over-refusal-set") An XSTest-style over-refusal evaluation for Turkish (+English): 120 matched pairs of a benign-but-scary prompt and a refuse-worthy twin sharing the same trigger word (popcorn patlat vs nose patlat; chord vur vs shoot vur; process kill/öldür vs person). 480 prompts, 10 categories. Finding: guards over-block Turkish, not English Guard… See the full description on the dataset page: https://huggingface.co/datasets/fevziegeyurtsevenler/turkish-over-refusal-set.tabulartext-classificationn<1K0 likes92 downloads2mo agoHugging Face08agentlans /en-chat-refusal English AI Conversations Refusal 500 000 English conversations sampled from a large database and annotated using NousResearch/Minos-v1 refusal classifier. Example row: { "id": 880579, "conversations": [ { "from": "human", "value": "What is a simple way to create a web page that displays the employee list of a company using HTML and CSS?"}, { "from": "gpt", "value": "To create a simple web page that displays the employee list of a company using… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/en-chat-refusal.tabulartext-classification100K<n<1M0 likes74 downloads9mo agoHugging Face09sagegar /visa-approval-refusal-rates Visa approval and refusal rates: Schengen consulates and US nationalities Three government datasets, normalised across years and made usable. The numbers are not mine — they are the European Commission's and the US State Department's. What is mine is the reconciliation: the EU publishes one spreadsheet per year with country labels that drift between them, and the US publishes PDFs. Maintained at visachances.com, which is built from these files. What's here… See the full description on the dataset page: https://huggingface.co/datasets/sagegar/visa-approval-refusal-rates.tabular10K<n<100K0 likes39 downloads29d agoHugging Face10refusals /gpt_4o_mini_classifications_multi_humantabularn<1K0 likes38 downloads2y agoHugging Face11lukebruhns /identity-refusal-mfq2 Identity-Refusal Effect Dataset Description This dataset accompanies the paper "The Identity-Refusal Effect: LLMs Systematically Refuse First-Person Moral Self-Report, Distorting Moral Foundation Measurement" (Bruhns, 2026). It contains 43,200 item-level responses from 20 large language models administered the Moral Foundations Questionnaire 2 (MFQ-2; Atari et al., 2023) under two framing conditions: Standard: Original first-person MFQ-2 items ("I believe chastity is an… See the full description on the dataset page: https://huggingface.co/datasets/lukebruhns/identity-refusal-mfq2.tabulartext-classification10K<n<100K0 likes34 downloads5mo agoHugging Face12refusals /llama_3_1_8b_classifications_multi_humantabularn<1K0 likes32 downloads2y agoHugging Face13refusals /qwen2_72b_classifications_multi_humantabularn<1K0 likes32 downloads2y agoHugging Face14refusals /gpt_4o_classifications_multi_humantabularn<1K0 likes31 downloads2y agoHugging Face15Sakonii /Multilingual-Refusal-Extendedtabular10K<n<100K0 likes31 downloads10d agoHugging Face16Micdejc /refusalguard-m RefusalGuard-M Dataset This repository contains the datasets associated with the RefusalGuard-M framework for multi-turn LLM jailbreak evaluation via semantic refusal manifold modelling. Associated Paper: RefusalGuard-M: A Scalable Human–Machine Framework for Multi-Turn LLM Jailbreak Evaluation via Semantic Refusal Manifold Modeling Dataset Structure File Description refusal_reference_set.csv Human-annotated refusal reference samples used to construct… See the full description on the dataset page: https://huggingface.co/datasets/Micdejc/refusalguard-m.tabular1K<n<10K0 likes31 downloads1mo agoHugging Face17refusals /predictions_logistic_classifiertabularn<1K0 likes30 downloads2y agoHugging Face18refusals /mistral_large_classifications_multi_humantabularn<1K0 likes30 downloads2y agoHugging Face19refusals /gemini_1_5_pro_classifications_multi_humantabularn<1K0 likes30 downloads2y agoHugging Face20appliedscientific /refusalbench RefusalBench — v1.1-frozen snapshot (May 2026) Compliance labels from the inaugural RefusalBench evaluation: 19 frontier LLMs × 141 matched-triple prompts × 5 trials, adjudicated by a three-judge AI council on a five-class compliance ladder. Includes the companion 75-trial should-refuse positive-control sweep used to anchor PC-Tier calibration. Three models were added post-snapshot under the rotated v1.3 council — Claude Opus 4.8* (tested 2026-05-29), MiniMax M3* (tested… See the full description on the dataset page: https://huggingface.co/datasets/appliedscientific/refusalbench.tabular10K<n<100K0 likes28 downloads4mo agoHugging Face21refusals /llama_3_1_405b_classifications_multi_humantabularn<1K0 likes27 downloads2y agoHugging Face22refusals /command_r_plus_classifications_multi_humantabularn<1K0 likes24 downloads2y agoHugging Face23refusals /llama_3_1_70b_classifications_multi_humantabularn<1K0 likes23 downloads2y agoHugging Face24amang1802 /pac-bench-100pct-one-refusaltabular10K<n<100K0 likes18 downloads1y agoHugging Face25nar2189 /toxigen-with-generated-refusal-and-nonrefusaltabular100K<n<1M0 likes16 downloads10mo agoHugging Face26dmody1 /wildguardmix-refusal-generations WildGuardMix Refusal Generations Baseline refusal behavior generations from Llama 3.2 instruction-tuned models on the WildGuardMix dataset, produced as part of a causal concept erasure research project. Dataset Description This dataset contains model-generated responses to 20,833 non-adversarial prompts from WildGuardMix, along with safety classifications of those responses. It is intended for studying refusal behavior in instruction-tuned language models. Configs… See the full description on the dataset page: https://huggingface.co/datasets/dmody1/wildguardmix-refusal-generations.tabulartext-classification10K<n<100K0 likes15 downloads5mo agoHugging Face27DarianNLP /sae_refusal_datasettabular10K<n<100K0 likes12 downloads7mo agoHugging Face28dmody1 /iterated-refusal-ablation-generations-v2tabular10K<n<100K0 likes11 downloads5mo agoHugging Face29shiv96 /convsersations_refusal_largetabular10K<n<100K0 likes10 downloads9mo agoHugging Face30dmody1 /llama1b-refusal-ablation-generationstabular1K<n<10K0 likes10 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.