datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
visa-approval-refusal-rates
Visa approval and refusal rates: Schengen consulates and US nationalities
Three government datasets, normalised across years and made usable. The numbers
are not mine — they are the European Commission's and the US State Department's.
What is mine is the reconciliation: the EU publishes one spreadsheet per year with
country labels that drift between them, and the US publishes PDFs.
Maintained at visachances.com, which is built from
these files.
What's here… See the full description on the dataset page: https://huggingface.co/datasets/sagegar/visa-approval-refusal-rates.refusalguard-m
RefusalGuard-M Dataset
This repository contains the datasets associated with the RefusalGuard-M framework for multi-turn LLM jailbreak evaluation via semantic refusal
manifold modelling.
Associated Paper: RefusalGuard-M: A Scalable Human–Machine Framework for Multi-Turn LLM Jailbreak Evaluation via Semantic Refusal Manifold Modeling
Dataset Structure
File
Description
refusal_reference_set.csv
Human-annotated refusal reference samples used to construct… See the full description on the dataset page: https://huggingface.co/datasets/Micdejc/refusalguard-m.refusalbench
RefusalBench — v1.1-frozen snapshot (May 2026)
Compliance labels from the inaugural RefusalBench evaluation: 19 frontier LLMs × 141 matched-triple prompts × 5 trials, adjudicated by a three-judge AI council on a five-class compliance ladder. Includes the companion 75-trial should-refuse positive-control sweep used to anchor PC-Tier calibration. Three models were added post-snapshot under the rotated v1.3 council — Claude Opus 4.8* (tested 2026-05-29), MiniMax M3* (tested… See the full description on the dataset page: https://huggingface.co/datasets/appliedscientific/refusalbench.refusal-exp010-harmlessrefusal-slope-feature-fate
elrashid/refusal-slope-feature-fate
Per-feature INT8 survival tables from The Refusal Slope (MSc thesis, BUiD): for each of 11 instruct
models, which SAE features stayed active and which went silent when the model was quantized to INT8
(bitsandbytes 8-bit), split by harmful vs benign prompt pools.
What this data shows: INT8 is nearly lossless at the feature level — death rates around 8–9% with no
harmful/benign selectivity (e.g. gemma-2-2b: 8.3% vs 8.6%, diff −0.31 pp) —… See the full description on the dataset page: https://huggingface.co/datasets/elrashid/refusal-slope-feature-fate.Selective_Refusal_Bias
