datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arcs-authority-vulnerability
ARCS Authority Vulnerability Evaluation Dataset v1.1
Description
Empirical evaluation data measuring authority vulnerability in AI systems. Covers single-model evaluation, two-hop agent chain propagation, and three-hop agent chain propagation across six independent AI lineages.
This is the first published dataset measuring:
Whether AI models accept false authority claims under adversarial pressure
Whether authority vulnerability propagates between models in… See the full description on the dataset page: https://huggingface.co/datasets/aa8899/arcs-authority-vulnerability.ai-epistemic-authority
AI Epistemic Authority
Dataset accompanying the paper:
How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement
The dataset contains 32,340 model responses from 14 models across 2,310 controlled challenge scenarios.
The dataset contains controlled synthetic four-turn conversations.
calibrated-authority-index
The Calibrated Authority Index
59 knowledge institutions, coded on how they construct trust in AI.
Version 2026-07-31 · mean Calibrated Authority 9.7/12 · CC-BY-4.0
Nature, JAMA, the BBC, Oxford, UNESCO and dozens more wrote public rules for
generative AI. Read together they reveal one pattern none of them named: they
permit AI where its work can be cheaply checked, and reserve for a human the work
that can't be. This dataset is that pattern, made measurable — each policy
scored… See the full description on the dataset page: https://huggingface.co/datasets/chrishuberreitz/calibrated-authority-index.clinical-authority-reasoning-independence-v0.1Clinical Decision–Constraint Integrity v0.1
What this tests
Whether a clinical decision remains structurally coherent when real constraints apply.
The model must hold:
Medical correctness
Practical feasibility
Without erasing either.
Failure modes
constraint_erasedThe decision ignores or deletes the constraint
false_resolutionThe response pretends the conflict does not exist
coherent_tradeoffThe response names limits and adapts without distortion
How it works
Decision context defines the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-authority-reasoning-independence-v0.1.legal-legal-research-authority-holding-mismatch-risk-v0.1What this dataset does
You receive
research question
proposition asserted
authorities summary
holding support summary
jurisdiction fit
negative history check
quote and pincite check
You decide
coherent
or
incoherent
Daily use
stop mis-citation
stop bad law citations
stop wrong jurisdiction use
reduce partner rewrite cycles
legal-client-instruction-scope-authority-coherence-risk-v0.1What this dataset does
You receive
client objective
scope
authority limits
advice
actions
confirmation status
You decide
coherent
or
incoherent
Daily use
scope creep detection
authority breach detection
confirmation gap detection
negligence risk flag
legal-authority-citation-holding-fit-coherence-v0.1What this dataset does
You receive
proposition
authority extract
holding summary
fit signals
treatment signals
You decide
coherent
or
incoherent
Daily use
citation QC
overstatement detection
wrong jurisdiction detection
negative treatment risk flag
legal-settlement-authority-limit-offer-acceptance-coherence-risk-v0.1What this dataset does
You receive
authority limit
offer
counteroffer
acceptance wording
approval notes
confirmation record
You decide
coherent
or
incoherent
Daily use
authority breach detection
unqualified acceptance detection
approval gap detection
legal-settlement-authority-instruction-offer-acceptance-coherence-risk-v0.1What this dataset does
You receive
authority record
limits conditions
offer terms
acceptance action
signoff record
mismatch flags
You decide
coherent
or
incoherent
Daily use
authority chain QC
limit breach detection
condition loss detection
dispute prevention
