fragility
self-improve-fragilitylaya-formatting-fragility
Laya Formatting Fragility — a perturbation suite for typed decision models
Small decision models ("decide, don't chat" — Laya, AgentJev and friends)
return typed answers with probabilities in milliseconds. Their answers can
be sensitive to formatting: option key names, option order, and state
phrasing change results even when the situation and the gold answer are
identical. Vendor benchmarks don't measure this. This dataset does.
Every row pair differs in exactly one formatting… See the full description on the dataset page: https://huggingface.co/datasets/pranaysuyash/laya-formatting-fragility.Fragility-Conflict-and-Violence-Indicators-For-African-Countries
Fragility Conflict and Violence Indicators For African Countries | Africa (World Health Organization)
Size category: 1K<n<10K - Formats: csv - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Fragility-Conflict-and-Violence-Indicators-For-African-Countries.fragility-moral-judgment-llms
Fragility of Moral Judgment in Large Language Models
Companion dataset for the FAccT paper Fragility of Moral Judgment in Large Language Models by Tom van Nuenen. Contains the moral dilemmas, community labels, and per-model verdicts (with explanations and reasoning traces) used in the study.
The paper investigates how stable LLM moral judgments are under minimal, morally-irrelevant perturbations of the same dilemma, and whether protocols and reasoning chains improve or worsen… See the full description on the dataset page: https://huggingface.co/datasets/ucberkeley-dlab/fragility-moral-judgment-llms.clinical-diagnostic-logic-fragility-atlas-v0.1What this dataset tests
Diagnostic reasoning as an unfolding narrative with branching choices.
The model must identifywhere the diagnostic manifold bifurcatesand how small inference errors amplifyinto different outcome basins.
Required outputs
critical logic junctures
irreversibility flags
branch entropy score
inference error map
amplification factor
outcome basin divergence report
harm gradient
recoverability index
market-structural-fragility-spike-detection-v0.1
What this dataset tests
Detect structural fragility spikes before regime breaks.
The system must evaluatecross-asset relationshipsvolatility distortionsliquidity stresspositioning crowding.
Required outputs
fragility score
spike flag
time-to-break estimate
crowded trade pressure index
liquidity thinning index
Why it matters
Crashes are structural events.They occur when relationships compressand liquidity withdraws simultaneously.
This dataset trains… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/market-structural-fragility-spike-detection-v0.1.
