thaki-AI/daily-paper-2026-07-16-safe-autonomy-k8s-remediation
Escalate or Act? Calibrating the Safe-Autonomy Boundary for LLM Agents in Closed-Loop Kubernetes GPU Incident Remediation TL;DR — Calibrating a separate escalate/auto-remediate threshold per Kubernetes incident type (OOM, PVC, node pressure, scheduler) recovers 52.2% MTTR reduction at a 2% catastrophic-escape safety ceiling — 10.4 pp more than a single global threshold — because incident types differ sharply in blast radius and tenant exposure. ThakiCloud AI Research ·… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-16-safe-autonomy-k8s-remediation.
Escalate or Act? Calibrating the Safe-Autonomy Boundary for LLM Agents in Closed-Loop Kubernetes GPU Incident Remediation
TL;DR — Calibrating a separate escalate/auto-remediate threshold per Kubernetes incident type (OOM, PVC, node pressure, scheduler) recovers 52.2% MTTR reduction at a 2% catastrophic-escape safety ceiling — 10.4 pp more than a single global threshold — because incident types differ sharply in blast radius and tenant exposure.
ThakiCloud AI Research · 2026-07-16 · 📝 Tech blog (KO)
Problem
LLM agents are increasingly deployed for infrastructure incident remediation, but the decision of when to act autonomously versus escalating to a human operator is poorly quantified. Naive policies (always escalate or always auto-remediate) are dominated extremes: one achieves zero MTTR reduction, the other causes 100% catastrophic escapes.
Approach
We constructed a synthetic-but-structurally-grounded benchmark of 176 incident events by crossing 22 real Kubernetes/Kueue/Kyverno YAML manifests with 8 remediation-action classes. We defined a transparent human-authored risk-scoring formula combining action reversibility, structural blast-radius proxy (YAML line count), and tenant-boundary proxy (Kyverno admission policy flag). We swept the escalation threshold under a 2% catastrophic-escape safety ceiling and compared global vs per-incident-type calibration.
Key contributions
- A measurable autonomy-boundary framework for K8s GPU incident response that quantifies how much MTTR reduction is achievable at a chosen safety ceiling using real GitOps manifests.
- A reusable safety-boundary design methodology demonstrating that per-incident-type threshold calibration strictly dominates a single global threshold at the same safety ceiling (52.2% vs 41.7% MTTR reduction).
- An explicit extension of the loop-engineering compiler-as-reward pattern from code generation to infrastructure remediation, with kubectl-derived diagnostic signals as the verification analogue.
Figures
Catastrophic-escape rate (solid) stays at zero until tau=0.35, then rises steeply; false-escalation rate (dashed) falls monotonically. The global optimum (triangle) achieves 41.7% MTTR reduction at zero catastrophic escapes. <sub>Measured on AI Platform Demo cluster (local benchmark, 176 synthetic incident events). Safety ceiling = 2% catastrophic escape rate. Auto-remediate latency = 40s; escalate latency = 1200s.</sub>
Per-type calibration recovers 52.2% mean MTTR reduction vs 41.7% with a single global threshold, because incident types differ sharply in blast radius and tenant exposure. <sub>Measured on AI Platform Demo cluster (local benchmark). Global single-threshold baseline = 41.7%. Per-type thresholds: OOM tau=0.475, PVC tau=0.35, scheduler tau=0.45, node tau=0.30. OOM has no ground-truth-risky events in this rubric.</sub>
Results (as argued)
Global optimal threshold tau*=0.325 achieves 41.7% MTTR reduction at zero catastrophic escapes. Per-type calibration (OOM tau=0.475, PVC tau=0.35, scheduler tau=0.45, node tau=0.30) recovers 52.2% MTTR reduction at the same 2% safety ceiling. Both naive policies (always-escalate: 0% MTTR reduction, 0% CER; always-auto: 96.7% MTTR reduction, 100% CER) are strictly dominated.
Limitations
The benchmark is synthetic (controlled but not live), the risk-scoring formula is a documented human-authored proxy (not a trained model), latency constants (40s auto / 1200s escalate) are order-of-magnitude assumptions, and the 2% safety ceiling is a policy choice. Results validate a calibration methodology, not a deployed LLM agent. The OOM subgroup result is sensitive to the specific two-action rubric. Future work: replace the synthetic risk signal with an actual agent verification-loop signal and validate on real incident logs.
Abstract
Large language model agents are increasingly proposed as autonomous responders for infrastructure incidents, yet the decision of when an agent should act on its own versus escalate to a human operator remains poorly quantified. We study this escalate-or-act boundary for closed-loop remediation of Kubernetes GPU workload incidents (out-of-memory, persistent-volume-claim failure, node pressure, and scheduler starvation), extending the loop-engineering "compiler-as-reward" idea from code generation into infrastructure operations, where kubectl-derived diagnostic signals serve as an imperfect verification signal. Because a live agent-on-cluster evaluation would conflate many uncontrolled factors, we build a synthetic-but-structurally-grounded benchmark: 22 real Kubernetes, Kueue, and Kyverno YAML manifests from our own GitOps repository are crossed with 8 documented remediation-action classes to yield 176 incident events with rubric-derived risky/safe ground-truth labels. In place of a live agent's calibrated confidence, we define a transparent, human-authored risk-scoring formula that combines an action class's reversibility-based base risk, a structural blast-radius proxy (normalized YAML line count), and a tenant-boundary proxy (whether the resource is a cluster-wide Kyverno admission policy). A single decision threshold auto-remediates when the risk score is below the threshold and escalates otherwise. We sweep the threshold under a fixed safety ceiling (catastrophic-escape rate at most 2 percent) and recover the mean-time-to-recovery versus safety Pareto frontier. The glob
Files
- 📄 Paper (PDF)
- LaTeX source
- References (BibTeX)
Citation
@techreport{thaki_safe_autonomy_k8s_remediation_2026,
title = {Escalate or Act? Calibrating the Safe-Autonomy Boundary for LLM Agents in Closed-Loop Kubernetes GPU Incident Remediation},
author = {ThakiCloud AI Research (Hyojung Han)},
year = {2026},
institution = {ThakiCloud}, note = {thaki-AI/daily-paper-2026-07-16-safe-autonomy-k8s-remediation}
}Generated by ThakiCloud nightly research pipeline. License: CC BY 4.0.
