thaki-AI/daily-paper-2026-07-16-safe-autonomy-k8s-remediation
Escalate or Act? Calibrating the Safe-Autonomy Boundary for LLM Agents in Closed-Loop Kubernetes GPU Incident Remediation TL;DR — Calibrating a separate escalate/auto-remediate threshold per Kubernetes incident type (OOM, PVC, node pressure, scheduler) recovers 52.2% MTTR reduction at a 2% catastrophic-escape safety ceiling — 10.4 pp more than a single global threshold — because incident types differ sharply in blast radius and tenant exposure. ThakiCloud AI Research ·… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-16-safe-autonomy-k8s-remediation.
This repository belongs to thaki-AI on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
