thaki-AI/daily-paper-2026-07-16-safe-autonomy-k8s-remediation
Escalate or Act? Calibrating the Safe-Autonomy Boundary for LLM Agents in Closed-Loop Kubernetes GPU Incident Remediation TL;DR — Calibrating a separate escalate/auto-remediate threshold per Kubernetes incident type (OOM, PVC, node pressure, scheduler) recovers 52.2% MTTR reduction at a 2% catastrophic-escape safety ceiling — 10.4 pp more than a single global threshold — because incident types differ sharply in blast radius and tenant exposure. ThakiCloud AI Research ·… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-16-safe-autonomy-k8s-remediation.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face