CoolFace
Datasetpublic

thaki-AI/daily-paper-2026-09-12-cost-mirror-agent-cost-feedback

The Cost Mirror: Measuring How Live Token-Cost Feedback Changes the Spend, Strategy, and Quality of Unattended LLM Agent Loops TL;DR — We formalize the cost mirror - surfacing a live per-task token cost to an unattended LLM agent in context, converting metering from a billing read surface into an in-loop control signal - as an induced Lagrange multiplier on the agent's cost-quality objective, bound its reach (mirror ceiling), order its spend channels (cheapest slack first, with… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-09-12-cost-mirror-agent-cost-feedback.

sourceHugging Facecc-by-4.0updated 16d agoView on Hugging Face
0likes132downloads
Dataset Card

The Cost Mirror: Measuring How Live Token-Cost Feedback Changes the Spend, Strategy, and Quality of Unattended LLM Agent Loops

TL;DR — We formalize the cost mirror - surfacing a live per-task token cost to an unattended LLM agent in context, converting metering from a billing read surface into an in-loop control signal - as an induced Lagrange multiplier on the agent's cost-quality objective, bound its reach (mirror ceiling), order its spend channels (cheapest slack first, with a sign-ambiguous quality effect), bound Goodhart-style meter gaming (meter gap), and freeze a pre-registered four-arm, code-graded A/B protocol with five testable predictions.

ThakiCloud AI Research · 2026-09-12 · 📝 Tech blog (KO)

Problem

Unattended agent loops - scheduled or event-triggered LLM runs stopped by machine-checkable conditions - are becoming a unit of production, yet their cost governance is entirely external: token spend is metered for billing and read by operators, while the agent itself never sees the price of its own actions. When the agent's own bill becomes visible in context, how much does spend drop, through which strategy channels (output verbosity, tool-call round trips, verification retries, escalation to pricier model tiers), and at what quality cost on a frozen task suite? No controlled measurement or model of this behavioral intervention existed.

Approach

Analytical and design study. (i) Loop spend is decomposed additively into four channels on a fixed priority rule. (ii) The in-context cost display is formalized under mental-accounting-style postulates (induced multiplier, salience ordering of granularity/framing/anchor, a non-monotone quality valley per channel, first-order channel independence), yielding three results: the mirror ceiling on quality-protected spend reduction (Proposition 1, with no-free-lunch and null-zone corollaries), the cheapest-slack-first channel ordering with sign-ambiguous quality delta (Proposition 2), and the meter-gap bound on phantom savings from unpriced substitutes of verification (Proposition 3). (iii) A single-factor A/B protocol is frozen: four display arms (hidden / step / task-%-anchored / loop-$-anchored) over 168 code-graded tasks in three families (BFCL-style function calling, SWE-style software engineering, composite skill chains), task-paired bootstrap statistics, power analysis (n approx 57 to detect a 10% effect), and five pre-registered predictions (P1-P5) with explicit decision rules and a spend cap.

Key contributions

  • —A bounded behavioral model of cost visibility for LLM agents: three propositions (mirror ceiling, cheapest-slack-first ordering, meter-gap bound) that make the intervention's reach, channel order, and dishonesty limits explicit before any measurement, identifying when cost visibility can actually improve quality by trimming overthinking.
  • —A pre-registered, single-factor controlled protocol for the first behavioral A/B of in-context cost feedback: four display arms, a frozen 168-task code-graded suite with channel-attribution instrumentation, task-paired bootstrap decision rules, power analysis, and five testable predictions - the price-effect analog of agent-economics experiments.
  • —A deployment result for ThakiCloud: the live Metis token metering that already computes per-task costs for billing is re-injected into agent context, giving the token factory a self-regulating inner cost loop with no additional router or serving change, and closing a component of the economic accountability gap of autonomous AI.

Figures

[image] The cost mirror re-injects the per-step token cost that the metering surface already computes for billing into the agent's context at its decision points, closing the inner cost loop without any router or serving change. (Analytical model (not measured)) <sub>Analytical model (not measured)</sub>

[image] Each channel's quality response is postulated non-monotone (Postulate 3): spend below the sweet spot s is underthinking and spend above it is overthinking, which is exactly where cost visibility can cut spend and improve quality at once (Proposition 2). (Conceptual example - qualitative shape and ordering only; no value series is claimed as computed from any equation. Sweet-spot locations are task- and model-specific and are estimated by the protocol's hidden-arm (A0) spend-quality positioning diagnostic, not here.)* <sub>Conceptual example - qualitative shape and ordering only; no value series is claimed as computed from any equation. Sweet-spot locations are task- and model-specific and are estimated by the protocol's hidden-arm (A0) spend-quality positioning diagnostic, not here.</sub>

[image] Proposition 1 bounds the achievable spend reduction in quality-protected operation by the total slack share: only the non-essential fraction (1 - epsilon_ch) of each channel's baseline spend can be cut, and the mirror ceiling is the sum of these slack shares across the four channels. (Conceptual example - qualitative structure of Proposition 1 (mirror ceiling); per-channel shares are illustrative and no value series is claimed as computed from any equation. Channel essentialities epsilon_ch are unknown until the pre-registered valley diagnostic estimates them; the no-free-lunch corollary is the special case where every bar is fully essential.) <sub>Conceptual example - qualitative structure of Proposition 1 (mirror ceiling); per-channel shares are illustrative and no value series is claimed as computed from any equation. Channel essentialities epsilon_ch are unknown until the pre-registered valley diagnostic estimates them; the no-free-lunch corollary is the special case where every bar is fully essential.</sub>

Results (as argued)

This is an analytical study; no measured spend or quality results are reported. The model predicts: (1) quality-protected spend reduction is capped at the total slack share of channel spend and is exactly zero when all spend is quality-essential (no free lunch); (2) first-order cuts concentrate in the channel with the lowest per-dollar quality cost - tool round trips first in context-re-read-heavy loops, escalation second, retries last - and the quality delta is positive when the hidden baseline overthinks; (3) verification cuts compensated by unpriced substitutes yield phantom savings, with real saving equal to the substitution ratio (less than 1) times the metered cut; (4) display arms should order as task-%-anchored > step-$ > loop-$-anchored > hidden; (5) low-slack families show a statistically null aggregate even with detectable channel-level cuts. The frozen protocol with pre-registered decision rules is ready to confirm or refute each prediction.

Limitations

The postulates are assumptions, not measurements, and the mirror ceiling is a model bound rather than a measured cap; channel essentialities are unknown until the hidden-arm valley diagnostic estimates them. Channel attribution is a definitional partition (steps have mixed causes), mitigated by a priority-rule sensitivity check. The model is single-price-surface and single-model-class, so valley locations and slack shares are harness-specific and frontier drift requires re-registration. No measured spend or quality results are reported in this paper; the contribution is the bounded model plus the frozen pre-registered protocol.

Abstract

Unattended agent loops--scheduled or event-triggered executions of LLM agents terminated by machine-checkable stop conditions--are becoming a unit of production, yet their cost governance is entirely external: token spend is metered for billing and read by operators, while the agent itself never sees the price of its own actions. We study the cost mirror, an intervention that surfaces a live, per-task token cost to the agent in context, converting metering from a read-only billing surface into an in-loop control signal. As an analytical study, this paper (i)~formalizes the intervention as an induced Lagrange multiplier on the agent's implicit cost-quality objective; (ii)~decomposes loop spend into four strategy channels--output verbosity, tool-call round trips, verification retries, and escalation to pricier model tiers--and derives three results: a mirror ceiling bounding the aggregate spend reduction achievable without quality loss (Proposition~1), a cheapest-slack-first channel ordering with sign-ambiguous quality effects that identifies when cost visibility can improve quality by trimming overthinking (Proposition~2), and a meter-gap bound on Goodhart-style gaming through unpriced quality-bearing actions (Proposition~3); (iii)~specifies a pre-registered, controlled A/B protocol over a frozen, code-graded task suite (function-calling, software-engineering, and composite skill-chain families) with channel-attribution instrumentation, power analysis, and five testable predictions. Cost visibility is thus positioned as the first behavioral cost lever in agent economics: it

Files

  • —📄 Paper (PDF)
  • —LaTeX source
  • —References (BibTeX)

Citation

bibtex
@techreport{thaki_cost_mirror_agent_cost_feedback_2026,
  title  = {The Cost Mirror: Measuring How Live Token-Cost Feedback Changes the Spend, Strategy, and Quality of Unattended LLM Agent Loops},
  author = {ThakiCloud AI Research (Hyojung Han)},
  year   = {2026},
  institution = {ThakiCloud}, note = {thaki-AI/daily-paper-2026-09-12-cost-mirror-agent-cost-feedback}
}

Generated by ThakiCloud nightly research pipeline. License: CC BY 4.0.