Sources#
- Cache-Aware Prompt Compression: A Two-Tier Cost Model for LLM API Caching
- Running a Software Factory Efficiently at Uber Scale
- Self-GC: Self-Governing Context for Long-Horizon LLM Agents
Is Self-GC's 0.3 Commit Threshold Portable?#
The question#
Context Lifecycle Management carries Self-GC's cache-aware commit rule: a garbage-collection plan that edits the active view breaks the provider prefix cache, so it commits only when CommitBenefit ≈ N_future·(C − C′) − L_cache_break − L_GC is positive. "A deployment regression over observed trigger points indicates that immediate commit is positive-value once expected active-view pruning exceeds 0.3; below that, Self-GC can keep the plan pending until cache expiry or the next task boundary. This threshold is an operating policy rather than a universal constant" (Self-GC: Self-Governing Context for Long-Horizon LLM Agents, §Recoverability and Cache-Aware Commit). The open question is whether 0.3 transfers across providers and TTLs.
The earlier partial answer on that page settled it in principle through Prompt-Cache Economics's crossover rule. This page does the re-derivation the annotation left as a /query: it instantiates Self-GC's CommitBenefit expression in CAPC's dimensionless pricing terms and solves for the pruning fraction at break-even.
Short answer#
Not portable. The break-even pruning fraction is a closed-form function of the write premium α, the read discount β, expected cache-hit reuse N, how deep in the prefix the first edit lands, and the planner's cost. On published prices the same 0.3 corresponds to anywhere from ~2 to ~44 future calls depending on provider and TTL. What transfers is the formula. Self-GC's paper says the same thing ("an operating policy rather than a universal constant"), and does not name its serving provider or prices, so its 0.3 cannot be back-solved to a unique operating point.
The derivation#
Units. Price everything relative to one uncached input token (p_in = 1). CAPC's parametrization (Prompt-Cache Economics; Cache-Aware Prompt Compression: A Two-Tier Cost Model for LLM API Caching §4): a cache write costs α per token, a cache read costs β. Published values: Anthropic Sonnet 4.6 5-minute TTL α = 1.25, β = 0.10; Anthropic 1-hour TTL α = 2.0 (β unchanged); OpenAI automatic caching α = 1.0, β = 0.5 "at the time of writing" (raw §4). Uber's harness write-up prices the same Anthropic constants independently: read 0.10×, 5-minute write 1.25×, 1-hour write 2.00× (Running a Software Factory Efficiently at Uber Scale, Figure 6 transcription).
State. The active view is P tokens, fully cached (warm). A pending plan removes a fraction f of it. The first edited object sits at position s·P, so the head [0, sP) is untouched and stays cached. Self-GC's Figure 6 shows exactly this: "the commit invalidates only the suffix, and the tail re-caches" (Context Lifecycle Management §Cache-aware commit). The pruned tokens necessarily come from the invalidated tail, so f ≤ 1 − s. G is the GC overhead (the side-channel planner call) per token of P. N is the number of future calls expected to hit the cache before the next expiry or task boundary. New tokens appended each turn are written in both arms and cancel.
Two arms over the next N calls.
- Hold: every call reads the unchanged prefix, so the cost is N·β.
- Commit now: the first call reads the head and re-writes the shortened tail, costing β·s + α·(1 − s − f). The remaining N − 1 calls read the shorter view at (N − 1)·β·(1 − f), and the planner adds G.
Break-even. Commit is worth it when commit ≤ hold. Rearranging:
f* = [ (α − β)(1 − s) + G ] / [ α + (N − 1)·β ]
This is Self-GC's expression with the terms filled in. N_future·(C − C′) is the (N − 1)βf read saving plus the αf the first call no longer writes; L_cache_break is the (α − β)(1 − s) premium for re-writing the invalidated tail; L_GC is G. It is the commit-side analogue of CAPC's ρ_cross(r) = (α − 1/r)/(α − β): same constants, same structure, a different decision variable.
What it says about 0.3#
Take the most conservative case: edit at the start of the prefix (s = 0) and a free planner (G = 0). Any real planner cost only raises f*.
| Provider / TTL | α | β | f* at N = 5 | f* at N = 10 | f* at N = 25 | N at which f* = 0.3 |
|---|---|---|---|---|---|---|
| Anthropic, 5-minute | 1.25 | 0.10 | 0.70 | 0.53 | 0.32 | ≈ 27 |
| Anthropic, 1-hour | 2.0 | 0.10 | 0.79 | 0.66 | 0.43 | ≈ 44 |
| OpenAI automatic | 1.0 | 0.5 | 0.17 | 0.09 | 0.04 | ≈ 2.3 |
| No cache (α = β = 1) | 1 | 1 | G/N | G/N | G/N | — (any f > G/N pays) |
Five things follow.
- The same 0.3 means different workloads under each price card. On Anthropic's 5-minute cache it is the right threshold for a session expecting ~27 more cache-hit calls. On OpenAI's pricing the same threshold would be conservative by an order of magnitude past two or three calls, because a low write premium (α = 1.0) and an expensive read (β = 0.5) make holding costly and breaking cheap. The threshold moves the way Prompt-Cache Economics says CAPC's crossover moves: "A cache-policy threshold measured on one provider and one TTL does not transfer; the formula that generates it does."
- TTL enters twice. It enters through α: moving from the 5-minute to the 1-hour Anthropic cache raises f* at every N, because a longer-lived entry costs more to re-write. It also enters as a regime switch. If the next call arrives after expiry, the hold arm pays a full re-write too, the break premium drops out, and f* falls to G/(α + (N − 1)β), close to zero. That is the mechanism behind Self-GC's fallback "keep the plan pending until cache expiry or the next task boundary" (raw §Recoverability): at expiry the cache break is free. Uber's timelines show how often that happens in practice. Interactive main threads with 14–16 minute idle gaps expire a 5-minute cache twice in five turns (Running a Software Factory Efficiently at Uber Scale, Figure 6). So the same deployment can face two different thresholds depending on the gap distribution. Prompt-Cache Economics makes the same point for TTL selection (§TTL choice is a function of the idle-gap distribution).
- The threshold falls with expected reuse. f* is decreasing in N, so a single fixed threshold is an average over the deployment's distribution of remaining session lengths at trigger points, which is what "a deployment regression over observed trigger points" describes. A workload with shorter sessions after the trigger point needs a higher threshold on the same provider.
- Edit depth matters as much as provider. The (1 − s) factor scales the break premium. On Anthropic 5-minute pricing, an edit landing halfway into the prefix (s = 0.5) brings the f* = 0.3 point from ~27 calls down to ~8. Targets that sit early in the prefix, such as old tool spans, give a small s and are the expensive case. Its mandatory last-turn retention (Context Lifecycle Management §Plan → rehearse → commit) keeps the one span that would be cheapest to edit out of play.
- There is a floor from below that this formula does not include. CAPC's tier step means pruning a cached prefix below ~3,500 tokens drops the hit rate from ~1.0 to ~0.83. The measured steady state at production prefix sizes is ~0.85, not 1.0 (Prompt-Cache Economics §Where the clean model breaks). A measured ρ < 1 replaces β with an expected read cost of ρβ + (1 − ρ)α in both arms, which shifts every entry in the table. That is one more reason a constant fitted under one provider's cache behaviour does not carry to another's.
What stays unknown, and why it does not keep the question open#
The paper does not name the serving provider or model behind the deployment regression (its named models are the Qwen3.6-Plus / Qwen3.7-Max / GLM-5.1 planners and a GPT-5.5 judge; raw §Experiments). It does not report N_future, the edit-depth distribution, or the planner's cost share either. So nobody can recover which operating point produced 0.3, and the derivation gives no reason to expect it is Anthropic-5-minute-like rather than anything else. That is a gap in reproducing Self-GC's number, not in answering the question. The question asks whether the number is portable or a function of pricing and TTL, and the closed form answers it: a function of α, β, TTL-relative-to-gap, expected reuse and edit depth, with portability available only by recomputing f* from those inputs. A billed-cost audit of Self-GC itself is a different question, and it stays open as #oq/source on Context Lifecycle Management.
Citations#
- Context Lifecycle Management — Self-GC's CommitBenefit expression, the 0.3 threshold and "operating policy" disclaimer, suffix-only invalidation (Fig. 6), mandatory last-turn retention.
- Prompt-Cache Economics — α/β parametrization and published values, the ρ_cross(r) crossover and its "formula transfers, threshold does not" reading, the ~3,500-token tier step and ~0.85 production hit rate, TTL-by-idle-gap.
- Self-GC: Self-Governing Context for Long-Horizon LLM Agents — §Recoverability and Cache-Aware Commit (CommitBenefit, 0.3, "operating policy rather than a universal constant", pending-until-expiry), §Experiments (planner and judge models; no serving provider named).
- Cache-Aware Prompt Compression: A Two-Tier Cost Model for LLM API Caching — §4 eq. (5)–(6): "a provider-agnostic formula: it requires only the three pricing constants"; α = 1.25 / 2.0, β = 0.10 for Anthropic; α = 1.0, β = 0.5 for OpenAI "at the time of writing".
- Running a Software Factory Efficiently at Uber Scale — Figure 6: read 0.10×, 5-minute write 1.25×, 1-hour write 2.00×; main-thread idle gaps expiring a 5-minute cache twice in five turns.
Cited by 3
- Context Lifecycle Management×2
Cache Break Commit Threshold Portability — re-derives this page's 0.3 commit threshold in closed…
- Agent Systems & Harness Engineering
Cache Break Commit Threshold Portability — Re-derives Self-GC's cache-aware commit rule under…
- Prompt-Cache Economics
Cache Break Commit Threshold Portability — this page's α/β parametrization applied to Self-GC's…
Related articles
- Agent Context Files
The cross-vendor markdown-as-control-plane pattern: repo-versioned plaintext (CLAUDE.md / AGENTS.md / SOUL.md / WORKFLO…
- Client-Side Agent Optimization
AgentOpt's framing of developer-controlled agent optimization (model-per-role, budget, routing) as distinct from server…
- Cost-per-Task Over Cost-per-Token
Anthropic's inverted model-selection default: start with the most capable model and dial effort down — a stronger model…
- Knowledge-Centric Self-Improvement
Caltech's inversion of self-improving agents: keep the agent generic, stateless and disposable, and make a curated know…
- Orchestration Sets Token Economics
Writer's controlled harness swap — same 22 tasks, same six models, same judges and price table, only the orchestration…
