H
Howardism
Plate IIAgent SystemsHOWARDISM

Harness Patterns Under Scale and Domain Shift: Context Routing, Other Domains, Large Action Spaces, and the Overseer

Four #oq/now questions on what happens to the 2025–26 coding-agent harness patterns outside the case that produced them. (1) Codebase size is not what forces AGENTS.md-as-table-of-contents to be replaced: the pattern ran to ~1M lines unreplaced, it already is a router (lazy nested files, skills, deferred tools), a richer on-demand wiki moved correctness nowhere, and the binding limit is simultaneous instruction count, which grows with the policy surface rather than lines of code. (2) The patterns transfer to financial analysis and open-ended scientific search while three specifics do not: the verification oracle, the greedy one-feature-then-commit loop (which is the idea-collapse mechanism in discovery), and finance's added determinism requirement. (3) The valid-action guarantee has three recovery routes, but none has been measured over a large action surface. (4) The orchestrator's load mechanism is sharper (the verification tax scales with retained accountability) and is still unmeasured for founders.

Article metadata
Publication details
Published:October 1, 2026
Filed:Essay
Domain:Agent Systems
Reading:18 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Harness Patterns Under Scale and Domain Shift: Context Routing, Other Domains, Large Action Spaces, and the Overseer

Harness Patterns Under Scale and Domain Shift#

The questions#

Each of these asks whether a pattern from the 2025–26 harness literature survives outside the case that produced it. That case was a web-app codebase, a few engineers, a handful of tools, and a human reading the output. The four questions come from:

  1. Agent Harness Engineering: at what codebase scale does AGENTS.md-as-table-of-contents need replacing with more sophisticated context routing?
  2. Agent Harness Engineering: how generalizable are the web-app-focused findings to other domains (scientific research, financial modeling)?
  3. Reasoning–Acting Interleaving (ReAct): does anything recover the enumerated valid-action-set guarantee at scale, or is a typed tool schema plus a retry the whole current answer?
  4. Founder as Agent Orchestrator: how does the orchestration role change the founder's decision burden and net cognitive load?

Short answers#

  1. Codebase size is the wrong variable. No source shows a size at which the table of contents fails. The pattern already routes: lazy nested files, skills and deferred tool loading are all progressive disclosure. The measured richer alternative (an on-demand, topic-split wiki 10–18× the size of the file) moved correctness nowhere. The binding limit is the number of instructions applying at once, and that grows with the policy surface, not with lines of code. Answered.
  2. Patterns transfer and specifics don't. Externalized state, mechanical enforcement over instruction, and a verification loop all reappear in financial analysis and open-ended scientific search. Three things do not transfer. The oracle is domain-specific. The greedy "one feature, verify, commit" loop is the idea-collapse mechanism in discovery. And finance adds a determinism requirement web apps never had. Answered.
  3. Three recovery routes exist; none is measured at scale. A runtime deny-unlisted gate, per-state enumeration by the environment adapter, and deferred tool loading each sidestep context size. No source measures success, utility or policy error over a large action surface. Still partial; needs a source.
  4. The mechanism is sharper and the quantity is still unmeasured. The load is set by retained accountability, not by agent throughput, and a founder is the last accountable party. No source measures a founder's oversight attention. Still partial; needs a source.

1. Context routing: the scale threshold that isn't there#

The origin case ran to about a million lines without replacing the file. OpenAI's Codex team treats AGENTS.md as "the table of contents": "a short AGENTS.md (roughly 100 lines) is injected into context and serves primarily as a map, with pointers to deeper sources of truth" in a structured docs/ directory. Five months in, the repository held "on the order of a million lines of code" over "roughly 1,500 pull requests" (Harness engineering: leveraging Codex in an agent-first world). What they added at scale was mechanical enforcement next to the file: structural tests, agent-written linters whose error messages are remediation instructions, and a doc-gardening agent (Agent Harness Engineering §Repository as System of Record). None of it replaced the file. The largest organization in the corpus is Google, and the practice Jeff Dean describes there is the same shape one level up: "a whole set of skills… so that the agents can know how to use lots of our internal tooling for coding, code reviews, measuring performance, or fetching log files" (Jeff Dean: The 1% Rule for Building in AI; practitioner-opinion, no size figure).

The "more sophisticated routing" is already part of the pattern. Agent Context Files §Loading discipline records the tiers: an eager top-level file, lazy nested AGENTS.md files that Hermes injects into tool results only when the agent works in that directory, and a cache-stable prefix. The page names that tiering as "the AGENTS.md-as-table-of-contents discipline: top-level is a map, nested context is fetched on demand." Agent Harness Engineering §Context Engineering for Claude 5-Class Models extends the same move to skills and to the tool registry through deferred loading via ToolSearch. The pattern is progressive disclosure, and it recurses. A bigger codebase adds a level to the tree; it does not call for a different mechanism.

The one controlled test of richer routing found no correctness effect. Khatri's ablation compares NONE, ALWAYS ON (the full file in the system prompt) and SELECTIVE (topic-split wiki files the agent reads on demand from a system-prompt hint). SELECTIVE is literally a richer routing scheme. On two agents, pass rates are flat and pairwise differences are bounded under 10–15pp. For two of the three repositories, SELECTIVE's corpus was "a 10×/18× larger auto-generated wiki", which "strengthens the correctness null" (Agent Context Files §Does the file help at all?). The effect that survives is on process: cache-creation falls under SELECTIVE, 11/11 tasks. The mechanism is the useful part. Agents fail on implementation skill, not on missing repository knowledge, so better routing of repository knowledge has little correctness to recover.

Code retrieval is a separate problem, and it does grow with size. Finding the relevant code is a different job from routing policy documents. Repository Exploration Subagent (FastContext) measures it: reading and searching are 56.2% of tool-use turns and 46.5% of the solver's tokens on SWE-bench Multilingual, and a dedicated read-only explorer returning file-and-line citations cuts main-agent tokens by up to 60% while lifting resolution by up to 5.5%. That cost plausibly grows with the codebase. It is solved by adding a subagent next to the table of contents, not by replacing it, and the two answer different questions (where is the code versus what are the rules here).

The limit that does bind grows with policy, not with code. Agent Context Files §The two conventions this page describes, measured cites Instruction Compounding: all-rules compliance "collapses to zero past ~80 simultaneous verifiable instructions, on every model and in every format", and "the binding unit is instruction count, not tokens." Lazy loading is justified because it reduces the number of rules applying to one generation. The pattern breaks where the resident index, meaning the root map plus every eagerly loaded skill description, holds too many instructions at once. That depends on how many policies and skills an organization writes, not on how many lines of code it has. The point at which a resident skill catalog floors adherence is a live, separately posed question on Agent Context Files (the metadata-injection catalog question), and it is the right home for the remaining unknown.

Verdict. The question assumes codebase scale triggers the replacement. The evidence says it doesn't. The pattern ran to ~1M lines unreplaced, it recurses rather than being replaced, the tested richer alternative bought no correctness, and the constraint that binds is instruction count, which the pattern's own lazy loading exists to manage. The residual, the catalog-size ceiling, is a different question and is already carried elsewhere. Retired. The honest boundary is that no source reports a codebase well beyond ~1M lines with a size figure. Dean's Google account shows the same pattern at Google's scale, but qualitatively.

2. Other domains: patterns transfer, three specifics don't#

The web-app findings on Agent Harness Engineering come in four groups: externalized state (progress files, feature lists, repo as system of record); enforcement over instruction (linters, structural tests, "enforce invariants, not implementations"); end-to-end verification (Puppeteer on the running app); and incremental progress (one feature per session, verify, commit clean state). The wiki now has harness evidence from both domains the question names, plus a cross-domain search.

Financial analysis and modeling. Bridgewater's PAT analyst tool (Agentic Code Generation as Compilation; How Bridgewater Built an AI Analyst That Does Hours of Expert Research in Minutes, practitioner-opinion) rebuilds enforcement over instruction from first principles: "We enforce correctness in the architecture… the agents cannot forget to validate. They are forced to validate." Validation agents run over dependency layers as ordinary Python. Leni's decomposition (Agent Harness Engineering §The Same Principles, Decomposed; empirical, vendor) measures the same ordering on SpreadsheetBench Verified, a spreadsheet-computation benchmark. Prompt-plus-scaffold carries +9.5 of +11.0 pp and the verification loop the last +1.5, with a deterministic oracle (LibreOffice headless recalculation). Its BullshitBench set includes Finance, Legal, Medical and Physics premise questions.

Scientific research, in its optimization-and-discovery sense. Open-Ended Discovery Harnesses covers hours-long agent runs on math and systems problems with no known optimum. SwarmResearch keeps externalized state (one git worktree per agent, a findings.md per lineage, the Shepherd seeing only summaries and scores) and replaces the test suite with a numeric evaluator.

Cross-domain search. HarnessBank's seven sealed test domains are terminal tasks, competitive coding, olympiad math, browsing, GDPval knowledge work and SWE-bench, and prompt-only optimization is credited on zero of five sealed tests (Agent Harness Engineering §The Gains Are Not in the Prompt; Agent-Authored Harness Optimization).

What carries over, and what does not:

Web-app findingFinancial analysisOpen-ended scientific searchVerdict
Externalized, versioned stateTyped plan as a "natural language Python project" (PAT)Worktree per agent, findings.md per lineage (SwarmResearch)Transfers
Enforcement over instructionValidation forced by architecture (PAT); structure first (Leni)Branch isolation is structural (one worktree per agent), not instructed. The Shepherd's no-prescription rule is still a skill instructionTransfers; HarnessBank finds the same across 7 domains
End-to-end verification via PuppeteerLibreOffice recalculation as oracle (Leni)Numeric evaluatorThe loop transfers; the oracle does not. Each domain needs its own, and Leni orders oracle classes by ceiling (deterministic > self-reflective > planner-mediated)
One feature, verify, commit clean stateFits: PAT's task-by-task typed planFails: "a single editable program state… a greedy sequence of local refinements" is one of SwarmResearch's three idea-collapse mechanismsDoes not transfer to open-ended search
(absent in web apps)Determinism as a requirement: 95% identical code across two agents, purchased as an eval substrateDiversity as a requirement: branch isolation, a Shepherd forbidden to prescribe ideasNew per-domain requirements the web-app harness has no slot for

The two rows that fail are the informative ones, and both come from each domain's objective function. A web app has a spec, so greedy incremental progress toward it is correct. A discovery problem has no known optimum, so the same loop anchors the agent in one basin. Open-Ended Discovery Harnesses notes that autoresearch "commits improvements and reverts regressions", which is the coding-agent loop's shape exactly. Finance needs reproducibility before a number can sit under a trade ("we can't have just vibe-coded analysis be the underpinning", How Bridgewater Built an AI Analyst That Does Hours of Expert Research in Minutes), and that pulls toward typed plans and determinism. Discovery pulls the other way. HarnessBank states the general form: harness patterns transfer, and specific settings are fitted to one model's dominant failure mode.

Verdict. As posed, the question is answerable from the corpus. The enforcement, externalization and verification-loop findings generalize. The oracle, the incremental-commit loop and the determinism-versus-diversity stance are domain-specific, and the table says which way each goes. Retired. Caveats that travel with it: every domain source is single-lab or vendor (PAT is self-reported with no ablation, Leni is a vendor on its own product, SwarmResearch is n = 1 run per task), and "scientific research" here means optimization and discovery. Literature synthesis is a different system class (Deep Research Agents).

3. The valid-action set at scale: three routes, none measured#

Guarantees That Degrade at Deployment: Action-Space Soundness, Admissibility Without Effect, and a Vendor-Coupled Security Framework already settled half of this. A typed schema is capability gating and strictly weaker than enumeration. The guarantee comes back by moving the enumeration into a fail-closed runtime: ScopeGate's "unlisted tools → DENY" makes an out-of-set action unexecutable rather than unlikely. The "at scale" half was left open because the bound moved from context capacity to policy coverage, and no gate had been measured over a large tool surface.

The wiki now has a second route and a sharper account of the first route's cost:

  • Per-state enumeration by the environment adapter. Browserbase rebuilt Stagehand's act() on the lecture's exact framing: "at any given moment a page offers only a finite set of interactive elements, so the next action is really a choice from a list." Stagehand marks the elements, Jev picks the action and element, and anything under 0.7 confidence falls back to an LLM (Building Prod with Jev and LangGraph, vendor-claim). The list is computed per state, so it is as large as one page rather than the whole tool surface. This recovers the guarantee on the model's side: the chooser can only pick a listed element. But Reasoning–Acting Interleaving (ReAct) notes that it "sidesteps the large-action-space limit rather than solving it". Jev's choice cardinality tops out at 255, beyond which the vendor runs a two-stage score-then-choose. The only reported number is median latency (1.97 s → 0.46 s, secondhand). There is no success rate, and nothing measures how often the correct action falls outside the enumerated set.
  • Deferred tool loading. A large tool surface costs no context until the agent fetches definitions via ToolSearch (Agent Harness Engineering §Context Engineering for Claude 5-Class Models). This keeps a large registry within context but guarantees nothing about validity.
  • What a gate costs when it blocks. The runtime gate recovers membership, but it does not stop the model proposing out-of-set actions. The closest measurement is for a risk gate rather than a membership gate. Harness-Induced Belief Divergence (via Agent Harness Engineering §The Same Principles, Instrumented) found the model re-proposing a same-class risky action within three steps in 42 of 60 blocked cases (UnsafeRetryRate 0.700) on destructive-command Terminal-Bench tasks. Enumeration steered the chooser; a gate only censors it. Over a large surface, that re-proposal rate is the utility cost, and nobody has measured it there.

Verdict. Still partial. Every recovery route is shown on small surfaces or reported without a success metric. The deciding measurement is task success, utility and policy-authoring error for a gate or per-state chooser over a tool surface large enough that prompt enumeration would fail, and that has to come from a source. Retag #oq/now → #oq/source.

4. The orchestrator's decision burden: the mechanism is named, the quantity isn't#

The Orchestrator's Real Workload: Decision Burden, Framing Discipline, and Whether Taste Scales answered the direction: the load is higher and reshaped, with execution traded for oversight at an unfavorable rate. It named the gap as founder-side oversight load measured directly.

Two sources since then sharpen the mechanism without closing the gap:

  • The verification tax is levied by accountability, not by output volume. Alami, Paja & Tiwari's case study (Psychological Costs of AI Adoption; The Psychological Costs of Artificial Intelligence Adoption in Software Engineering, 21 interviews at one firm) defines it as "the additional cognitive and engineering effort required for practitioners to remain accountable for their work when AI-assisted." A participant puts it this way: "if we weren't responsible for the code it produced, it would be a lot faster." The page draws the consequence: the load "falls only when the practitioner needs less reconstructed understanding to stand behind the output." For the founder question, that means the load scales with how much the founder must personally stand behind, not with how many agents run. A founder is the last accountable party for everything the company ships, so orchestration moves them to the highest tax rate unless accountability is delegated to other humans or to deterministic gates. This is the "bounded only by deliberate structure" conclusion of the earlier synthesis, now with a mechanism attached.
  • The strain changes shape at the org level rather than easing. Faros's September 2026 report (AI Brain Fry Connections; The Speed Trap: 8 takeaways from our latest AI engineering research, vendor-claim) finds parallelism/too-many-threads pressure subsiding while work restarts rise 66.7%. These are period-over-period rates on developers, and "neither report measures a cognitive variable directly."

Verdict. Still partial. Both new sources study engineers inside firms, not founders, and neither measures attention. A net-load answer needs a direct measure of a founder's oversight time or error rate under agent orchestration, and only a new source can supply that. Retag #oq/now → #oq/source.

Citations#

§ end
Cited by 4
Related articles
  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…

  • Agent Systems & Harness Engineering

    Map of Content for the agent-systems domain — 53 concepts. Harness engineering, agent loops and orchestration, context…

  • Context Lifecycle Management

    Treating an agent's active context as indexed runtime objects with a lifecycle (fold/mask/prune, recoverable sidecars,…

  • Harness Shrinkage as Models Improve

    Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…

  • Agent Harness Engineering

    Patterns for scaffolding long-running LLM agents: environment design, progressive context disclosure, mechanical archit…