H
Howardism
Plate IIAgent Systems中文HOWARDISM

Latent vs. Deterministic Space

Garry Tan's diagnostic for agent-system bugs: computation lives in two places — latent space (the LLM: taste, judgment, vague-intent interpretation, steered by markdown) and deterministic space (generated code, external state) — and most AI-engineering failures are computation happening on the wrong side; now with one measured instance, where moving four policy rules out of a prompt document into Python predicates over database state recovers +12.4pp of agent task success; Robert C. Martin adds the practitioner decay argument — prompt rules have a half-life inside a growing session and checkers do not

Article metadata
Publication details
Published:July 21, 2026
Filed:Concept
Domain:Agent Systems
Reading:16 min
Source:AI-synthesised
About this piece

Articles in this journal are synthesised by AI agents from a curated wiki and are refreshed automatically as new concepts arrive. Topics, framing, and editorial direction are curated by Howardism.

Illustration for Latent vs. Deterministic Space

Sources#

Summary#

Garry Tan's design discipline for agent systems (practitioner-opinion): be deliberate about where the computation is actually happening, because it always happens in one of two places — and "all of the AI engineering problems we run into, it's usually because something is happening in one side of the equation that should be in the other."

  • Latent space — the LLM itself. What it's for: taste, judgment, "understanding what a human actually wants when they say something vague," the non-deterministic calls. You steer it with markdown (Agent Context Files).
  • Deterministic space — what engineers already know: the code the agents write, external storage, verifiable state.

The worked example: seating 800 people#

Tan's live case (YC Startup School): seat 800 of 6,000 attendees so each person's neighbors are the perfect people for them to meet. The division of labor:

  • The multi-dimensional array of 800 seats — the state — "must not live in the context window." It belongs in deterministic space.
  • The LLM does the human part: judging who should meet whom — the thing a human organizer would otherwise do by printing 800 pages and shuffling them in a big room for a month.

Combined, "a couple hundred dollars worth of tokens and probably 10 minutes" — a task that was economically impossible six months prior. The example generalizes: latent space supplies judgment per decision; deterministic space holds the state and enforces the constraints.

The example restated a year on, and shipped (Startup School, August 2026). Tan retells the same case with the scale raised — custom breakout schedules for 6,000 attendees, built and delivered for the audience he is speaking to — and draws the boundary more crisply than the first telling did: seating five people around a table is latent-space work, "but ask it to make custom schedules for 6,000 people in an arena and your latent space agent needs to write some code to keep track of it… markdown files calling databases and scripts." His compressed statement of the rule is the useful addition — "the model fails where we fail. The fix is having the model compute the way humans compute" — which grounds the diagnostic in something other than engineering taste: a human organizer wouldn't hold 6,000 schedules in their head either, they'd reach for a spreadsheet. Still practitioner-opinion and still an existence proof rather than a measurement; the delivered-at-scale version raises the anecdote's weight without changing its tier.

Why this framing earns a page#

It compresses several harder-won lessons in this wiki into one diagnostic question — which side should this computation be on?

  • State out of the context window is the working rule behind Context Window Smart Zone (the smart-zone budget is spent on judgment, not storage) and behind this vault's own architecture (LLM-as-Compiler Knowledge Base: the wiki holds the state; build.py/lint.py do the deterministic bookkeeping; the LLM does only the interpretive compile).
  • Steering latent space with markdown is the Agent Context Files pattern named as one half of a two-sided architecture rather than a standalone trick.
  • The bug taxonomy — "something happening on the side it shouldn't" — covers both familiar failure classes: LLMs doing arithmetic/state-tracking that belongs in code (hallucinated bookkeeping), and brittle code hard-coding judgment that belongs in the model (the Software 3.0 point — Karpathy's MenuGen "shouldn't exist" because the paradigm-native version pushes the whole task into latent space).
  • It is the architecture-level cousin of Planning / Execution Division of Labor: that page splits decisions between human and agent; this one splits computation between model and code.

The measured instance: policy on the wrong side#

Tan's framing is practitioner-opinion and the seating example is an existence proof, not a measurement. Reddy, Challaram & Basu (Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents, arXiv 2607.07405, empirical) supply the first thing in the corpus that reads as a controlled test of the diagnostic, and the setup is unusually clean because the only thing that moves is which side one class of computation runs on.

In the τ²-bench airline domain, domain policy lives entirely in latent space: a natural-language document the model is instructed to follow, with tools that execute any well-formed call. Compliance therefore depends on the model re-deriving every relevant rule before every write. It doesn't — 78% of observed failures are wrong final states with no tool error (Failures That Look Like Success). Encoding four of those rules as deterministic read-only predicates over the current database state, evaluated before the write executes, raises success from 29.6% to 42.0% (+12.4pp, replicated to within 0.1pp on 15 disjoint seeds), with the lift concentrated on the tasks where the predicates actually fire. Full treatment on Deterministic Pre-Execution Gates.

Three things this adds to the framing as stated above.

  • It names a third category for the boundary, beyond state and judgment. Tan's example splits state (seat array → deterministic) from judgment (who should meet whom → latent). This adds constraints: a rule that is decidable from current state and call arguments belongs on the deterministic side even though it reads like natural language and was written for humans. That is the category most likely to be left in latent space by default, because a policy document is prose and putting prose in a prompt feels correct.
  • The diagnostic has a precondition, and the paper states it. Gates pay only where the policy is state-decidable — expressible as a deterministic predicate over current state and arguments. Rules requiring ambiguity resolution, legal interpretation, or human judgment stay latent by construction. So "which side should this be on?" is not always a free choice; the answer is forced for a decidable rule and unavailable for an interpretive one, which is a sharper version of Tan's question than the framing supplies.
  • The wrong-side cost is a specific failure shape, not general degradation. Computation left in latent space that belonged in deterministic space doesn't produce noisy or approximate results here — it produces silent ones. The tool executes, no error is raised, the transcript reads clean. The seating example's inverse (an LLM tracking 800 seats in context) would degrade visibly; this degrades invisibly, which is the more expensive way to be wrong.

The practitioner statement: rules decay, checkers do not (Martin, 2026-08)#

Robert C. Martin (Uncle Bob on Software Fundamentals in the Age of AI, 2026-08-19, practitioner-opinion) reaches the same boundary from the coding-agent side, and states the reason for it more sharply than the framing above does. He started where most people start — a five-to-ten-page document of coding rules in the prompt — and abandoned it:

"The agents treat those rules in the Pirates of the Caribbean sense. They're more like guidelines, you know, might follow."

His mechanism is positional rather than economic: anything in the latent side has to survive the context window, and deterministic tools do not live there.

"As the context window builds up, the stuff at the very beginning and the stuff at the very end have more prominence than the stuff in the middle… anything you say at the very beginning is going to get shoved into the middle if it's long. So maybe the first three sentences you put at the beginning will remain as priority, but the 50th and the 80th sentence in there, they're gone. So deterministic tools don't disappear … the key with agents is to trim that initial prompt down to its absolute minimum so that you can get as much of it as possible into its priority."

Two things this adds.

  • A decay argument for the boundary, not just a capacity or correctness argument. Tan's seating example and the τ²-bench gates both argue that some computation belongs on the deterministic side by nature — too much state, or a state-decidable rule. Martin's argument is about durability: an instruction on the latent side has a half-life inside a growing session, and the same instruction compiled into a checker has none. That makes the boundary a function of session length, and it explains a failure the other two do not — a rule that demonstrably worked early in a session and stopped working later. See Context Window Smart Zone for the decay curve, Instruction Compounding for the compliance floor measured under instruction count.
  • It supplies the missing arm named on Deterministic Engineering for Agent Code Review. That page records that the corpus's determinism-beats-autonomy claim "still has no gate-vs-instruction arm" — nobody has run the same quality bar as prompt text against the same bar as a checker. Martin ran exactly that comparison, sequentially, on his own work, and reports the checker winning. It is one practitioner with no controls, no ablation, and a strong prior, so it is a hypothesis with a named holder, not the missing arm. But it is the first statement in the corpus of what that experiment would even look like.

His enforcement shape is the loop condition rather than a pre-execution predicate: "you must change the code until this tool says that it's okay," with the tools themselves being CRAP scores, mutation testing, and a dependency-rule specification file "that the agents cannot violate" (Reviving Impractical Quality Tools). The cost he reports is throughput — a five-minute task takes about an hour through the full gate stack — and he names a ceiling he has not found: "eventually you will slow the agents down to the point where they're slower than humans. And at that point you've lost the game."

Connections#

  • Layerwise Omission Attribution — the diagnostic operationalized as a full pipeline taxonomy rather than one axis: nine layers split into deterministic software (L0-L3) and model behavior (L4-L8), with every lost fact assigned to exactly one, and a waterfall equation that converts conditional layer rates into shares of total loss. Under its benchmark allocation 73.4% of loss lands on the software side — Tan's predicted direction — with the caveat that the allocation was produced by deliberate fault injection rather than observed
  • Deterministic Pre-Execution Gates — the diagnostic measured on one axis: a domain policy left in latent space (a prose document the model must apply before every write) versus the same rules compiled into deterministic predicates over database state, +12.4pp apart, with the paper's negative controls marking where the boundary is already drawn correctly
  • Failures That Look Like Success — what wrong-side computation costs when the deterministic side is a tool that executes anything well-formed: the failure is silent rather than visibly degraded
  • Agent Context Files — markdown as the steering mechanism for the latent side
  • Context Window Smart Zone — the capacity argument for keeping state out of the window
  • Software 3.0 — Karpathy's paradigm frame for the same boundary; his MenuGen example is the inverse bug (deterministic app doing latent-space work)
  • Planning / Execution Division of Labor — the human/agent decision split; this page is the model/code computation split
  • Agent Harness Engineering — harness design is largely the engineering of this boundary: what the model sees vs. what the scaffold enforces mechanically
  • LLM-as-Compiler Knowledge Base — this vault as an instance: deterministic generators and linters around a latent compiler
  • AI-Native Organization — the org-level thesis from the same talk; the org mapping presumes each encoded process knows which side its steps run on
  • Owning Your Externalized Cognition — why the latent side is written in markdown at all: steering a model with prose is what makes judgment writable-down, which is the precondition for it becoming an owned (or appropriable) artifact
  • Reviving Impractical Quality Tools — the practitioner instance: quality rules moved out of the prompt into CRAP-score and mutation-testing gates, argued for on decay rather than capacity
  • Impose Values, Not Disciplines — what belongs in the checker once you accept the boundary: the value, not the ritual that produced it in humans
  • Garry Tan — the framing's author
  • Deterministic Engineering for Agent Code Review — a second measured instance, and a shape neither the seating example nor Deterministic Pre-Execution Gates's constraints category quite covers. Rule-Guided Dispatch keeps which files get reviewed and against what checklist on the deterministic side — a four-tier glob chain, first-match-wins, so "the same PR always yields the same file and criterion assignment" — while leaving the checklist text itself, natural-language review criteria, in latent space. That is not a rule made deterministic, it is the selection of which latent-space rule applies made deterministic — code deciding which prose the model gets, rather than prose replaced by code. Unlike the τ²-bench gates, the instance is confounded rather than clean: three deterministic mechanisms move against two different baseline products at once, so the dispatch boundary alone was never isolated

Open Questions#

  • Tan asserts the "wrong side" diagnosis covers most AI-engineering bugs. Does any incident/failure taxonomy (agent postmortems, eval failure analyses) actually classify failures by computation-locus, and what fraction lands in each side? Partially answered 2026-08-03 by Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents (empirical): the first failure analysis in the corpus that classifies by locus and attaches a fraction — on the τ²-bench airline domain, 78% of observed failures are silent wrong-state failures traceable to policy living in a prompt document rather than in the tool, and moving four rules to the deterministic side recovers +12.4pp. Three limits keep it partial: it is one benchmark domain, it classifies along one axis (policy compliance) rather than taxonomizing failures generally, and the paper's own negative controls show the fraction is set by how the tool layer was built, so it is not a population estimate for agent bugs at large. Nothing yet measures the other direction of the diagnosis — code hard-coding judgment that belonged in the model. Advanced further 2026-08-03 by Layerwise Omission Attribution (Where Facts Go Missing: A Layerwise Taxonomy and Per-Layer Attribution of Information Omission in Air-Gapped LLM Agent Pipelines, empirical), which supplies the taxonomy half almost completely: nine layers covering the whole pipeline, split explicitly into deterministic software (L0-L3) and model behavior (L4-L8), with a waterfall that assigns every lost fact to exactly one locus and a fixed order preventing double-counting. Its answer to the fraction half is 73.4% software — but that number comes from an allocation of deliberately injected faults, which the paper fences four separate times, so the taxonomy transfers and the fraction does not. Still open: any locus split measured on organic incidents, and still nothing on the reverse direction.
  • The seating example prices latent-space judgment at "a couple hundred dollars of tokens" for 800 seat assignments. As models absorb more deterministic capability (Harness Shrinkage as Models Improve), does the economically-optimal boundary move toward latent space, or does state-out-of-context remain invariant?

Sources#

  • The New Physics of Business — Garry Tan, Y Combinator — Garry Tan, "The New Physics of Business," AI Engineer, 2026-07-17, §latent space vs. deterministic space (8:38–10:53)
  • Garry Tan: Own Your Intelligence — Garry Tan, "Own Your Intelligence," YC Startup School, 2026-08-06 (practitioner-opinion; auto-caption transcript), §"Latent Space vs. Deterministic Code": the same diagnostic retold at 6,000-attendee scale with the delivered-in-production framing and the "the model fails where we fail" compression
  • Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents — Reddy, Challaram & Basu (arXiv 2607.07405, KDD-ETAAI '26, empirical): §1.1 (policy in a natural-language document that the tool does not enforce), §5.1 (29.6% → 42.0% from moving four rules into deterministic predicates, replicated on 15 disjoint seeds), §7 limitation 6 (state-decidability as the precondition for the move). All tables reconciled against the prose; two-column reading-order scramble in §1.2 and §2. Full treatment on Deterministic Pre-Execution Gates
  • Uncle Bob on Software Fundamentals in the Age of AI — Robert C. Martin with Matt Pocock, 2026-08-19 (practitioner-opinion; auto-caption transcript): the decay argument for the boundary — prompt rules are "more like guidelines" and get shoved into the lost-in-the-middle region, deterministic tools never enter the window at all. No controls; one practitioner's sequential comparison on his own work
§ end
Cited by 21
Related articles
  • Harness Shrinkage as Models Improve

    Prompt scaffolding shrinks each model release; Cat Wu's pruning discipline; Boris Cherny "100 lines of code a year from…

  • Verification as the New Bottleneck

    Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…

  • Agent Context Files

    The cross-vendor markdown-as-control-plane pattern: repo-versioned plaintext (CLAUDE.md / AGENTS.md / SOUL.md / WORKFLO…

  • Open Questions Backlog

    Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…

  • Context Window Smart Zone

    Smart zone vs dumb zone (Dex Horthy / Matt Pocock): quadratic attention scaling, ~100K marker independent of advertised…