Sources#
Summary#
Brian Houck (Applied Scientist, DX) proposes a diagnostic vocabulary for agent failures that stop at "the model got confused" (Your agent doesn't have a model problem, DX newsletter, 2026-09-16, practitioner-opinion). A context smell is modelled on a code smell: not a bug, but a named signal of trouble that turns vague unease into something you can point at in review. Houck names six, and claims that "none of them are model limitations, and all of them survive a model upgrade." The load-bearing argument is historical rather than technical:
"I don't think AI created a context problem. I think it withdrew the error correction that had been hiding one."
Every organization runs on documentation "that would fail inspection". It works anyway because a layer of human judgment sits between bad context and the work: a new engineer asks what an ambiguous ticket means, notices that a runbook does not match production, or asks in a channel which of two documents is current. Agents don't reliably run that repair loop. "Handed the same ambiguous ticket, the same stale runbook, the same contradictory pair of documents, they proceed. Confidently, quickly, and at a scale that used to be impossible."
Evidence posture. One practitioner's argument, with anecdote and citations but no data. It comes from a vendor that sells developer-productivity measurement and uses the post to trail its CAFE(S) context-quality framework (webinar 2026-09-24; a measurement approach "over the coming weeks"). The taxonomy is useful as vocabulary. The causal headline ("your agent doesn't have a model problem") is stronger than the vault's measurements allow. See What the corpus's measurements say below.
The six smells, mapped onto the vault#
| Smell | Houck's definition | Where the vault already treats it |
|---|---|---|
| Confident hallucination | Untruths stated with the same fluency as truths; a model can separate supported from unsupported claims only if its context gives it a way to. Houck's example is the Mata v. Avianca sanctions, which he reads as a workflow failure: nothing downstream of the fabrication was designed to catch it | Failures That Look Like Success: fluency survives skim review |
| Specification ambiguity | Faced with an obviously underdetermined request, an agent asks. Faced with a semi-ambiguous one, "it picks an interpretation and proceeds", and every later step compounds the first misreading | Unknowns as the Agentic Bottleneck (the too-vague failure); Configurable Human Participation (measured: clarification wins only when the agent decides when to ask) |
| Stale guidance | Context that accurately describes a system that no longer exists. A human senses the mismatch; an agent treats it as current | Code as Source of Truth (docs stale at high throughput); Agentic Work Systematization (stale skills "executed rather than read") |
| Lost in the middle | Signal buried under too much context; "the agent has everything it needs and still fails" | Context Window Smart Zone (the Martin section: position, not occupancy) |
| Lost in the details | Implementation detail without the reasoning behind it. Houck cites Margaret-Anne Storey's intent debt, recorded rationale disappearing from a system's artifacts. His example: a 300 ms timeout, and whether it came from a benchmark, an incident or "somebody's afternoon guess" | Agentic Technical Debt (re-derived intent); Rationale as a Dated Record: Where the Why Lives for the Next Reader |
| Weekend runaway | No stated definition of success, so the agent retries, widens scope and burns tokens on an objective that may not be achievable | Loop Engineering (/goal: a written, separately checked stop condition); Evals as Product Spec (evals as "what done looks like") |
None of the six is new to the vault. What is new is the single frame: they are one failure class seen from the context side, not six unrelated problems. Houck's own point is the same one made about the field: requirements engineering, documentation drift, information retrieval and knowledge management each own one smell, and they have always existed.
The new-colleague test#
Houck's only prescription is deliberately low-tech. Before delegating, ask whether a competent new colleague could complete the task with what you are making available: "Not a senior engineer who already knows your system. A capable person with no additional access to what's in your head." Each smell produces a question that colleague would ask: Which interpretation did you mean? Is this runbook still current? Which of these fifty pages matters? What does done look like? Those questions are the repair loop. The test works by making the delegator ask them before the agent fails to.
This is the pre-delegation cousin of Thariq Shihipar's unknowns catalog. Thariq has the agent interview the human. Houck has the human simulate an outsider. Both place the binding constraint on what the human failed to write down rather than on the model, and they are independent practitioners (one frontier-lab, one measurement vendor) arriving at it within three months. They differ on time: Thariq says the bottleneck became the human with Fable-class models, while Houck says it was always the context and humans were hiding it.
What the corpus's measurements say#
The vault has three controlled or large-sample results that bear on the headline claim. They do not line up behind it.
- For: context is the modal cause of rejected agent review comments. In Cynthia et al. (
empirical, 470 argued-but-unresolved comments), the largest genuine-rejection category is Intentional Design Decision (23.8%): the code was deliberate and the agent could not see why. That is "lost in the details" measured, and hallucination proper is 4 of 470. - Against: removing the context file doesn't lower correctness. Khatri 2026 (
empirical, 288 test-evaluated runs) ablatesAGENTS.mdentirely with no measurable correctness loss. Its near-misses fail on implementation skill: a precision bug, a wrong retry pattern, a rule read and then miswired. Those are model limitations in exactly Houck's excluded sense. The two findings reconcile narrowly: Khatri tests whether adding a context file fixes failures, not whether bad context causes them, and his one file that paid carried an environmental fact (test-suite cost) that the new-colleague test would have demanded. But "your agent doesn't have a model problem" as a universal does not survive this. Weighted by tier, the empirical ablation beats the opinion: some agent failures are model failures, and adding more context will not fix them. - Mixed: "lost in the middle" is not a context property. Houck files it as a context smell that survives a model upgrade. The vault's evidence treats it as a per-model property: Eliav 2026 (
empirical) finds the recall wall moves by model at identical token counts, and the failure near the ceiling is refusal rather than fabrication. The context-side remedy (shorter, more selective context) is real. The claim that the smell is model-independent is contradicted, and it is the weakest of the six as a pure context problem.
Net reading: the taxonomy survives as vocabulary and the repair-loop argument is plausible, but the headline is a framing, not a finding. Treat "check the context before blaming the model" as a good triage order, not as a result.
The DORA amplifier as an information problem#
Houck offers one reinterpretation worth keeping. DORA 2025's finding that AI amplifies an organization's existing strengths and weaknesses is usually read as an organizational-capability claim. He suggests part of it is informational: teams with clearer specs, better docs and more accessible knowledge "have less damage for the human buffer to absorb in the first place." He labels this a speculation, and it is untested. It adds a candidate mechanism to the DORA-versus-Faros dispute on Telemetry vs. Survey Measurement, where whether maturity protects at all is the contested point.
Connections#
- Unknowns as the Agentic Bottleneck: the elicitation protocol for the same constraint. Specification ambiguity is Thariq's too-vague failure, and the new-colleague test is his interview run by the human on themselves
- Code as Source of Truth: stale guidance is the failure Fung's rule exists to prevent; the repo is the one artifact inside the update loop
- Agentic Technical Debt: lost in the details is the debt's mechanism. Storey's intent debt and the playbook's re-derived intent are the same disappearance of rationale
- Failures That Look Like Success: confident hallucination and specification ambiguity both produce work that "looks confident the entire way down", the signature that page catalogs
- Agent Context Files: the artifact where most of these smells live; Khatri's null bounds how much fixing the file can buy
- Agentic Work Systematization: stale guidance at the skill layer, where Gao et al. find reused skills rarely maintained
- Configurable Human Participation: HAS-Bench measures the repair loop in reverse, an agent choosing when to ask. Houck's semi-ambiguous case is the one where it chooses not to
- Agent Review Comment Resolution: the empirical case for the thesis. The modal rejected agent comment misread intent the agent could not see
- Loop Engineering:
/goal's written, separately checked stop condition is the direct cure for weekend runaway - Evals as Product Spec: evals as the definition of done; the missing artifact behind weekend runaway
- Telemetry vs. Survey Measurement: Houck's informational reading of DORA's amplifier finding enters that page's maturity-protection dispute
- Context Window Smart Zone: where lost in the middle is measured, and where it turns out to be a per-model property
- Verification as the New Bottleneck: the smells name the upstream failure. Verification asks whether output is right, while the smells ask whether the input ever said what right was
Open Questions#
- Do the six smells predict failure? Houck promises a measurement framework (CAFE(S)). Falsifiable: rate context for the six smells before delegation and test whether flagged tasks fail at a higher rate than unflagged ones, controlling for task difficulty.
- Does the new-colleague test actually "catch a surprising amount"? No source measures how many agent failures a pre-delegation outsider check would have prevented, or its cost in delegator time against the failures avoided.
Sources#
- Your agent doesn't have a model problem: Brian Houck (Applied Scientist, DX), "Your agent doesn't have a model problem", DX newsletter (Substack), 2026-09-16, ~1,600 words,
practitioner-opinion. No data; citations are to Mata v. Avianca, an LLM-ambiguity paper (arXiv 2304.14399), Meta's tribal-knowledge mapping, Liu et al.'s lost-in-the-middle (TACL), Storey's intent debt (ACM Queue), Cemri et al.'s multi-agent failure taxonomy (arXiv 2503.13657) and DORA 2025. COI: DX sells engineering measurement and uses the post to promote its CAFE(S) framework. The one figure (six cards) is transcribed in the raw
Cited by 13
- Unknowns as the Agentic Bottleneck×3
your agent doesnt have a model problem — Brian Houck (DX), 2026-09-16, practitioner-opinion, vendor…
- Agent Context Files
Context Smells — a six-item vocabulary for what goes wrong inside the file (stale guidance, lost in…
- Agent Review Comment Resolution
Context Smells — the 23.8% Intentional Design Decision category is Houck's lost in the details…
- Agentic Technical Debt
Context Smells — lost in the details is this page's mechanism named as a context smell:…
- Agentic Work Systematization
Context Smells — stale guidance at the skill layer: a skill that accurately describes a workflow…
- Code as Source of Truth
Context Smells — stale guidance (context that accurately describes a system that no longer exists)…
- Configurable Human Participation
Context Smells — Houck's specification ambiguity smell is the case this benchmark's clarification…
- Evals as Product Spec
Context Smells — weekend runaway (no stated finish line, so the agent never stops) is what a…
- Failures That Look Like Success
Context Smells — two of Houck's six smells produce this page's signature from the input side:…
- Loop Engineering
Context Smells — /goal's written stop condition is the direct cure for Houck's weekend runaway…
- AI Coding Practice
Context Smells — Brian Houck's (DX) vocabulary for recurring agent-context failures, by analogy to…
- Open Questions Backlog
Context Smells ×2 (oldest 4d) — Do the six smells predict failure?
- Telemetry vs. Survey Measurement
Context Smells — Houck's informational reading of DORA's amplifier finding: part of what "strong…
Related articles
- Verification as the New Bottleneck
Fiona Fung: coding is no longer the bottleneck — verification, review, maintenance are; shift-left; TDD loses its tax;…
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Security Debt of Agent-Generated Code
Sakib, Banik & Jadliwala (UTSA, arXiv 2607.12428): LLM-as-judge + manual coding over 16,112 high-risk file changes in 4…
- Agent Context Files
The cross-vendor markdown-as-control-plane pattern: repo-versioned plaintext (CLAUDE.md / AGENTS.md / SOUL.md / WORKFLO…
- Agentic Technical Debt
Debt that *compounds* (not just accumulates) because each agentic-coding session re-derives architectural decisions wit…
