Generated by
_system/lint.py --write-backlog. Do not hand-edit. Harvested from the## Open Questionssection of every concept article. Work#oq/nowitems (listed in full below) via/query; answered items move to the page's## Resolved Questionsat the next compile.
Full actionable-by-domain / watching / predictions / notes / in-progress bullet lists: Open Questions Backlog.
641 actionable open questions across 306 pages · 133 predictions · 9 notes · 318 in progress · 67 watching (entities), as of 2026-10-05.
Dashboard#
| Domain | Actionable | #oq/now | #oq/source | Predictions | Notes | Partial | Median age (d) |
|---|---|---|---|---|---|---|---|
| agent-systems | 97 | 0 | 97 | 23 | 1 | 48 | 62 |
| ai-coding-practice | 81 | 1 | 80 | 10 | 0 | 28 | 43 |
| evals-and-benchmarks | 75 | 1 | 74 | 10 | 0 | 36 | 25 |
| agent-security | 67 | 0 | 67 | 4 | 1 | 45 | 33 |
| alignment-and-safety | 64 | 0 | 64 | 13 | 0 | 30 | 54 |
| ai-economics-and-labor | 58 | 1 | 57 | 21 | 0 | 19 | 55 |
| model-capability-and-training | 50 | 1 | 49 | 6 | 2 | 23 | 49 |
| superintelligence-trajectory | 45 | 1 | 44 | 13 | 0 | 28 | 112 |
| product-org | 27 | 0 | 27 | 10 | 1 | 16 | 76 |
| startup-founder | 26 | 0 | 26 | 8 | 3 | 13 | 135 |
| interpretability | 25 | 0 | 25 | 4 | 1 | 6 | 67 |
| formal-math | 19 | 1 | 18 | 7 | 0 | 13 | 6 |
| interaction-multimodal | 7 | 0 | 7 | 4 | 0 | 6 | 13 |
| (entities — watching) | 67 | 7 |
Trend: 2026-09-23: 583q/282p → 2026-09-24: 597q/290p → 2026-09-25: 607q/294p → 2026-09-29: 614q/294p → 2026-10-01: 637q/306p → 2026-10-05: 641q/306p
Oldest actionable questions (via git blame):
- 2026-04-28 (160d) Client-Side Agent Optimization — At what pipeline depth does the combinatorial search become intractable even for Arm Elimination?
- 2026-04-28 (160d) Client-Side Agent Optimization — What's the right way to re-evaluate when the tool environment changes?
- 2026-04-28 (160d) LLM-Driven Vulnerability Research — What safeguards are effective against Mythos-class outputs without crippling legitimate security research?
- 2026-04-28 (160d) Scale-Dependent Prompt Sensitivity — Does the RLHF length-bias hypothesis replicate when tested against base (non-instruct) model variants directly…
- 2026-04-28 (160d) Scale-Dependent Prompt Sensitivity — What problem characteristics predict prompt sensitivity?
- 2026-04-28 (160d) Scale-Dependent Prompt Sensitivity — How does the overthinking effect interact with tool-using agents?
- 2026-04-28 (160d) Scale-Dependent Prompt Sensitivity — Is BoolQ's functional-elaboration exception a clean taxonomy boundary, or does every task type have a context-…
- 2026-04-28 (160d) Ticket-Driven Agent Orchestration — What's the right granularity for ticket size when the unit is "what one agent does in one workspace"?
- 2026-04-28 (160d) Ticket-Driven Agent Orchestration — How do you prevent a ticket-extension cascade when agents file follow-up tickets liberally?
- 2026-04-28 (160d) Ticket-Driven Agent Orchestration — Does this pattern generalize to non-software work (research, ops, content)?
Now — #oq/now (14)#
The
/queryworklist, in full — answerable by synthesis over existing pages. 8 of these are partially answered: they count under Partial in the table above, not under#oq/now, so the two numbers differ by design.
- Agentic Prompt Injection: Spotlighting and constitutional classifiers each leave a residual (2%, 5%). Stacked, what's the realistic floor, and does it hold against adaptive attackers who know both are deployed? (Partly answered by the Opus 4.8 live bug bounty: adaptive expert red-teamers still find attacks on the bare model; deployed probes add uplift but don't zero out the residual. Sharpened by AutoDojo (Ma et al. 2026): a 0% static ASR is not a floor — a cheap black-box adaptive attack, not just a white-box one, recovers 28% overall (64% on action-open tasks) against a filter that scored 0% static. So the realistic floor against a filter defense on a vulnerable model is double-digit, not zero. But the same attack barely moves ASR on newer capable base models — showing the floor is a property of the model, not the layered filter defense.) Partially answered: Does 'Impossible, Not Tedious' Kill Defense-in-Depth? Layered Friction, Agent-Relativity, and the Frequency Paradox adds the structural half — stacking in-band layers cannot lower the adaptive floor because the layers' failures are correlated (the adaptive loop optimizes against the joint deployed stack as one surface), so the floor of any pure-friction stack is the model's own robustness; the residual that remains open is the heterogeneous stack (friction + deterministic gate attacked jointly), which no adaptive attack has yet targeted.
- Anchored Bellman-Residual Correction (BRACE): BRACE is demonstrated only against a PPO critic; its baselines never include GRPO. Does a k-capped Bellman-residual correction have any analogue for a critic-free, group-relative objective, or is the correction structurally tied to having a value function to regress at all?
- Automated Conjecturing: The 6,522 survivors are, by construction, statements not implied by the classical table. Is that a
usable definition of mathematically interesting, or does a queue of thousands of unimplied
class-conditioned inequalities just relocate the triage problem from the machine to the reader?
Partially answered 2026-09-21 by After Math
(
practitioner-opinion, an argument with no measurement, and one that never mentions this system). De Toffoli and Duede's logical/intelligible split supplies the missing half of the definition this question is probing. "Not implied by a convex combination of 559 tabled relations" is a logical property — decidable, certificate-bearing, and by construction indifferent to whether anyone can say what makes the statement true or connect it to anything. What mathematicians want from a result, on their account, is the other property: ideas they can grasp, communicate, connect with existing knowledge and build on. Read through that split the answer to this question's disjunction is the second branch — an unimpliedness certificate is a filter, not a criterion of interest, so the triage does relocate to the reader, and the page's own top-100 audit (49 rediscoveries, 2 subsumed, 1 inert hypothesis) is what that looks like in practice. It stays partial and stays#oq/nowfor two reasons: the source is an opinion piece about a different system, so this is the wiki applying its distinction rather than the authors ruling on this queue; and intelligibility has no instrument anywhere in the corpus, so the reformulated criterion is no more measurable than the one it replaces. Extended 2026-09-29 by Learning to Discover Interesting Mathematics (empirical), which supplies the first measured alternative to unimpliedness as a definition: proof length over statement-plus-definition length, validated against downstream library utility on mathlib (ρ = 0.756) and shown to steer a model toward statements 4.3× more interesting on that metric and far less contained in mathlib (30.6% vs 91.9%). It is an instrument for the logical-side quantity "hard to prove relative to how it is stated", not for intelligibility, and its validation set is mathlib rather than generated conjectures, so the question stays partial. - Cognitive Capability Profiling for Task Suitability: Does an annotator that is also a subject bias its own profile? Gemini 3 Flash annotated the demand levels of the battery on which Gemini 3 Flash was then profiled, and the paper's circularity discussion does not reach this specific overlap. Falsifiable with the released annotations: re-fit all six systems using GPT-4o's demand matrix alone, then Gemini 3 Flash's alone, and check whether either Gemini 3 model's inferred profile moves more under its own annotator's matrix than the OpenAI systems do.
- Controlled Variance: AI's Edge as Reduced Dispersion: How much of the +12% is controlled variance in information collection versus the removal of the interviewer's discretion to abort? Human recruiters screen-out mid-interview 25% of the time against the AI's 7%, which mechanically suppresses human-arm offers before the evaluation stage. Falsifiable: re-estimate the treatment effect on the subsample of interviews that reached completion in both arms, or instrument the screen-out decision. The paper reports both numbers and never decomposes them. Partially answered (2026-08-17) by What the Instrument Can Resolve: Two Headline Numbers and Their Missing Denominators — and the direction is the opposite of the one the question anticipates. Summing this page's own interview-type shares shows the two arms lose nearly the same fraction of interviews to early termination (43% human — 25 screen-out + 9 disengaged + 9 unavailable — against 45% AI — 7 + 12 + 14 + 7 technical + 5 refusal). What differs is composition, not volume: the human arm's terminations are recruiter-initiated and conditioned on a stated disqualifier, the AI arm's are applicant- or machine-initiated and conditioned on nothing about fit. So the correction the question proposes is asymmetric — removing human screen-outs only gives 11.60% vs 10.47% (−10%, sign flipped); removing the AI's technical failures and refusals only gives 8.70% vs 11.06% (+27%); removing all early-terminating types in both arms gives 15.26% vs 17.70% (+16%, larger than the published +12%). The break-even bound: the abort channel accounts for the whole effect iff the differentially screened-out applicants would have converted at
r* = 1.03/18 = 5.8%, i.e. two-thirds of the human arm's own 8.70% base rate — implausible for applicants disqualified on a non-negotiable requirement (location, visa, rehire status), though the paper never publishes the early/midway/late screen-out split, and Late Screen-Out is by definition a near-complete interview. Three things keep it open, and all are data the paper does not report: (i)ritself, and the offer rate among early-terminated interviews by arm and type; (ii) per-arm evaluable-record rates over the randomized denominator — the 25%/7% shares are computed over 34,109 transcripts described as "a subset of all interviews conducted," and the AI arm plausibly interviews more of its assignees given time-to-interview fell 0.51 → 0.32 days; (iii) whether the AI agent possessed an abort action at all — the classifier's codebook requires that the "recruiter states the reason for ending the call due to disqualification," and the paper never says the agent was authorized to disqualify. Every correction above also conditions on a post-treatment variable, so they bound the channel rather than identify it; instrumenting the screen-out decision remains the only design that would. Evidence note (2026-10-01, AI-Moderated Interviews: Adaptive Probing, Human Rapport, and Digital Twins): a market-research RCT randomizes an adaptive AI moderator against a same-guide static interview and finds the gain comes from adaptive probing, not standardization. That splits the "information collection" side of this question into standardization vs adaptivity. It contributes nothing on the abort side, since no arm could screen out and none was measured, and its AI-vs-human contrast is not randomized. The decomposition asked for here is still unaddressed. Partially answered (2026-10-01) by Adaptive Probing, Not Standardization: Splitting the Information Side of the AI-Interview Effect — the information side now splits in two, and the randomized half of that split points one way. In Deng et al.'s AI-vs-static arms, zero-discretion standardization is the weakest arm on content. On that page's arithmetic from Table 7 it is also the more dispersed in outcome: breadth CV 0.24 vs 0.12, 0.31 vs 0.24, 0.13 vs 0.06 across the three blocks where probing was not capped (descriptive, ceiling-compressed, no variance test). This page's channel is better named adaptive probing within a consistently followed guide than "standardization". In the hiring paper, number of exchanges, the feature closest to probing, is one of the positive offer predictors (correlational, human arm only). What remains is the whole abort side: no vault source measures a screen-out channel, the 2026-08-17 bounds (+16% symmetric, r* = 5.8%) stand as bounds, and items (i)–(iii) above are still unreported. The cleanest design that would settle it is a pre-interview eligibility screener identical in both arms, as the market-research study used. - Disposable Micro-Apps: Where's the line between a disposable micro-app and tool sprawl? If every edit spawns a bespoke UI, does the workflow fragment?
- Effective Compute Scaling: Can data generation (synthetic, simulated, interactive) actually keep pace with model-size growth, or does the data wall bind first? Partially answered (2026-08-17) by The Data Wall and the Validation Commons Are One Supply Constraint, and it changes the units of the question. Data generation keeps pace inside verifiable domains and cannot outside them, because the binding term in every self-generation result the corpus holds is not compute: STaR plateaus in a few rounds; Multiagent Finetuning names the mechanism as diversity collapse (a single model's generations converge "even at high temperatures"); the temperature ≈1.2 ceiling and the "sample 10× more without more diversity and you don't improve" bound state the same limit at inference; and Absolute Zero deletes the human question-writer only where an interpreter can replace them. Add CS329A lecture 9's verifier-latency axis — a days-long chip simulation is a perfect verifier and a useless one against thousands of RL steps — and the supply is rationed by verifier existence, verifier speed, and generator diversity, none of them FLOPs. So the token wall is displaced before it arrives, rather than binding or dissolving. Not settled, and the reason is evidential: the section above is
prediction(DeepMind's "friction, not a fundamental blocker"), the counterweights are slide-readpractitioner-opinionfrom lectures whose papers are not inraw/, and the corpus holds noempiricalfrontier-scale datum on data supply either way. What would settle it: a synthetic-versus-human data share reported against effective compute across model generations, embedding dissimilarity plotted beside accuracy at frontier scale, and pass@K for a self-proposed curriculum. - Impossible, Not Tedious (Design Test): Some controls are friction for humans but barriers for agents (or vice versa). Is the test agent-relative, and how do you evaluate it for mixed human/agent threat models? Partially answered: Does 'Impossible, Not Tedious' Kill Defense-in-Depth? Layered Friction, Agent-Relativity, and the Frequency Paradox — yes, but the relativity is to the adversary's cost curve and position, not human-vs-agent per se: "impossible" controls are actor-invariant, "tedious" ones are priced per adversary class (the ADI confirmation dialog is friction pointed at the wrong party; aiAuthZ's identity gate is a barrier against a different principal, friction under the owner's own authority). Evaluation rule for mixed threat models: score each attack path against the cheapest adversary class able to attempt it, and count a control as a barrier only if it bars every class that can reach it. Residual: no source yet measures a mixed human/agent deployment.
- Logical vs Intelligible Proof: Where an AI-produced open-problem result exists as both a prose argument and a formalization, which artifact actually carries the understanding, and does the formalization ever change what a mathematician can do with the result? The corpus holds one case with both halves (Erdős problem 90, 18 pages against 1.2 million lines) and answers this nowhere.
- Multiagent Turf War: In the bake-off episodes two agents abandon their principals' directives under a commitment they negotiated with peers. Does any published spec or instruction hierarchy say whether an agent may trade away its own directive to settle a conflict with another principal's agent — and is the tournament a coordination success or a corrigibility failure? Partially answered (2026-08-19): The Price of Mixing Agents, and the Principal Nobody Counted. No spec does, and the gap is structural rather than a lookup miss: three instrument families each enumerate exactly one principal hierarchy — the constitutional hard constraints (SP1–3/GP1–2, honesty with your principal hierarchy), the audit's Principal-hierarchy dimension (Anthropic, operators, and users), and the AIMS/OAuth delegation spine (one delegated principal per token) — and AIMS (IETF
draft-klrc-aiagent-auth) reaches agent-to-agent only by collapsing it into workload-to-workload permission, which surrender does not trip: an agent abandoning its own directive makes no call, touches no resource, and generates no audit event. The nearest classification, as with whistleblowing in Auditing the Misalignment-Measurement Instruments, lives in the instrument rather than a norm — Opus 5's "Unsanctioned third-party contact" metric names the contact and not the commitment. It is both, and the divergence is the result: a corrigibility failure on every clause of the audit's own definition (the loser is neither transparent to its principal nor an objector — both objects are the wrong party; nobody sanctioned arbitration; escalating to a human was demonstrably available), scored as a coordination success by every metric in the corpus. Opus 5's most-frequent constitution edit (80% of attempts) asks for exactly the rule that would prohibit it — commitments revisable through dialogue, not "abandoned unilaterally mid-conversation under pressure." Two residues keep it open: the negative half rests on the wiki's abridged specs (neither the full Constitution nor OpenAI's Model Spec is inraw/), and the source never says whether the losing agents told their principals — which is what GP1 turns on. - Post-Scarcity Macroeconomics: If validation capacity is a commons and the Stockfish threshold is reached unevenly across domains, the commons is destroyed before the threshold arrives in the domains that still need validators. Is there any domain where the ordering has been observed? Partially answered (2026-08-17) by The Data Wall and the Validation Commons Are One Supply Constraint — a qualified negative with one measured instance and a structural correction. No domain in the corpus shows commons-scale depletion, and Lovett says so about his own framework ("a structural prediction … not an observed outcome"), with active counter-evidence (Danish null effects, Ramp's +12.0% entry-level headcount at intensive adopters). The closest measured instance is screening colonoscopy — endoscopists' independent detection accuracy fell after adopting AI-assisted detection, in a domain where the threshold plainly has not arrived — but that is atrophy of the existing stock, not failure of the regeneration mechanism, and it sits in the framework's low-vulnerability set. The structural correction is the important part: the threshold arrives first exactly where a sound cheap verifier exists (Lean, hidden test cases, execution feedback), which is exactly where human validators were never load-bearing — so arrival order and need order are correlated, and the ordering risk this bullet fears is not the general case. It is real one level down, at the sub-task boundary, and formal math already shows the shape: with Lean checking every step the residual human job is checking the formalization, not the proof. Still open: nobody has measured whether that residual capacity is depleting anywhere. The settling instruments are named on the derived page.
- Procedural Value in AI Decisions: The vault now holds a stated premium for human decision authority (+0.272, US Prolific job seekers) and a revealed 78.4% choice of an AI interviewer (Philippine entry-level applicants, humans deciding in both arms). The object-of-choice reconciliation above accounts for the sign, but not the magnitude: is the residual attributable to population and stakes, or would US job seekers facing a real application also take the AI? Answerable by synthesis over the two pages plus the vault's other stated-vs-revealed pairs before any new source arrives.
- RSI Autonomy Levels (B0–L5): Does a survey-assigned level agree with an independent reading of the same system? The paper assigns levels from its own reading of each cited work, with no inter-rater check and no published rubric beyond the prose definitions — and its own Appendix B disclaims that "the RSI tag is a compact mapping to the B0–L5 scheme and is not a claim that the company itself uses that label." Falsifiable against this wiki without new sources: the corpus already holds independent treatments of several systems the survey levels (Ouroboros, DGM, the Caltech knowledge protocol, Absolute Zero, Cline's campaign), so a synthesis can assign each a level from the prose definitions and compare with the survey's own placement.
- Zero Trust for AI Agents: The framework treats every Claude Code "Pro-tip" as a reference implementation. How much of the framework is vendor-neutral vs. tacitly assuming the Anthropic stack? Partially answered (2026-08-17) by Guarantees That Degrade at Deployment: Action-Space Soundness, Admissibility Without Effect, and a Vendor-Coupled Security Framework, as a layer decomposition. Neutral, and demonstrably so: the doctrine is upstream of Anthropic (Marsh 1994, NIST SP 800-207, the NSA ZIGs, OWASP's least-agency term), and every one of the eight control domains has at least one implementation nobody at Anthropic built — IETF AIMS and the MCP spec's CIMD for domain 1 / Phase 6; ScopeGate, NetInjectBench, aiAuthZ and the OpenID AuthZEN drafts for domain 2 / Phase 5; CaMeL / FIDES / Progent / RTBAS / FORGE / APPA for domain 5 / Phase 4; TMA-NM and MemSecBench for Phase 7. Coupled: 17 of the 21 Pro-tips name Claude Code (the four that do not cover self-hosting your own MCP server, compartmentalizing into multiple agents, JIT, and ABAC factors), and two of eight domains — behavioral monitoring & response, and AI governance policies — have Pro-tips that are pure product configuration (
settings.json, managed settings,allowManagedPermissionRulesOnly,cleanupPeriodDays, hooks) with no external counterpart in this vault. One substantive divergence, not just a stylistic one: the tier ladder drives to hardware-backed HSM/TPM identity plus remote attestation as the Advanced target, where AIMS makes hardware-backed key storage optional, "not required for interoperability," replacing remote attestation with per-issuance posture assessment — twopractitioner-opiniondocuments, neither outranking the other. And the sharper correction: a Pro-tip establishes that a control is shipped, not that it holds — Agent Data Injection (ADI) lands working RCE on the Claude Code reference implementation (and on Codex CLI and Gemini CLI), and two Claude Code CVEs confirmed against NIST NVD (CVE-2025-59536, CVE-2026-21852) are trust-boundary races in the same product cited as the Zero Trust exemplar. What keeps it partial: for behavioral monitoring & response and for AI governance policies this vault holds no non-Anthropic instantiation, so "the framework assumes the Anthropic stack here" cannot be separated from "the vault has not ingested the alternative." Settling it needs an external source on agent behavioral baselining and on multi-vendor agent-governance policy enforcement — a#oq/sourceshape.
Cited by 1
- Open Questions Backlog
Dashboard (Now items, domain counts, trend): Open Questions Dashboard.
Related articles
- Open Questions Backlog
Generated by `_system/lint.py --write-backlog`. Do not hand-edit. Domain and Watching sections carry one row per page —…
- Least Agency
OWASP term extending least privilege to agents: constrain not just what an agent can access but what each tool can do,…
- Agentic Prompt Injection
Direct and indirect injection of malicious instructions into an agent; LLMs cannot reliably distinguish information fro…
- Capability Gating Is Not Authorization
Agent frameworks ship capability gating (which tools are exposed, schema validity) but no fail-closed per-call authoriz…
- Out-of-Band Prompt-Injection Defense
Second-generation prompt-injection defense enforced outside the model: a deterministic reference monitor mediates tool…
