READING PATH
- MAIN ISSUE — A correctly scoped agent can still take an unacceptable path inside its scope.
- TABLE — Does "we scoped it right" hold up as a defense, and who answers for it when it doesn't?
- VSR-01 — Identify how much of an agent's granted access has never been exercised, and treat the unused portion as unrealized incident exposure.
- VSR-02 — Determine whether a given permission actually constrains the specific action being taken, or only constrained the agent's eligibility at some earlier point in time.
- VSR-03 — For a given failure category, determine whether access/permission control is even the right lever before spending remediation effort there.
- VSR-04 — Catch agents whose granted access has outlived the task, owner, or purpose that originally justified it.
- SOURCES — Inspect source backbone and claim-control notes.
PACKAGE MEDIA
Video briefing, slide deck, and field diagnostic for this DFEI package. The Signal Briefing video and deck are distillations of the issue and its VSRs.
Method
SURFACE STRUCTURE.
LABEL NOISE.
SIGNAL IMPLICATION.
DFEI issues originate within a human framework, evolve through machine-assisted research and reasoning, and pass through The Table — a structured human-machine roundtable where the signal undergoes scrutiny prior to publication.
ISSUE CONTENTS
- 01 SIGNAL
- 02 HIGHLIGHTS / FIELD SPOTLIGHTS
- 03 TERRAIN
- 04 VECTOR
- 05 OPERATOR IMPLICATIONS
- 06 ZEITGEIST
- 07 THE FAILURE THAT WASN'T A HALLUCINATION
- 08 SIGNAL GRID
- 09 TREND REPORT
- 10 FREE / USEFUL TOOLS
- 11 PAID TOOLS WORTH CONSIDERING
- 12 OVERHYPED / UNDER-TESTED FRAMES
- 13 WATCHLIST / UPCOMING DEVELOPMENTS
- 14 TABLE SECTION
- 15 VECTOR // SPECIAL REPORTS
- 16 SOURCE NOTES / CLAIM BOUNDARIES
- 17 APPENDIX / DOWNLOADS
- 18 CLOSING
Core position: A correctly scoped agent can still take an unacceptable path inside its scope.
01 SIGNAL
In the same week the EU AI Act's transparency obligations became enforceable (August 2, 2026), the identity-security industry arrived at Black Hat USA with a different answer to a question DFEI.009 spent an issue asking. SailPoint unveiled a unified "Agentic Fabric" for governing human and AI-agent identity together. Rubrik launched Agent Identity, which authorizes agent access one tool call at a time. Okta published a survey of more than 300 CISOs naming AI agents a board-level governance crisis. A coalition of security vendors launched the Open Secure AI Alliance, scoped around agent identity, permissions, isolation, and guardrails. None of the launches documented here establishes reliable trajectory-level judgment of whether an otherwise-authorized sequence of actions was acceptable. They govern what an agent is allowed to reach.
That distinction is not a technicality. DFEI.009 established that a correct outcome can be produced by an unacceptable trajectory: that scoring the endpoint of an agent's run tells you the goal was reached, not what was done to reach it. The industry's response, in the two weeks since, has been to build infrastructure that constrains the space an agent can act inside, on the reasonable premise that a narrower space produces fewer catastrophic paths. The evidence backs the premise: Teleport's 2026 Infrastructure Identity Survey found a 17 percent incident rate for least-privileged AI systems against 76 percent for over-privileged ones, a 4.5x gap, the kind of number that ends procurement arguments.
What none of this year's identity-security launches claim, and what none of the surveys behind them measure, is whether a correctly scoped agent still took an unacceptable path inside its scope. Permission scoping answers "what can this agent reach." It does not answer "was this specific run, through what it was allowed to reach, an acceptable one." Those are different questions, and the industry currently has fast, fundable, demoable answers to the first and none to the second.
This is not an argument against access control: the incident-rate gap is real and the practices behind it (governed identity, runtime-bound approval, lifecycle review) are good discipline, independent of anything else in this issue. It is an argument that access control is being sold, and increasingly bought, as if it were trajectory evaluation. It narrows where the danger can occur. It does not evaluate what happened where it was allowed to occur. A permission is a boundary. A trajectory is what happens inside the boundary. Governing the boundary is necessary. It is not the same project as judging the path.
02 HIGHLIGHTS / FIELD SPOTLIGHTS
Issue Highlights
-
Regulation arrived, but only its shallowest layer. The EU AI Act's transparency obligations became enforceable August 2, 2026. High-risk classification, the part that would actually constrain deployed agentic systems, was pushed to December 2027 (Annex III systems) and August 2028 (AI embedded in regulated products) under the AI Omnibus agreement.
-
The identity-security industry moved faster than the regulator, in the same window. SailPoint, Rubrik, Okta, and Cisco all shipped or announced agent-identity products around Black Hat USA 2026, days after the EU's enforcement date.
-
The data supports access control's incident-rate case, narrowly. Teleport's 2026 Infrastructure Identity Survey: least-privileged AI systems see a 17 percent incident rate versus 76 percent for over-privileged systems. That is a real, attributable, vendor-run but corroborated finding. It is a statement about general incident rate, not about trajectory-level harm specifically.
-
In one large dataset, the failure picture has moved past hallucination. ChatSee.ai's State of Enterprise AI Failures 2026, drawn from more than 10,000 observed failure events across 150+ categories, found hallucination responsible for under 10 percent of failures. The largest category, 31.1 percent, is resolution and escalation breakdowns: systems that appear to provide service without resolving anything. Execution and action failures rose 62 percent against a Q2 2024 baseline.
-
A federal review framework is under discussion, not yet policy. The White House has floated a plan for pre-release cybersecurity review of new AI models, capped at 30 days, likely exempting open-weight models. Reported by both the New York Times and Politico, both describing it as a proposal still being shaped.
-
The workforce conversation is getting more specific. Apollo Global Management's analysis found real wage decline, not just headcount loss, in roles with high AI exposure. The New York Fed's Liberty Street Economics blog is now tracking the same shift in hiring practices directly.
Field Spotlights
Black Hat week becomes agent-identity week
SailPoint's "Agentic Fabric" (paired with an updated Human Fabric identity platform), Rubrik's Agent Identity, and a fresh Okta CISO survey all landed within days of each other at Black Hat USA 2026. Cisco used the same window to extend its Duo Security identity platform and AI Defense shadow-AI inventory tooling. None of these is a new category. They are established identity-and-access-management vendors repositioning existing product lines around "agent identity" as the label of the moment. That repositioning is itself a signal: whoever owns the term first shapes what the term ends up meaning.
The Open Secure AI Alliance launches, scope not yet fully public
NVIDIA and 36 other technology, cloud, cybersecurity, and enterprise software organizations announced the Open Secure AI Alliance, covering the full stack around agents: identity, permissions, isolation, harnesses, guardrails, logs, and evaluation systems, rather than model security narrowly. Membership count confirmed via direct source check; the full member list and any published deliverables remain unconfirmed. Treated here as a watchlist item, not a settled fact.
Synopsys ships an autonomous, long-running agentic design workflow
At the 2026 DAC Chips to Systems Conference, Synopsys showcased new autonomous agentic AI workflows for chip design built with Microsoft and used by AMD, including a fully autonomous debug closure workflow (AgentEngineer™) with initial results showing up to a 40 percent reduction in cycle time. It is a useful field spotlight precisely because it is outside this issue's usual software/enterprise lane, a reminder that the permission-substitute problem shows up anywhere an agent is granted standing authority to act across a long-running task, not only in customer-facing enterprise deployments.
OpenAI ships Frontier, a cross-vendor agent management platform
OpenAI launched Frontier as an open platform for building, deploying, and managing agents across models from any vendor, with early adopters reported to include HP, Intuit, Oracle, State Farm, Thermo Fisher, and Uber. DFEI has not yet traced this adopter list to OpenAI's own primary announcement; treat the list as reported, not confirmed, until then.
03 TERRAIN
The terrain this issue sits on has two layers moving at different speeds, and the gap between their speeds is most of the story.
The slow layer is law. The EU AI Act's transparency obligations (marking and disclosing AI-generated content, disclosing certain AI-system interactions) became enforceable on August 2, 2026. That is real: it is the first time a major jurisdiction's AI-specific rules have actual legal force. But the AI Omnibus agreement that accompanied it pushed high-risk classification, the part of the Act that would require actual conformity assessment for consequential deployed systems, to December 2027 for most high-risk (Annex III) systems and August 2028 for AI embedded in already-regulated products. The enforceable layer, in other words, is the disclosure layer. The control layer is still years out.
The fast layer is the market, and it did not wait. In the same week the EU's disclosure rules took effect, four established identity-and-security vendors (SailPoint, Rubrik, Okta, Cisco) either launched or substantially extended agent-identity products, timed to Black Hat USA 2026. A fifth development, the Open Secure AI Alliance, positioned itself as an industry-wide standard-setting effort covering the same territory: identity, permissions, isolation, guardrails, logs, evaluation. Reading these together, the pattern is not five independent bets. It is a single, fast-moving consensus: the identity-security segment has decided the answer to "how do we make agentic systems safe enough to deploy" is access control, and it has decided this well ahead of any regulator, and largely without waiting to see whether access control actually closes the gap it is being sold to close.
The evidence for access control's value is real and specific. Teleport's 2026 Infrastructure Identity Survey, a named, traceable, vendor-run but corroborated study, found least-privileged AI systems carry a 17 percent incident rate against 76 percent for over-privileged systems, a 4.5x difference. That is the strongest single number in this issue's terrain, and it should be taken seriously: least-privileged access is associated with a materially lower incident rate. What the survey measures is general incident rate. It does not measure, and nothing in this terrain measures, whether the incidents avoided are the trajectory-level failures DFEI.009 was concerned with (unauthorized, irreversible, or unsafe paths) as opposed to simpler failures (an agent reaching a system it had no business touching at all, regardless of what it would have done once there). Reducing the incident rate by narrowing what an agent can reach is a genuine, welcome improvement. It is a different achievement than evaluating whether the paths taken inside the narrowed scope were acceptable.
The empirical failure data sharpens the point. ChatSee.ai's State of Enterprise AI Failures 2026, the most methodologically substantial source in this terrain, drawing on more than 10,000 observed failure events across 150+ categories, 10 industry verticals, and seven AI lifecycle stages, found hallucination accounts for under 10 percent of observed enterprise AI failures. The largest single category, at 31.1 percent, is resolution and escalation breakdowns: an agent that behaves correctly by every visible signal and simply never resolves the thing it was asked to do. Execution and action failures rose 62 percent against a Q2 2024 baseline. None of ChatSee's largest failure categories are access-control failures in the sense the identity-security industry is solving for. In this dataset, the dominant failure mode is not an agent reaching somewhere it shouldn't. It is an agent staying inside its lane and still not doing the job, or doing damage while technically staying in its lane. Access control does not touch that category at all. If this dataset's distribution generalizes, it implies a major share of operational failure sits outside what access control was ever built to reach; DFEI treats that as a plausible reading of one substantial dataset, not an established fact about the field.
04 VECTOR
The mechanism worth naming precisely is this: permission scoping is a design-time decision, made once, about what an agent is categorically allowed to reach. A trajectory is a runtime sequence of specific actions, taken in a specific order, in response to a specific and often unanticipated state of the world. Scoping constrains the first. Evaluation judges the second. The identity-security buildout this issue tracks is entirely, and understandably, focused on the first: it is the tractable half of the problem. You can build a product that enforces least privilege. You cannot yet buy a product that reliably judges, at runtime, whether a specific sequence of otherwise-authorized actions added up to an acceptable path.
There is a genuine terminology collision hiding inside this, and it is worth surfacing because it is exactly the kind of gap that lets a partial fix get sold as a complete one. ChatSee's failure taxonomy includes a category it calls "scoping failures," meaning task or goal scoping: an agent misunderstanding or drifting from what it was actually asked to do. The identity-security industry's "scoping" means something structurally different: least-privilege access scoping, what systems and data an agent is permitted to touch. Both are real failure classes. Neither substitutes for the other. When a vendor says "scoped," ask which sense. A perfectly access-scoped agent can still goal-drift its way through a bad trajectory using only privileges it was legitimately granted; a perfectly goal-scoped agent operating with excessive access can still do trajectory-level damage the access layer should have prevented. The two failure modes need two different controls, and treating either as a proxy for the other is the substitution this issue is named for.
Mapping ChatSee's failure categories against what identity/permission scoping plausibly addresses makes the gap concrete. This mapping is DFEI's own synthesis, not a finding from any single source, and should be read as interpretation:
| ChatSee failure category | Share of observed failures | Addressed by permission/identity scoping? |
|---|---|---|
| Resolution / escalation breakdowns | 31.1% (largest category) | No — a completion/quality failure, not an access failure |
| Execution / action failures | Up 62% vs. Q2 2024 | Partially — scoping limits blast radius of a wrong action; does not prevent the wrong action |
| Hallucination | <10% | No — output-quality failure |
| Retrieval failures | Reported, share not isolated in available excerpt | No — data/retrieval quality failure |
| Task/goal-scoping failures | Reported, share not isolated | No — this is the other meaning of "scope"; access control does not touch goal drift |
| Governance failures | Reported, share not isolated | Yes — this is the category the identity-security buildout is actually built for |
| Response failures | Reported, share not isolated | No — output-quality failure |
Read across the table, permission and identity infrastructure squarely addresses one category out of roughly seven, and partially addresses a second. It is real, valuable work on a real slice of the problem, and it is being marketed, through the sheer concentration of investment and press attention this issue's terrain documents, as if it were the whole answer. The industry has built a strong lock for one door in a house with seven doors, and is currently the loudest voice in the room about home security.
05 OPERATOR IMPLICATIONS
For an operator running or overseeing agentic systems, this issue's frame changes what "we've handled the safety problem" should mean.
First: treat access control as necessary infrastructure, not as evaluation. The practitioner guidance converging across this terrain is consistent and worth adopting on its own merits: inventory every agent as a governed identity with a named owner; bind approval gates to runtime execution rather than design-time trust (an approval given once at deployment is a different thing than an approval that holds at the moment of a specific action); apply lifecycle discipline (reviewing ownership, purpose, connected tools, and revocation) on a recurring cadence rather than once at launch. None of this is optional, and the Teleport incident-rate gap (17 percent versus 76 percent) is a strong enough number to justify doing it regardless of anything else in this issue.
Second, and separately: build or buy trajectory-level evaluation as its own line item, not as something access control will eventually cover once it matures. If ChatSee's dataset is representative of your own environment, the largest failure category it found (resolution and escalation breakdowns, at 31.1 percent) is one that a correctly scoped, correctly permissioned agent can still produce at full volume. An operator who has fully implemented least-privilege access and stops there has addressed the governance-failure slice of the problem and left the largest slice (completion quality, judged across the whole trajectory rather than at the endpoint) untouched.
Third: when a vendor's product description uses the word "scoped," ask which sense of scope is meant, out loud, in the procurement conversation. Access-scoped and goal-scoped are different properties, addressed by different tooling, and conflating them in a pitch deck is either an honest simplification or a sales tactic. An operator's job is to tell the difference before signing.
Fourth: know what your own audits actually attest to before you let anyone cite them as more. A passed access review is a real, specific claim: this agent's permissions match a documented, approved scope, the grant has an owner, observed access is consistent with it. It is not a claim that the agent's work is correct, and it was never designed to be. The gap between what an audit certifies and what an incident postmortem, a customer, or a regulator ends up assuming it certified is where accountability quietly goes missing, not because anyone lied, but because a signed document is easier to point to than an honest account of what wasn't checked. Say the boundary of your own attestations out loud before someone else defines it for you after something goes wrong.
06 ZEITGEIST
The mood underneath this issue's terrain is confidence arriving ahead of proof. The identity-security industry's Black Hat week had the texture of a category settling into place: vendors racing not to solve a new problem but to own the vocabulary of an already-agreed-on one, the way "zero trust" or "cloud security posture management" settled a decade earlier. That settling happens fast when the underlying number is good enough to end an argument, and 17 percent versus 76 percent is exactly that kind of number. It is easy, watching a room agree that quickly, to mistake consensus on the value of access control for consensus on the size of the problem it solves.
The labor conversation running in parallel has developed the same kind of specificity, moving past the abstract "AI will change jobs" framing toward measurable claims: real wage decline in high-exposure roles per Apollo, not just headcount attrition; the New York Fed's own research arm now tracking hiring-practice shifts directly rather than citing third-party studies. Both threads, security and labor, show the same pattern in this window: the conversation is getting more granular and more evidence-backed, and the granularity is exposing gaps that the broader, vaguer version of the conversation used to paper over. That is, on balance, progress (a more precise problem statement is a more solvable one), even when, as in this issue's central case, the precision reveals how much of the current response is aimed at the tractable half of the problem rather than all of it.
07 THE FAILURE THAT WASN'T A HALLUCINATION
The framing habit worth retiring this issue is the assumption that "AI risk" means bad output (a hallucinated fact, an off-brand response, a factual error a human should have caught). TERRAIN already gave the numbers that retire it, in ChatSee's dataset hallucination is under 10 percent of failures, and the largest category by a wide margin, at 31.1 percent, is a system that says the right things, follows the right output rules, and simply never resolves what it was asked to resolve.
That category, not hallucination, is where this issue's central subject runs out of reach. Access-control and identity infrastructure is well matched to a subset of execution and governance failures: an agent reaching a system, credential, or dataset it should not have. It is not matched at all to resolution and escalation breakdowns. A perfectly scoped agent, touching only systems it was explicitly authorized to touch, can still fail to resolve a customer's issue, silently, at the largest observed rate in the entire failure taxonomy, and no identity platform will show up in that failure's postmortem, because nothing about the failure involved unauthorized access.
The honest reading is not that access control is misdirected effort. It is that "we've secured our agents" and "our agents work" have become, in the current wave of enterprise deployment, two different claims that are being marketed and often budgeted for as if they were one. An organization can be simultaneously well-governed by access-control standards and failing at the primary thing agents are deployed to do. Naming that gap plainly is this issue's contribution, and it is the gap the four Vector Special Reports that follow are built to work through, one layer at a time.
08 SIGNAL GRID
Tier 1, Active / Confirmed
- EU AI Act transparency obligations are enforceable as of August 2, 2026; high-risk classification rules are not (pushed to Dec 2027 / Aug 2028).
- Major identity-security vendors (SailPoint, Rubrik, Okta, Cisco) have shipped or substantially extended agent-identity products, concentrated around Black Hat USA 2026.
- Least-privileged AI access is associated with a materially lower incident rate than over-privileged access (17% vs. 76%, per Teleport's 2026 Infrastructure Identity Survey).
- The dominant enterprise AI failure category is resolution/escalation breakdown, not hallucination (ChatSee.ai, State of Enterprise AI Failures 2026).
- A federal pre-release AI model review framework is under active discussion in Washington, not yet adopted policy.
Tier 2, Emerging
- A multi-vendor Open Secure AI Alliance (NVIDIA + 36 other organizations) has launched around agent identity/permissions/isolation/guardrails/evaluation; full member list and deliverables not yet independently confirmed.
- Cross-vendor agent management platforms (OpenAI Frontier, Google ADK 2026) are maturing as a distinct infrastructure layer beneath identity/permission tooling.
- Workforce-impact research is shifting from job-loss framing toward wage-suppression-in-high-exposure-roles framing.
- Agentic workflows are expanding into non-software verticals (chip design, per Synopsys at DAC).
Tier 3, Watchlist
- Whether the Open Secure AI Alliance publishes an actual standard versus remaining a press-release coalition.
- Whether the White House pre-release review framework advances past discussion stage, and how the open-model exemption is finally scoped.
- EU high-risk classification final guidance, still pending (carried from DFEI.009).
- Whether any vendor or researcher builds a trajectory-level (not general-incident-rate) evaluation of access-scoping's effect.
09 TREND REPORT
The trend worth naming in this coverage window is repositioning, not invention. Every major identity-security move in this issue's terrain (SailPoint's Agentic Fabric, Rubrik's Agent Identity, Okta's CISO survey and product posture, Cisco's Duo/AI Defense extension) is largely an extension of an existing IAM, PAM, or security-posture capability into agent identity, rather than a wholly new control architecture built from scratch for the agent problem. None of it is a new architecture built from scratch for the agent problem; all of it is "we already do identity and access management, and now agents are identities too." That is not a criticism: extending mature identity infrastructure to a new class of actor is exactly the right instinct, and the incident-rate data backs it. It is a description of where the investment is actually landing: on extending a solved problem (identity and access management) to a new population, rather than on solving the comparatively unsolved problem (runtime trajectory evaluation) this issue has been tracking.
The second-order trend is timing. Every major move documented here happened within roughly two weeks, clustered around Black Hat USA 2026 and immediately following the EU AI Act's first enforcement date. That is not coincidence: it is a category moving to establish itself at the exact moment attention is highest, before either the regulatory floor (EU high-risk rules, still two years out) or the technical alternative (trajectory-level evaluation, still nascent per DFEI.009's terrain) is ready to compete with it for the same budget line.
10 FREE / USEFUL TOOLS
See the boundary. See the path.
OpenFGA — Fine-grained authorization. Open-source, CNCF-hosted authorization system built to answer granular questions: may this actor perform this action on this object. It formalizes and enforces the boundary about as cleanly as a free tool can. What it does not answer is whether the authorized action was the right one in the circumstances. OpenFGA can answer "may this actor do this?" It cannot answer "should this particular trajectory have done it?" Top free pick.
Open Policy Agent — Policy-as-code, runtime decisions. A general-purpose, open-source policy engine that separates policy decisionmaking from application logic and can incorporate contextual, external data into a decision. That puts it usefully between simple static permission and genuine runtime binding. The boundary still holds: a more context-aware permission check is still a policy decision. It doesn't automatically become a judgment about whether the resulting multi-step action sequence was correct. Top free pick.
Langfuse — Agent tracing and evaluation. Open-source and self-hostable, tracing LLM applications and evaluating traces, sessions, and datasets, built to detect regressions and inspect behavior, not just assign access. This is the counterweight to OpenFGA and OPA: they show the boundary, Langfuse shows the path. Permission tells you what was allowed. Tracing tells you what actually happened. Evaluation asks whether what happened was acceptable. Top free pick.
AgentOps — Agent execution tracing and replay. A free tier covering up to 5,000 events, with monitoring, replay analytics, cost tracking, and framework integrations. More operational than theoretical: if an agent completed the task, what did it actually do between start and finish. Directly supports this issue's trajectory argument. Useful free tool, practical operator pick.
Google ADK 2026 — Substrate, not a control. The layer agents get built on, not an access-control or evaluation tool. Worth naming only so the stack is visible end to end: ADK underneath, OpenFGA and OPA constraining what's built on it, Langfuse and AgentOps observing and evaluating it once it runs. Reference resource, not a headline pick.
No single tool on this list closes the whole gap, and that's the point. Authorization and trajectory evaluation stay different layers no matter how many of these an operator runs together.
11 PAID TOOLS WORTH CONSIDERING
Know which problem you're actually paying to solve.
The paid tier forces the same procurement question this issue keeps asking. Are you buying identity governance, runtime authorization, observability, or trajectory evaluation? Those are not interchangeable products, and a vendor who blurs the line between them is making exactly the substitution this issue is about.
Permit.io — Fine-grained, agent-aware authorization. Positions itself around authorization across apps, APIs, agents, and data (RBAC, ABAC, ReBAC, policy-as-code, approval workflows, decision traces), with enforcement reaching down to the individual agent/tool interaction. Entry pricing starts around $25/month. The vendor itself acknowledges that traditional authorization alone is insufficient for agents, emphasizing delegation, purpose, and goal-scoped permissions, more self-aware framing than most of this issue's terrain offers. Worth examining because it pushes permission closer to runtime, narrowing this issue's gap. It still doesn't erase the distinction between authorizing an action and judging the completed trajectory. Strong paid candidate.
SailPoint Agentic Fabric — Enterprise identity governance. Central to this issue's own terrain; its Black Hat USA 2026 positioning is the clearest evidence of an established IAM vendor extending into agent identity. Consider it if the problem is governing agent identity and entitlement. Do not buy it expecting trajectory evaluation. It was never built to deliver that. Enterprise / identity-governance tier.
Rubrik Agent Identity — Runtime agent access governance. Belongs here because VSR-02 is built directly around the failure mode it targets: approval that isn't runtime-bound. The open editorial question is whether the product materially binds approval to the specific runtime action or only narrows standing access at a coarser grain, which is this issue's exact operator test, not a rhetorical one. Worth evaluating, not yet a blanket endorsement.
Galileo — Agent evaluation and runtime protection. The category this list would otherwise be missing. Positions its platform around evaluating multi-step agent behavior, tool selection, action completion, and flow, not just endpoint output, a useful foil to the identity products. Worth remembering that Galileo is itself a vendor making claims about its own evaluation capability. Interesting because it attempts to evaluate the path, not just the permission. Not evidence that automated trajectory judgment is a solved problem. Professional evaluation tier, test before trusting.
Langfuse Cloud — Managed evaluation and observability. The paid layer on top of the free tool above. Pay when managed infrastructure, collaboration, retention, and production-scale evaluation save more operational burden than self-hosting costs. The underlying conceptual fit, tracking and evaluating traces, is already excellent for free. Worth paying for at production scale.
Buying all five would not prove that every in-scope action was correct. It would give an organization real, layered coverage across identity, permission, tracing, and evaluation, and coverage is not completion. That's the same distinction the Attestation Coverage Matrix builds into its scoring: even a perfect run through this entire list caps out at Scoped and Governed, not Assured.
12 OVERHYPED / UNDER-TESTED FRAMES
Overhyped: Individual frontier model releases (Gemini 3.5 Pro, Meta's Muse Spark 1.2, OpenAI's Astra, DeepSeek-V4-Flash) drew the bulk of mainstream coverage in this window. None of them changes this issue's argument. The structural story is one layer down, in the infrastructure being built around agents rather than in the agents' raw capability.
Under-tested: The causal link this issue's whole argument depends on has not actually been tested by anyone in this terrain. Teleport's survey shows a correlation between privilege level and general incident rate. No source here isolates whether narrowing access specifically reduces trajectory-level harm (unauthorized, irreversible, or unsafe paths, in DFEI.009's vocabulary) as opposed to simpler categorical incidents (an agent reaching a system it had no reason to touch, independent of what it would have done there). That gap is not a citation this issue is missing. It is a study nobody appears to have run yet.
13 WATCHLIST / UPCOMING DEVELOPMENTS
- Open Secure AI Alliance's actual output. Watch whether it produces a published standard or stays a coalition press release. This issue treats its scope as unconfirmed.
- White House pre-release review framework. Watch whether the 30-day cap and open-model exemption survive into an actual policy, and who ends up doing the review.
- EU AI Act high-risk classification, final guidance. Carried from DFEI.009, still pending, now formally delayed to Dec 2027 / Aug 2028 under the AI Omnibus.
- Trajectory-specific evaluation of access control. Watch for the first study that isolates whether permission scoping reduces trajectory-level harm specifically, rather than general incident rate: the gap this issue identifies as unfilled.
- Identity-vendor consolidation. Watch whether SailPoint/Okta/Rubrik/Cisco's agent-identity repositioning produces genuine new capability over the next two issues, or settles as a rebrand.
- Runtime binding as a compliance requirement. THE SUBSTITUTION TEST (this issue's Table) forecast that disclosure and transparency rules will likely pull runtime-bound approval checks forward as a compliance expectation before market incentive does it voluntarily. Watch whether that shows up in any vendor's product roadmap or any regulator's guidance within the next two issues.
- Automated judgment of action-correctness. The Table's harder finding: nothing on the horizon, market or regulatory, is close to mandating or reliably delivering automated evaluation of whether a specific in-scope action was actually correct. Watch for the first credible attempt, research or product, and treat anything claiming to have already solved it with the skepticism this issue applies to "scoped means safe."
14 TABLE SECTION
THE TABLE // THE SUBSTITUTION TEST
The DFEI.010 Table took one question: if an agent stayed entirely within access it was correctly and legitimately granted, and the resulting trajectory was still harmful or unresolved, did the access-control layer succeed, and who is accountable when "properly scoped" and "actually safe" turn out not to be the same claim?
The Table is a synthetic multi-persona roundtable, a structured reasoning artifact, not an empirical study. The personas are analytical constructs designed to surface and pressure-test competing positions on a live question. No persona represents an individual, organization, or institutional view. The Accelerationist and the Hype Agent voice full-conviction positions deliberately, to expose the assumptions the question must price, not because DFEI holds those positions.
The question the Table was designed to hold
This issue's terrain arrived with a genuinely good number behind it: least-privileged agents see a 17 percent incident rate against 76 percent for over-privileged ones. Nobody at the Table disputed that finding. What the Table pressure-tested was everything people are quietly allowed to assume that number also proves, that scoped means safe, that a signed access audit is a defense against any failure, that the fundable half of a problem is the whole problem because it's the half anyone can measure.
What the Table produced
A guest persona new to this issue's roster, The Compliance Auditor, drew the sharpest line: a passed access review attests that permissions match an approved scope, nothing more. It does not attest that the agent's work was correct, and it never claimed to. What turns that narrow attestation into a broader, unearned claim of safety isn't fraud, it's habit: a signed document is easier for an organization to point to than an honest account of everything that wasn't checked. The Ethics Examiner traced where that habit puts the cost, on whoever relied on the broader claim without the standing to question it, usually the customer or worker on the other end of the agent's action, never the party whose signature closed the transaction.
Two guest personas seated live via redirect, The Infrastructure Engineer and The Regulator, split the remaining gap in a way nothing else in this issue's terrain had: runtime-bound approval checking is a build decision organizations haven't prioritized, not an unsolved problem, and compliance pressure will likely reach it before market incentive does. Automated judgment of whether a specific in-scope action was actually correct is a genuine research gap, not currently buildable at will, and nothing on the regulatory horizon is close to mandating it.
The closing position
Scoped is not safe. It's one input into safe. Treat it as anything more, and the next incident is the one that proves it.
"We scoped it right" survived the Table as a real, specific defense, against unauthorized access. It did not survive as a defense against the failure category the data actually shows is largest. The Table's operator closed by putting the weight on the gap between what gets certified and what gets assumed, not on any single villain in the chain.
Field artifact
The Table's accountability framing and the four VSRs' operator tools are combined in The Attestation Coverage Matrix, a six-dimension diagnostic whose sixth dimension, whether a specific in-scope action was actually correct, is deliberately never scored Pass. It stays visible on every scorecard as a named, currently unassessable gap, so the other five dimensions can never be mistaken for the whole picture. Available in the DFEI.010 resources package.
15 VECTOR // SPECIAL REPORTS
The permission substitute has four layers where the same gap plays out, each one region of a single permission's life. They can be read independently or in sequence; read in order, they chain, from whether access was granted correctly, to whether it's enforced when it matters, to whether access was ever the right lever for the failure in question, to whether any of it stays true over time.
VSR-01 — THE SCOPE THAT WAS TOO WIDE Finding the access an agent never needed before it becomes the access that hurt you
Layer: Grant level — the access-provisioning decision made before any trajectory runs Applied tool: Grant-vs-Use Gap Audit Field rule: If a granted capability hasn't been exercised in the review window, its risk is already priced in and its benefit is not.
Provisioning defaults wide because narrowing later is friction. The Grant-vs-Use Gap Audit compares what an agent was granted against what it actually uses, turning unused scope from invisible headroom into a line item that has to justify itself.
VSR-02 — THE APPROVAL THAT WASN'T RUNTIME-BOUND A grant made once at deployment is not the same control as an approval checked at the moment of action
Layer: Runtime level — the gap between a design-time approval and the runtime action it's presumed to cover Applied tool: Runtime Binding Test Field rule: An approval that isn't re-checked at the moment of action isn't a control at the moment of action, it's a memory of one.
A narrow grant checked once at deployment can still authorize an unacceptable specific action inside that scope, if nothing checks the action itself. The Runtime Binding Test classifies any approval gate as design-time trust or genuinely runtime-bound.
VSR-03 — WHAT SCOPING DOESN'T TOUCH Before tightening access, confirm the failure you're fixing is an access failure
Layer: Taxonomy level — which failure categories access control can and cannot reach Applied tool: Control Coverage Map Field rule: Before tightening access, confirm the failure you're fixing is an access failure.
The largest category of enterprise AI failure has no access dimension at all. The Control Coverage Map sorts an operator's own failure log by whether access control is even the right lever, before remediation budget goes to a fix that structurally cannot help.
VSR-04 — THE IDENTITY THAT OUTLIVED ITS PURPOSE A permission is only as current as its last review
Layer: Lifecycle level — ongoing identity/ownership discipline across an agent's operational life Applied tool: Agent Lifecycle Register Field rule: A permission is only as current as its last review. Treat an unreviewed grant as an expired one.
An agent correctly scoped at creation can drift out of alignment with its own purpose as owners, tasks, and tool stacks change, with nothing structurally forcing a re-check. The Agent Lifecycle Register applies joiner-mover-leaver discipline to agent identities specifically.
16 SOURCE NOTES / CLAIM BOUNDARIES
This issue draws on a mix of institutional, vendor-run, and vendor-adjacent sources; posture is disclosed here rather than left implicit.
Directly citable, institutional: the European Commission's own AI Act enforcement status; Goodwin Law's transparency-obligations alert; the New York Fed's Liberty Street Economics coverage of AI's labor-market effects.
Reported claims, traced to a named primary source: ChatSee.ai's State of Enterprise AI Failures 2026 (10,000+ observed failures, 150+ categories); Teleport's 2026 Infrastructure Identity Survey (the 17%/76% incident-rate figures).
Vendor claims, attributed as such: SailPoint's Agentic Fabric, Rubrik's Agent Identity, OpenAI's Frontier platform and its reported adopter list, all company announcements, not independent findings.
Vendor-adjacent surveys, sponsor disclosed: Okta's CISO survey, SAP LeanIX's adoption figures, Upwork's Future Workforce Index, all run by parties with a commercial interest in the framing.
Explicitly not used as stated: a 44%/13% policy-adoption figure (conflicting across sources, not resolved to a single primary number); Fiddler AI's "88% fail in production" headline (the firm's own synthesis, not a named study, though the studies underneath it, WebArena, a Carnegie Mellon evaluation, an MIT pilot-outcomes report, and a Princeton reliability study, are independently citable); an unnamed Futurism claim about older workers being displaced first.
Known gap: no source in this issue's terrain isolates whether access scoping reduces trajectory-level harm specifically, as distinct from general incident rate. This issue's central argument treats that as a real, unfilled gap rather than papering over it with an imperfect citation.
Full claim-layer breakdown is available in the companion Source Notes; source-by-source reading path: Source Backbone
17 APPENDIX / DOWNLOADS
- Main issue — this document
- VSR-01 — The Scope That Was Too Wide (Grant-vs-Use Gap Audit)
- VSR-02 — The Approval That Wasn't Runtime-Bound (Runtime Binding Test)
- VSR-03 — What Scoping Doesn't Touch (Control Coverage Map)
- VSR-04 — The Identity That Outlived Its Purpose (Agent Lifecycle Register)
- THE TABLE — THE SUBSTITUTION TEST (full transcript)
- Field Artifact — The Attestation Coverage Matrix
- Source Notes — claim layers, what this issue claims and doesn't
- Source Backbone — full public source list
18 CLOSING
The access-control industry did real work in this window, and this issue has tried not to undersell it: a 4.5x incident-rate gap between narrow and broad access is not a marketing artifact, it is exactly the kind of measurable win that should get built, funded, and shipped without waiting for the harder half of the problem to catch up. The mistake is not building it. The mistake is letting "we built it" answer a question it was never positioned to answer.
The Table's closing line is the plainest version of this issue's argument available: scoped is not safe, it's one input into safe. The permission substitute isn't a conspiracy or a scam. It's what happens when one half of a problem is measurable, fundable, and demoable, and the other half isn't yet, and an organization lets the first half's confidence stand in for the second half's absence. Four VSRs in this issue give an operator a way to close the gaps that are actually closable right now, grant, runtime, coverage, lifecycle, and are explicit about the one gap that isn't: whether a specific action, taken inside a perfectly authorized scope, was the right one to take. Nobody at this issue's Table solved that, and this issue's research didn't turn up a reliable, general-purpose way to do so. The honest position is to say so, keep building the half that's buildable, and treat any claim that the whole problem is already handled as the exact substitution this issue is named for.