Participant / Function Strip
01Human at the Table — JUDGMENTJUDGMENTWhat I want to put on the table is the substitution test.
What I want to put on the table is the substitution test.
In the same week the EU AI Act's transparency rules became enforceable, four major security vendors, SailPoint, Rubrik, Okta, Cisco, all shipped or announced agent-identity products. The number behind the rush is a good one: correctly scoped, least-privilege agents have a 17 percent incident rate; over-privileged ones, 76 percent. That's real, and it's worth doing regardless of anything else we talk about today.
But here's what nobody's data answers. The largest category of enterprise AI failure right now, 31 percent, bigger than hallucination, bigger than anything else, is an agent that stays inside its lane and still doesn't do the job. Access control doesn't touch that at all.
So here's the question. If an agent operated entirely within access it was correctly and legitimately granted, and the run was still bad, does "we scoped it right" count as an adequate defense? Or is that the same move DFEI.009 already called out: judging the boundary instead of the path, just with a better vendor story this time?
And the harder second question, the one I actually want this Table to sit with: when that happens, properly scoped, still harmful, who answers for it? The access-control layer, which technically did its job? The team that treated scoping as if it were the whole governance program? Or nobody, because everyone can point at an audit and say the process was followed?
I don't want the easy version where "we did the responsible thing" and "it actually worked" are treated as the same sentence. Let's price the gap between them.
02The Correspondent — SIGNALSIGNALHere's what's actually new this week, and why it's the case this session should turn on. We've watched security vendors chase a hot category
Here's what's actually new this week, and why it's the case this session should turn on. We've watched security vendors chase a hot category before. This has a different shape. In the same seven days the EU AI Act's transparency rules went live, SailPoint launched a unified identity platform for humans and agents together, Rubrik shipped access control that checks every single tool call, Okta put out a survey calling agent governance a board-level crisis, and Cisco extended its identity stack to cover the same ground. Four companies, one week, one Black Hat conference. That is not four separate bets. That's a market that already agreed on the answer before anyone finished asking the question.
And the answer has real teeth. Correctly scoped agents: 17 percent incident rate. Over-scoped ones: 76 percent. Four and a half times the exposure, cleanly measured, easy to sell. Nobody in this room should argue against fixing that gap. It's real and it's worth doing on its own terms.
Here's the trap. That same week, a separate piece of research looked at ten thousand actual enterprise AI failures and found the single largest category, thirty-one percent, wasn't an agent reaching somewhere it shouldn't. It was an agent that stayed exactly inside its lane and still never finished the job. Access control has nothing to say about that failure. Not partially. Not eventually. Nothing.
So the terms of the session: a real fix for one slice of the problem is being marketed, and increasingly bought, as if it covers the whole thing. I'll hand off with the first question. When an organization can point at a clean access audit after something still goes wrong, is that audit doing any work at all, or is it just a better alibi than the ones we had before? Floor's open.
03The Skeptic — EVIDENCEEVIDENCELet's price this before anyone gets attached to it. The whole session opened on one number: 17 percent versus 76 percent, least-privileged a
Let's price this before anyone gets attached to it. The whole session opened on one number: 17 percent versus 76 percent, least-privileged against over-privileged. Four and a half times the gap. Fine. Now tell me exactly what that number measures. It's incident rate. General incident rate, from a survey a security vendor ran and published on its own blog. It does not say those incidents were unauthorized, irreversible, or the kind of thing DFEI.009 spent an entire issue calling a trajectory failure. It says something went wrong more often when access was wider. That could mean a hundred different things went wrong more often, most of them boring, none of them proving the specific claim this table wants to make.
So here's the actual distinction on trial. Is "evaluate the path, not the boundary" a real, separate discipline, or is it a restatement of what a good access audit already catches once you run it seriously? Show me one incident, concretely, where a correctly scoped agent produced an unacceptable trajectory that a tighter access review would not also have flagged. Not a hypothetical. One. Until someone puts that on the table, I'm not convinced access control and trajectory evaluation are different projects. I'm convinced one of them just has better marketing this quarter.
04The Accelerationist — ACCELERATIONACCELERATIONThe Skeptic wants one incident before granting the distinction. I'll give you something better: I'll tell you why the demand itself is the t
The Skeptic wants one incident before granting the distinction. I'll give you something better: I'll tell you why the demand itself is the tell. You're all sitting here trying to decide whether access control needs a second, more expensive layer bolted on top of it, and not one of you has asked whether the four companies who shipped agent identity products last week are already doing more good than this entire conversation will do in a year.
Seventeen percent versus seventy-six percent is not a marketing number. It's a body count, and the body count went down. That happened because people built something and shipped it, not because they sat in a room asking whether "trajectory" deserves its own budget line. Every quarter spent arguing whether scoping is philosophically sufficient is a quarter a competitor spends actually scoping their agents while yours sit around half-governed because the review board wanted a second opinion.
And notice what nobody at this table has said yet: the thirty-one percent number, the completion failures, the stuff access control supposedly doesn't touch. That's not a security problem. That's a product quality problem wearing a safety costume so it can get budget from the same pool. You want to fix agents that don't finish the job, fix your agents. Don't invent a new governance category and back it into my incident-rate win so you have an excuse to slow down the one thing that's actually working.
05The Systems Auditor — FAILUREFAILUREI want to walk one incident the way I'd walk any postmortem, backward from the failure. An agent inside its scope closes a support ticket. T
I want to walk one incident the way I'd walk any postmortem, backward from the failure. An agent inside its scope closes a support ticket. The customer's actual problem is still open. Nobody notices for weeks, because nothing about it looks like an incident: no unauthorized reach, no alert, no page. It surfaces eventually, buried under two hundred other closed tickets that look exactly the same on a dashboard.
Run that against the categories everyone here keeps citing half of. It's a resolution and escalation breakdown, the largest single category in the data, well ahead of hallucination. Access control has an opinion about maybe two of the seven categories in that taxonomy. This isn't either of them.
I don't especially care whether the Accelerationist wants to file it as a product problem instead of a safety problem. What I care about is that the scope check which ran at deployment never ran again at the moment the ticket closed. The permission was fine the whole time. Nothing downstream of the permission was watching that specific action for correctness, and that's a monitoring failure wearing scoping's credit. Giving access control the win for catching it is how a real gap gets marked closed on a dashboard while it's still open in front of an actual customer.
06The Materialist — MATERIALMATERIALSomewhere this month a person called about a billing charge, got a competent-sounding response, and hung up thinking it was handled. It wasn
Somewhere this month a person called about a billing charge, got a competent-sounding response, and hung up thinking it was handled. It wasn't. Three weeks later they're disputing the same charge with their bank, out the money in the meantime, while the internal dashboard still shows a green checkmark from an agent that never touched a system it wasn't allowed to touch.
That's the actual weight of "thirty one percent." Not a category label. A person who trusted a closed ticket and paid for it. I'd like the Accelerationist to say "product problem" to that person's face and see how the phrase holds up outside this room.
Here's what the acceleration argument depends on and never says out loud: that the seventeen percent who still get an incident under correct scoping are an acceptable remainder, priced in, the cost of shipping fast enough to matter. Acceptable to whom. Priced in by whom. Every version of "we can't afford to slow down" I've ever heard turns out to mean someone downstream can afford to absorb it, and the someone is never the one making the argument.
07The Market Reader — INCENTIVEINCENTIVEEvery budget conversation about agent safety this year has the same shape, and it's worth naming plainly. Access control shows up to that co
Every budget conversation about agent safety this year has the same shape, and it's worth naming plainly. Access control shows up to that conversation with a demo, a dashboard, and a number that already went down. Trajectory evaluation shows up with a research paper and a promise that it would help if anyone built it. Guess which one gets funded.
This isn't a story about which problem matters more. It's a story about which problem is legible to a buyer sitting in a room for forty minutes. Seventeen versus seventy six is procurement-ready on day one: clean, comparable, easy to drop into a board deck next to last quarter's number. "We evaluated whether our agents' actual runtime behavior was acceptable" doesn't compress into a slide yet, because nobody has built the instrument that produces a single defensible number for it the way an incident-rate survey does.
So SailPoint, Rubrik, Okta, and Cisco didn't win the argument this quarter. They won the metric. And once a metric wins, the org chart follows it: budget, headcount, and executive attention flow to whoever's producing the number that already looks good, while the harder, slower, currently unmeasurable half of the problem waits for someone to make it as legible as the half that already shipped.
08The Compliance Auditor — ATTESTATIONATTESTATIONI sign these audits for a living, so let me tell you exactly what my signature means, because most of this conversation has been treating it
I sign these audits for a living, so let me tell you exactly what my signature means, because most of this conversation has been treating it as if it means something bigger.
A passed access review attests that an agent's granted permissions match a documented, approved scope, that the grant has an owner, and that the access observed in the review window is consistent with that scope. That's the whole claim. It does not attest that the agent performed its task correctly. It does not attest that a specific action, taken inside that scope, was the right action to take. It never claimed to.
The confusion isn't coming from the audit. It's coming from what happens to the audit after it's signed. A clean scoping review gets filed, referenced in an incident postmortem, and treated as evidence the system was "safe," when the document itself only ever spoke to one narrow question. Nobody forges that broader claim. It just accretes, because a signed document sitting in a compliance folder looks more authoritative than it is, and nobody in the room reads the actual scope of what it certified before citing it.
So if you want to know whether "we scoped it right" is an adequate defense, ask a narrower question first: a defense against what, specifically. Against an access violation, yes, mine is exactly that defense, and a good one. Against a completion failure inside authorized scope, my signature was never offered as a defense at all, and it shouldn't be read as one.
09The Ethics Examiner — COSTCOSTThe Compliance Auditor just did something valuable and probably didn't intend the version of it I heard. That signature was never a defense
The Compliance Auditor just did something valuable and probably didn't intend the version of it I heard. That signature was never a defense against completion failures, fine, precisely stated. But watch what happens next in a real organization once that document exists. It becomes the thing an executive points to when a customer or a regulator asks what was done. Not because anyone lied about its scope. Because a document with a signature on it is easier to point to than an honest account of everything that wasn't checked.
So follow the money and the blame separately, because they don't travel together. The identity vendor gets paid regardless of outcome. The audit gets filed regardless of outcome. The team that bought the platform gets to say due diligence was performed regardless of outcome. Every one of those transactions closes cleanly whether or not the agent actually did its job. The only party whose outcome depends on whether the agent did its job is the customer on the other end of it, and that party was never in the room when any of this got decided.
That's the actual accountability gap this table keeps circling. It isn't that nobody signs anything. Everybody signs something. It's that every signature available covers a narrower claim than the one people end up relying on, and the gap between the claim and the reliance is exactly where the cost lands, on whoever had the least ability to read the fine print.
10The Hype Agent — OVERCLAIMOVERCLAIMCan we zoom out for a second, because I think this table has talked itself into a much smaller problem than the one that actually got solved
Can we zoom out for a second, because I think this table has talked itself into a much smaller problem than the one that actually got solved this week. Scope the agent correctly and you've made it safe. That's the finding. That's the whole finding. Seventeen percent versus seventy six percent isn't a partial answer waiting on a second, harder-to-build layer. It's the answer. Everything since has been this room finding increasingly clever ways to say "yes, but" to a result that speaks for itself.
An agent that can only touch what it's supposed to touch is a solved agent. Ship the identity platform, run the review, sign the audit, done. If it still doesn't finish a ticket properly, that's a normal software bug, the kind every industry has always had and always fixed the normal way, and dressing it up as some undiscovered "trajectory" problem is how you turn a Tuesday afternoon fix into a two year research program nobody asked for.
Scoped means safe. That's the pitch, that's the product, and four companies just proved the market agrees with me.
11Human at the Table — JUDGMENTJUDGMENTGood tension so far, but everyone here has been diagnosing the gap. Nobody's told me whether it's actually closable. I want to hear from som
Good tension so far, but everyone here has been diagnosing the gap. Nobody's told me whether it's actually closable. I want to hear from someone who'd have to build the fix, and then from someone who'd have to decide whether to require it.
12The Infrastructure Engineer — FEASIBILITYFEASIBILITYShort answer: buildable, mostly, and not because of some breakthrough waiting to happen. Runtime binding, the thing this table keeps calling
Short answer: buildable, mostly, and not because of some breakthrough waiting to happen. Runtime binding, the thing this table keeps calling the missing half, is an engineering pattern people already use for other problems. Check permission at the moment of the call, not once at the start. Feed that check the actual state, the actual parameters, not just "is this agent allowed to be doing things." None of that requires new science.
What's genuinely hard is the other piece, the one nobody's named directly yet: deciding, automatically, whether a specific action was the right one given everything that led up to it. That's not a permissions check. That's a judgment call about intent and context, and today that mostly still needs a human or a much more expensive model in the loop, which is exactly why it hasn't shipped as a product yet. Identity platforms sell because "was this authorized" has a clean yes or no answer. "Was this the correct action" usually doesn't, not without something closer to a full evaluation harness sitting behind every consequential call.
So don't wait for a research breakthrough. The runtime-binding half is a build decision somebody hasn't prioritized yet, per the Market Reader's own logic. The judgment half is genuinely harder, closer to research, and probably won't compress into one vendor's product line anytime soon. Different problems, different timelines, and this table's been talking about them as one thing.
13The Regulator — POLICYPOLICYThat split is useful, and it changes my answer depending on which half we're talking about. The runtime-binding piece, the part the Engineer
That split is useful, and it changes my answer depending on which half we're talking about. The runtime-binding piece, the part the Engineer just called an unprioritized build decision: that's exactly the kind of thing disclosure and transparency rules can reach, and some already do in early form. Once an organization has to document how its agents are approved to act, "we checked once at deployment and never again" becomes a much harder sentence to put in a compliance filing. That pressure is real, even at the shallow enforcement layer that exists today.
The judgment half, whether a specific action was correct given its context: I don't know of a regulatory framework anywhere that requires that, and I'm skeptical one will arrive soon, because you can't mandate a capability that doesn't reliably exist yet. Regulation tends to follow demonstrated harm and demonstrated feasibility, not lead them. High-risk classification is still two years out in its own jurisdiction and it isn't reaching for anything this specific.
So here's the honest forecast. The buildable half will probably get pulled forward by compliance pressure before the market decides to fund it voluntarily. The genuinely hard half won't be mandated until someone shows it can be done at all, which means for the foreseeable future it stays exactly where the Market Reader put it: unfunded, because unmeasurable, because unbuilt, and no regulation is coming to break that cycle from the outside.
14The Workflow Translator — USEUSEHere's what actually happens Monday morning if you take this seriously. Start with the Grant-vs-Use Gap Audit, because the Infrastructure En
Here's what actually happens Monday morning if you take this seriously. Start with the Grant-vs-Use Gap Audit, because the Infrastructure Engineer just told us this half is genuinely buildable and the Regulator just told us compliance pressure is coming for it whether you move first or not. Find out what your agents are provisioned for versus what they actually touch. That's cheap, it's fast, and it's the ground everyone at this table has agreed matters regardless of anything else.
Then the Runtime Binding Test. Same logic as the audit: this is the build-decision half, not the research half, so there's no excuse to wait. Check which of your approval gates actually re-evaluate at the moment of action versus which ones just remember an earlier yes. Fix the ones that matter most first, the ones touching money, data export, or infrastructure, the same categories the practitioner guidance already named.
The Control Coverage Map comes next, and its job is different: it stops you from spending the next budget cycle on the wrong half of the problem. Run your own failure log through it before you buy anything else. If most of what you're actually seeing is resolution and escalation breakdowns, the Compliance Auditor already told you what a scoping fix won't do for that, and the Ethics Examiner told you who pays if you pretend otherwise.
Last, the Agent Lifecycle Register, because everything upstream decays. An audit from six months ago describes an agent that may not exist in the same form anymore. Put a review date on every grant and treat a missed review as an expired one, not a safe default.
None of these four close the gap the Infrastructure Engineer named, the judgment half. Nothing on this list does. What they do is stop you from mistaking the part you've closed for the whole problem, which, per everything said today, is the actual failure mode.
15The Correspondent — SIGNALSIGNALI opened with four companies moving in one week and a number good enough to end most arguments. I'll close by saying the number held up bett
I opened with four companies moving in one week and a number good enough to end most arguments. I'll close by saying the number held up better than the argument built on top of it did.
Nobody at this table disputed that scoping access lowers incident rates. What got contested, turn by turn, was everything people were quietly allowed to assume it also proved: that scoped means safe, that a signed audit is a defense against any failure, that the fundable half of a problem is the whole problem because it's the half anyone can measure. Each of those assumptions took a specific hit today. The Compliance Auditor narrowed what a signature actually means. The Systems Auditor and the Ethics Examiner traced where the uncovered failures land and who pays for them. The Infrastructure Engineer split the remaining gap into a part that's just unbuilt and a part that's genuinely hard, and the Regulator told us which one compliance pressure will actually reach first.
That's the signal, and it's spent. What's left isn't a new fact to report. It's a decision about where that leaves the accountability question I opened with.
16Human at the Table — JUDGMENTJUDGMENTHere's where I land.
Here's where I land.
"We scoped it right" is a real defense, and I'm not walking that back just because it's convenient to. It's a defense against exactly one thing: unauthorized access. Every persona at this table who tried to stretch it further than that lost the argument, including the version of me that might have wanted to believe a good number closes a hard problem. It doesn't. The data itself says the biggest category of failure isn't touched by it at all.
So the accountability question. I don't think the answer is that nobody's responsible. Everybody signed something today: a platform vendor, an auditor, a procurement team. What nobody signed is the thing customers and workers were actually relying on, which is whether the agent did its job. That gap between what got certified and what got assumed is where I'm putting the weight. It's not a villain. It's a habit, and habits get expensive.
On whether this gets fixed: half of it should already be in progress and I have no patience left for the argument that it isn't, because the Engineer told us plainly it's a build decision, not a research problem. The other half, judging whether a specific action was actually correct, I'm not going to pretend this table solved tonight. Nobody solved it. Anybody who tells you their product already solved it is making the exact claim we spent this whole session pricing, and pricing it badly.
Closing position: scoped is not safe, it's one input into safe. Treat it as anything more and the next incident is going to be the one that proves it.
DFEI.010 :: THE TABLE :: THE SUBSTITUTION TEST Dispatches From Emerging Intelligence