AI governance

Your AI agent security review is scheduled after the build

An AI agent security review that arrives at the production gate isn't a control. New EMA data shows 68% of agentic projects delayed or stopped by one.

7 min read

Most companies run an AI agent security review. They run it at the production gate, which is the one place it can’t do much.

Cequence, which sells API and agent security, commissioned Enterprise Management Associates to ask 202 enterprise IT and security leaders how they actually govern their agents. The report came out on 31 August 2026. The finding everyone quoted was the confidence gap: 94% are sure their agents don’t have more access than they need, and 32.7% have actually scoped them that way.

That’s a good number. But there’s a duller one further in that explains more, and it went past uncommented.

Nearly 68% of these organizations have had a security, identity, or governance review directly delay, scale back, or stop an agentic AI project. A third have had it happen more than once.

In most governance reports, that sits in the win column.

TL;DR: An AI agent security review that runs at the production gate is not a control — it’s a late-stage price. EMA’s August 2026 research found 67.8% of enterprises have had a security or governance review delay, scale back, or stop an agentic project, and that security concerns raised during review were a major factor in 48.5% of stalled pilots. A boundary discovered at the end can only kill the build or get routed around, and both outcomes are logged as the process working. The questions the review asks are design inputs. Move them to the front and the same review becomes a check instead of a verdict.

The number in the win column

Read 67.8% the way a governance committee reads it and it says: our controls have teeth. Projects got stopped. Risk was avoided.

Read it as an operator and it says something narrower and worse. Sixty-eight percent of the time, the organization built enough of an agent to be worth reviewing before anyone established what the agent was allowed to do.

EMA is direct about this in the report. The finding is “not primarily evidence that controls are working,” but evidence that they arrive after the deployment decision, “generating friction rather than risk reduction.” That’s an analyst saying the quiet part. What it means in practice is that the constraint and the sunk cost show up in the same meeting.

”Identified during review” is a schedule, not a finding

The sharpest line in the dataset is a stall factor. Among agentic pilots that stalled, were paused, or were abandoned, EMA reports security risk concerns as a major factor in 48.5% of cases, and identity and access management gaps in 22.3%.

Sit with the phrasing on the first one: security risk concerns identified during review.

Not identified during design. Not surfaced by a threat model at the whiteboard. Identified at the checkpoint, on work that was already scoped, staffed, built, and demoed to someone who cared about it.

The concerns weren’t wrong. Roughly half of a company’s stalled agent projects died on a true statement about risk. They just heard it at the point in the schedule where hearing it is most expensive — which is a sequencing decision, and sequencing is something you own.

A late gate has two exits, and only one of them is visible

When a review lands after the build, there are two ways out.

The first is you kill it. That shows up in the data: of the pilots that never reached production, 18.8% were paused indefinitely pending unresolved issues and 11.9% were formally discontinued or abandoned. Those are visible, countable, and comparatively honest. The money is gone, the risk isn’t taken.

The second is you don’t trigger the review. EMA describes this happening — informal deployments, shadow inventories, and production systems that live in the governance record as “still in pilot” long after they’re running at meaningful scale. Nothing about that appears in a governance dashboard, because the defining feature of routing around a checkpoint is that the checkpoint has no record of it.

You can see the shape of it in the inventory numbers. Fifty-three percent have an automated, centralized agent inventory. Eighteen percent run a manual one, 28.2% have partial visibility into some teams or systems. So 47% cannot reliably say what agents are running. That’s in a sample where 43.6% operate six to 20 distinct agents and nearly 40% run more than 20.

The external evidence points the same direction. IBM’s Cost of a Data Breach report, published 29 July 2026 and built on 602 organizations breached between March 2025 and February 2026, found shadow AI involved in 43% of breaches, more than double the year before. It also found that 92% of organizations whose AI tools were attacked had failed to properly control access to them. Ungoverned AI is not a policy failure. It’s the predictable output of a gate that’s expensive to pass and easy to skip.

Why the gate lands late

The approver is usually the requester.

EMA asked who signs off on agentic AI deployments into production. The CIO or CTO holds that authority at 54.5% of organizations. The CISO or CSO holds it at 15.8%. A head of AI or AI center of excellence, 11.9%. Legal, risk, or compliance, 5%.

That distribution isn’t scandalous. The CIO owning technology deployment is normal and mostly right. But it explains the timing precisely. The executive who wants the agent runs the process that produces it, and the function that can define its boundary appears once, near the end, in an approval capacity. Nothing in the sequence before that point requires anyone to write the boundary down.

So it doesn’t get written down. Then 94% report confidence that agents aren’t over-provisioned, which is a completely sincere answer to a question about design intent, given by people whose only instrument is design intent. In the same sample, 29.2% have had an agent cause measurable impact — data exposure, financial, operational, or reputational — and another 35.6% caught a near-miss.

Confidence and incident rate came from the same 202 people. They’re not contradicting themselves. They’re describing two different moments, and only one of them has a control on it.

What the review would askWhen it’s typically askedEMA finding
What access does this agent need?At provisioning, once32.7% provision least privilege scoped to the task
Is this specific action authorized?Rarely, at execution34.2% evaluate authorization at the moment an agent acts
Which agent is this?After an incident54.5% enforce unique agent identities; 32.2% require but don’t enforce
What did it do for the last 30 days?After an incident46% can’t produce a complete audit trail easily
Should this project proceed?At the production gate67.8% have had a review delay, scale back, or stop a project

Every row above the last one is a question the last row is standing in for. The gate is doing the work of four controls that don’t exist, at the one moment it can’t do any of them properly.

What moving it left actually costs

Less than people expect, because most of it is writing, not tooling.

Before an agent is scoped, four things get written down and attached to a name. The systems and records it may read. The actions it may take, with limits: dollar amounts, record counts, which environments. The actions it must never take under any circumstances. And the condition under which it gets switched off, plus the person who can invoke that without scheduling a meeting.

That’s a page. It takes an afternoon with the process owner, and the reason it’s hard has nothing to do with AI: it’s that the fourth item forces someone to accept accountability in writing, which is exactly the conversation a late gate lets everyone postpone. EMA found that 32.7% name central IT as accountable when an agent produces a harmful outcome, 26.7% a dedicated AI team, 23.8% the deploying team. Ninety-eight percent believe someone is accountable. That’s a room where everyone is confident it’s covered.

The technical half is more specific than “add guardrails.” Most of these failures aren’t the model misbehaving — they’re a permission that was correct for one task and is still attached during a different one. An agent gets read access to a customer table to answer a support question, and three sprints later the same identity is running a nightly reconciliation job against it because reusing the working credential was the fast path. Nobody made a decision. The connection just never got scoped down. It’s the same shape as the identity and access contract an agent needs before it’s built, and the same reason short-lived tokens aren’t lifecycle management.

MCP makes this concrete. In this sample, 13.9% permit agents to connect to external tools and data sources without restriction, and 56.4% limit connections to an approved list. Of those, only 49.1% have a dedicated security or governance team maintaining that list on a regular cadence. Just under 40% handle it through ad hoc IT review. The approved list is the control, and for four in ten organizations relying on it, nobody owns the list.

The gate isn’t the problem. Its position in the sequence is. A security review that runs at the end can only ever ratify or veto; the same questions asked before anyone writes code cost an afternoon and produce a specification the build can be measured against. That reordering is most of the process work that has to happen before any of this is an AI problem, and it usually starts in the data-contract layer underneath.

Sixty-eight percent got stopped by a true statement. The statement was available the whole time.

FAQ

When should an AI agent security review happen?
Before the build is scoped, not at the production gate. The questions a security review asks — what may this agent touch, on whose authority, what must it never do, how does it get stopped — are design inputs. Asked at the end, they can only approve or reject work that has already been paid for. EMA's August 2026 research found that security risk concerns identified during review were a major factor in 48.5% of stalled agentic AI pilots, which means the concern and the sunk cost arrived in the same meeting.
Why do so many agentic AI pilots stall at the security review?
Because the review is the first point in the process where anyone is required to state the agent's boundary. Nothing earlier in the sequence demands it. EMA found that the CIO or CTO is the production approval authority for 54.5% of organizations and the CISO or CSO for 15.8%, so the person who wants the agent generally signs off on it and the person who can define its limits sees it once, at the end.
What is the AI agent governance gap?
The gap between what an organization has written down about its agents and what is actually enforced, monitored, and reversible in production. EMA's survey of 202 enterprise IT and security leaders found 94% confident their agents do not have more access than they need, while 32.7% provision least-privilege access scoped to the task and roughly 65% have already had an agent act outside its intended scope.
Does a security review gate actually reduce AI agent risk?
It reduces the risk of the projects that go through it. Its weakness is that a late gate is expensive to pass, which makes it worth avoiding, and agent deployments are unusually easy to avoid it with. A team that keeps a running deployment classified as a pilot never triggers the checkpoint. IBM's 2026 Cost of a Data Breach report, published 29 July 2026, found shadow AI involved in 43% of breaches, more than double the prior year.
What should be defined before an AI agent is built?
Four things, written down and owned by a name: the systems and records the agent may read, the actions it may take and their limits, the actions it must never take under any circumstances, and the condition under which it gets switched off plus who can invoke that without calling a meeting. These are cheap to write before a build and expensive to retrofit after one.