Who Authorized That Agent to Touch the Grid?

AI agents are already operating against enterprise and OT systems — holding credentials nobody issued deliberately, crossing boundaries that were supposed to be controls. A practitioner framework for governing non-human identity before the regulator asks.


Most organizations deploying AI agents today are running identity controls that cannot produce the evidence they are about to be asked for.

Not will struggle to. Cannot.

And if you run a corporate network alongside a segmented one that actually operates something, that evidence isn’t only for a financial statement. It’s for the moment somebody asks why a valve closed, why a breaker opened, why a line stopped, or why a terminal started talking to a system it had never reached.

Agent governance today is usually assembled from whatever was already in place — a registry someone maintains, a gateway that agents are meant to route through, a secrets vault, cloud entitlement tooling, logs in a SIEM, quarterly certification. Add whatever else is on your list — the argument doesn’t depend on which pieces you have. Whatever the mix, it rests on one assumption: that the thing being governed will still be there when the control gets around to looking.

Which is why the same three failures appear regardless of toolset. That isn’t a maturity gap time closes. It’s an architectural mismatch, and adding another tool built on the same assumption doesn’t move it.

Three questions, three failures

Audit has asked the same three things of every actor in an enterprise system for twenty-five years. The questions aren’t new. What’s gone is the ability to answer them.

Completeness — can you produce the population?

A registry lists the agents somebody remembered to register. A gateway log lists the agents that chose to route through it. Neither is a population, and that difference isn’t academic: completeness is a property of the population, not of the instrument measuring it.

In environments like these it’s unusually hard to say where that population even ends. It runs across corporate IT, cloud platforms, remote and field sites, store and plant floors, contractor tooling, and equipment in a substation that predates your identity program by a decade. Segmentation doesn’t help you count it — it guarantees part of the population is invisible to whatever is counting.

An auditor has a term for a control that can’t capture its whole population no matter how well it runs. Not “needs improvement.” Deficient by design — and you don’t fix a design deficiency with a better schedule.

Traceability — can you follow an action back to its authority?

A setpoint changes at 2:14 a.m. The log records that a service account did it. That’s the entire entry.

Which agent? Acting for whom? Triggered by what? The log doesn’t know, because several agents share that credential and the credential is all the target system ever saw. The attribution wasn’t buried — it was destroyed at authentication, and nothing downstream recovers it.

In a financial system that question surfaces during an audit. Here it surfaces during an incident review, and the person asking is a regulator, an insurer, or a safety investigator.

Accountability — is there a human who answers for it?

The pain in an access review usually isn’t excessive access. It’s the orphan: a credential with no accountable owner. Nobody approves it because nobody knows what it does. Nobody removes it because nobody can say what breaks. It survives review after review by being nobody’s problem.

Distributed operations manufacture orphans structurally. Contractor engagements end; the accounts don’t. Acquisitions arrive with an identity estate inherited whole. And termination is triggered by HR — there is no HR feed for agents, and certainly none for an agent whose owner works for your integrator.

Two populations, two hard problems

Ephemeral agents expire before the evidence does. An agent assumes a role at 1:15 a.m., queries a system, terminates at 1:18. Your nightly reconciliation runs at 4:00. No account to certify, no owner to ask, no session to reconstruct — and shrinking the interval doesn’t help, because nothing runs faster than three minutes.

Persistent agents accumulate authority. It gets shared. Every team sharing it needs one more permission added, and permissions are added far more often than removed. Nobody rotates it, because nobody can say who would break.

Between them, that’s most of the agent population — and it’s why detection is running out of road.

A quarterly review has a ninety-day detection window. An agent can enumerate a system and act on it in under a minute. Detective controls were built against human misuse, which is slow enough to interrupt. Against a machine actor, even a daily control isn’t a control — it’s a post-mortem with better formatting. And when it fires, what’s the remediation? The credential expired. The agent terminated. The valve is already closed.

You cannot supervise machine speed. You can only constrain it.

Detective controls don’t disappear; they change jobs, from catching the bad transaction to proving the preventive control held. But the center of gravity has to move. If the actor may not exist when you look, the control has to operate when it acts.

What’s actually being deployed

Not chatbots drafting emails.

Agents already in production monitor wells and pipelines against live SCADA data, optimize field production against control systems, and run safety and environmental compliance checks. Arriving next is a different category: write access to pipeline shutoff valves, real-time control of refinery yield, autonomous drilling with no human in the loop — and, the one worth pausing on, agents filing regulatory reports on the organization’s behalf, exercising its legal authority.

Nor is this specific to energy. In retail, agents reconcile inventory, adjust pricing and resolve payment exceptions — work running directly against, or beside, the cardholder data environment PCI segmentation exists to protect. In manufacturing, they schedule production, tune machine parameters against live quality data, and place supplier orders on their own authority.

Different floor, same structure: a non-human identity, holding a credential nobody issued deliberately, reaching across a boundary that was supposed to be the control.

And the timeline is compressing. In August a major model developer published a specification for AI agents to operate physical instruments directly — built on read and write primitives whose own worked examples are “get temperature” and “set temperature,” with devices and agents discovering each other across the network. It’s model-agnostic, it runs over standard protocols, and it’s already in the hands of manufacturers.

Read that as an identity problem and it says something uncomfortable: the interface between agents and physical equipment is being standardized right now. The governance layer isn’t.

The frameworks already agree. We’re already failing them.

Whichever regimes apply to you — and they vary enormously — they converge on a requirement written long before anyone said “AI agent.”

IEC 62443 draws a line most enterprise programs never did: SR 1.1 covers identification and authentication of human users, and SR 1.2 requires the control system to identify and authenticate all software processes and devices. Non-human identity as a distinct, named obligation, in a standard published in 2013. PCI DSS arrived at the same place from the opposite direction — version 4 added explicit requirements for application and system accounts, including that their credentials not sit hardcoded in scripts or config files, enforceable since March 2025. NERC CIP has moved steadily toward governing vendor and remote access; TSA’s directives require access control over both local and remote access to critical systems; CISA and NIST both speak of credentials for users, devices and processes.

Different bodies, different scopes, different decades. The same instruction: know every actor, authenticate it uniquely, give it an owner, scope what it can do, and be able to revoke it.

Not one of them mentions AI agents. All of them already cover them — because they were written around the idea of an authenticated actor, not the idea of a person. Agents didn’t create a new obligation. They created a much larger population subject to one you already have.

And consider how we’re doing. State-affiliated actors have spent two years taking over internet-exposed controllers at water and energy utilities using the credentials those devices shipped with. Not exploits — default passwords. It has run since late 2023 and was still claiming US utility victims this spring.

A requirement to uniquely identify and authenticate every device has existed since 2013, and devices are being commandeered because nobody changed the password. That is the distance between what we’ve agreed to and what we’ve deployed — today, at human speed. Now add a population that provisions itself and acts in three-minute bursts.

One comfortable assumption worth retiring: segmentation is a real control, but the air gap is largely an illusion in an industry with cloud EMS, remote operations centers, historians replicating into corporate analytics, and vendors dialing in for maintenance. Agents don’t respect that boundary, and cloud identity tooling often can’t reach across it to govern them.

What changes

Preventive, at runtime, enforced in the entitlement rather than asserted in a prompt.

  • Access that doesn’t exist until the work does. Not a smaller standing grant — no standing grant. The credential is minted for the task and expires with it.
  • Scope written into the permission. Put the constraint in the entitlement and it’s a steel barrier bolted into the rock above the drop. Put it in the agent’s instructions and it’s a painted line and a sign. In daylight the two look identical, and your policy document records them identically. You find out which one you built exactly once — at speed, in the dark.
  • Ownership bound at creation, because there’s no feed coming later to catch it.

Underneath all of it, discovery. You can’t enforce across a population you can’t see, and you can’t scan your way to it — a scan has an interval, and the interval is where the problem lives.

So you stop looking for the agent and start watching the moment it asks for permission. An agent can avoid your registry, your gateway and your scan window. It cannot avoid authenticating. It cannot reach one sensor, one historian or one control system without first obtaining a credential — and that moment is already instrumented.

We don’t take attendance. We watch the door.

Three ways in, three different problems

What to do next depends on how the agent got its access, and each pattern breaks traceability and accountability differently.

When it borrows an employee’s identity

The target system logs a person. The trail looks immaculate, right up until somebody asks whether the employee changed that setpoint or an agent did it while acting as them. Nobody can tell — including the employee. The problem isn’t that nobody owns it; it’s that the record names an owner who didn’t act, may not have known, and couldn’t have stopped it. A false owner is worse than no owner, because it satisfies the question and stops anyone looking further.

Two things resolve it: records carrying both principals, so the log reads this agent, acting for this person; and access scoped to the intersection of what the employee can reach, what the agent is approved to do, and what the task requires — not everything the badge opens.

When several agents share a service account

Attribution is destroyed at authentication. The fix isn’t only splitting or rotating the credential — it’s enumerating the consumers, which agents are holding it right now. Most tools can tell you the credential exists and who owns it on paper; far fewer can say who is actually using it. Owning a credential is not the same as being accountable for what’s done with it, and access reviews have quietly conflated the two for years.

When it’s issued a short-lived token

The best pattern, with the subtlest failure: the token is scoped and expires, which is right, but nobody approved this issuance. The approval happened once, when someone wrote the trust policy that lets this workload mint credentials — possibly years ago, possibly someone who has left. So the token’s claims must carry agent and run identifiers into the target log, or you’ve built perfect short-lived access with no durable record of what it did. And the trust policy itself has to be a reviewed object. That’s where the standing authority now lives, and almost nobody certifies it.

Fix attribution, and watch what happens next

Shared credentials across autonomous agents won’t survive scrutiny for long, and audit thinking is heading toward distinct attribution between an agent and the credential it executes with. No standard-setter has published that yet — but it’s where every conversation about traceability ends up.

That direction is right. Satisfied naively, it detonates. Machine identities already outnumber people substantially. Give every departmental, site-level and ephemeral agent its own static account and you have hundreds of thousands of credentials. Nobody certifies that quarterly. Nobody rotates it without breaking a pipeline. Every finished project leaves standing privilege behind.

Three ways out, and it’s worth knowing which you’re on: dedicated static accounts (compliant, brutal, viable only with automated discovery, owner binding and dormancy cleanup); federated workload identity with short-lived tokens (the right answer, and only if claims carry attribution); or shared accounts with runtime context restored out of band (a transition, acceptable only while you can correlate sessions to what happened at the target).

The strategic point: attribution at scale isn’t a certification problem you can staff your way out of. It’s an architecture decision, and the operating load lands on automation or nowhere. Which makes three things non-optional — identities bound to an owner at issuance, dormancy decommissioned before anyone thinks to look, and review effort tiered by risk. Reviewing two hundred thousand things at the same depth is functionally the same as reviewing none of them.

Two directions, and most tools govern one

Inbound is who may invoke this agent — which people, systems and other agents can prompt it into action. Outbound is what it can reach — credentials, tools, data, and how sensitive that data is.

Identity providers and gateways are built for inbound. Vaults and cloud entitlement tools cover part of outbound. Very little does both against the same identity, and the interesting risk is the combination: a low-privilege agent anyone can invoke, which can call a high-privilege agent nobody reviewed.

That’s also where separation of duties has quietly relocated. If agent A can call agent B, A’s effective authority includes everything B can do. An agent that reads sensor data is unremarkable until it can invoke the agent that writes setpoints. Which points to a requirement most programs haven’t absorbed: agent-to-agent delegation is an access path and has to be governed as one. Today it’s built as an integration, reviewed as architecture, and never appears in an access review.

Worth asking any vendor which direction they govern. Most will answer one and hope the question doesn’t come back.

What the answer actually has to do

Those paths are governed by different tools today — cloud entitlement management for roles, identity governance for accounts, a secrets manager for credentials, a SaaS posture tool for OAuth grants, and usually nothing for delegated employee identity. Each reports confidently on its own slice. None can answer what actually gets asked: what can this agent reach, and on whose authority?

That question crosses every boundary, and it doesn’t decompose. Effective access is a join — the role assumed, the entitlements on the account used, the secret checked out, a grant approved three years ago, and the sensitivity of the data underneath. Five partial answers don’t add up to one complete answer, and an incomplete picture doesn’t return “unknown.” It returns a confident, wrong answer, which is worse, because you act on it.

So the requirements get specific enough to be a checklist.


Before your next agent goes live

What a governed agent population requires

  1. See every non-human identity the moment it authenticates. Cloud control plane, secrets layer, endpoint, pipeline. Not on a schedule. Not only where a gateway was deployed.
  2. Cover every path to access. Workload identity, shared service account, borrowed employee identity, a secret nobody vaulted, a consent granted years ago, the assistant installed this morning.
  3. Bind an owner at issuance and keep it current. Including automated dormancy detection and decommissioning.
  4. Preserve attribution through the joins that destroy it. Both principals on delegated actions, the real consumers of a shared credential, agent and run identifiers carried in the token.
  5. Enforce preventively, at runtime. Access that expires with the task, scope written into the entitlement, authorization evaluated at the call.
  6. Govern both directions. Inbound and outbound — with agent-to-agent delegation treated as an access path, not an integration.
  7. Know what the data underneath is worth. So “what can this agent reach” is a question about risk rather than permissions.
  8. Return all of it as one current answer. Including for an agent that no longer exists.

Now notice what that list is not: it is not eight products.

Each requirement is individually available somewhere. They only work when they resolve to the same object. The moment ownership sits in one system, secrets in a second, cloud entitlements in a third and data classification in a fourth, you’re reconciling — and every seam is a place an agent operates unobserved. In this sector that isn’t theoretical: the seam between corporate and operational identity is the exact lateral path we’ve been losing to for a decade.

Which is the honest reason this converges on a platform rather than a portfolio. Not because fewer vendors are tidier, but because what did this thing do, and who authorized it is one question — and only something holding the whole chain, from human sponsor through agent and credential to the data touched, answers it in one move. Anything less is reconciliation, and reconciliation is a project, not a control.

Six questions worth asking your own team

  1. Can we produce a complete list of the agents operating against our environment — including the ones a vendor deployed and the ones running on somebody’s laptop?
  2. Which of our agent controls are preventive and which are detective? For the detective ones: what’s the remediation when the actor no longer exists?
  3. When an agent acts on an employee’s behalf, does our record name the employee, or the employee and the agent?
  4. Can we list which agents are holding a given service account right now — not who owns it on paper, but who is using it?
  5. Which of our agents can invoke another agent, and does that delegation appear anywhere in an access review?
  6. Could we tell a regulator — or an incident responder — who authorized a specific action taken by an agent that ran for three minutes last Tuesday, as one answer rather than five exports?

If the last one doesn’t have an answer, it’s worth an hour of somebody’s time.

Because right now, agents are acting against your systems — holding credentials nobody issued deliberately, on the other side of boundaries that were supposed to be the control. The question is coming, and it will not arrive in a planning session.

It arrives at 3 a.m., with an outage running, a regulator on the line, and somebody scrolling a log that says a service account did it.

Every organization will eventually produce a full account of what its agents could reach and who authorized them. The only thing still open is whether you build that answer deliberately, in advance — or assemble it afterward, under the lights, with counsel in the room.

One of those is a control. The other is a disclosure.