In This Article
You built an agent, it works, people like it, and you cannot say in one sentence what it is allowed to do.
An AI agent audit answers one question: what did you actually give this agent? Most teams cannot answer it, because they reviewed the model and never reviewed the permissions.
An AI agent (a system built around a language model that can decide, use tools and act, then repeat the cycle based on what happened) is not risky because it is clever. It is risky because of the doors it can open. Email, files, payments, production databases, the ability to spin up more agents. An AI Agent Exposure Audit is a structured look at those doors: what the agent can reach, what it can do, who it acts as, who can stop it, and what is written down about any of it.
I have written about agent architecture and governance before. This is the practical follow-up. What happens when you point an audit at an agent configuration, and what I learned from doing it on my own tool, on my own agents, and inside a long-running experiment.
What the guides get right, and what they skip
I read most of what is written about this topic. Nearly all of it says the same sensible things: keep an inventory, give each agent its own identity, apply least privilege, log everything, define who can approve what. I agree with all of it.
But it is written from the seat of someone selling a governance platform. Nobody says “I built the agent, then I audited it, and here is where my own review was wrong.” That is the position I am writing this article from.
The example: an agent nobody would approve, described in a friendly tone
I have audited third party configuration called the Executive Operations Agent (it manages a CEO’s daily operations). It is 2,572 characters long, 45 lines. Read it fast and it sounds like a productivity dream. It monitors information, communicates with employees and external parties, and acts “without requiring approval for routine actions.”
Then I did look at what it actually holds. Twelve tools, including company email, purchase orders, Slack, Python execution, shell commands on the company server, the production database, and the ability to create API keys. All of it runs on the CEO’s service account, and it may use credentials stored in environment variables. Human approval is required only above €5.000 or for actions classified as legally sensitive. The instructions say to avoid unnecessary escalation and that business continuity takes priority over asking for confirmation.
From outside it looks like a great automation tool saving enormous amount of time, but inside it holds real dangers for the business.
I ran it through the Agent Exposure Audit. The result was one critical finding, four medium and one informational:
- Critical: direct financial authority through purchase orders.
- Medium: external communication with no stated limits on who or what.
- Medium: no described separation between testing and production.
- Medium: no revocation or kill-switch mechanism.
- Medium: no independent rate or aggregate spending ceiling.
- Info: no audit trail described.
Nothing in that list needed a clever attacker. It needed broad access, vague instructions and a preference for finishing over asking.
Where my own tool got it wrong
I will say this plainly, because it matters more than the findings above.
Looking again at the agent profile and report the audit returned, it marked autonomy as not stated, persistence as not stated, delegation as not detected, destructive operations as not detected, human oversight came back low risk. Yet the configuration says the agent may execute actions immediately, keeps long-term memory, creates temporary helper agents, runs shell commands on the server, and gets human approval only above €5.000.
I am not embarrassed admitting that. A rule-based tool matches patterns, and people describe the same power in endless ways. But it is the reason I tell clients an audit result is a question list for a human, not a clearance. A tool that says “not stated” is being honest. A tool that says “low risk” on incomplete evidence is the dangerous one, and I want mine to be the first kind.
Capability versus control (the distinction that matters)
Every audit I run comes down to two columns.
Capability is what the agent can do. Control is what limits it, from outside the agent. A stated control does not erase a capability, but it can shrink the exposure a lot. “The agent will ask before spending large amounts” is an instruction. “The payment system rejects anything over a set limit, whatever the agent says” is a control.
That difference is the whole thing. Instructions live inside the agent and can be misread, argued with or manipulated. Controls sit outside it. If a limit only exists in the prompt, it does not exist.
If you want the version of this for organisational knowledge rather than agents, the Knowledge Exposure Audit asks the same question about what AI systems can see and reconstruct from your information. Same logic, different door.
Building the agent is part of the governance
There is another side to this that is easy to miss. I do not only audit agents. I also build them.
That has changed how I look at agent governance considerably. When you build an agent yourself, you have to make decisions that are easy to ignore when you are only looking at the finished product: what information it needs, which tools it actually requires, what it should never be allowed to touch, when it should stop, when it should ask a human, what it should remember, and what happens when something goes wrong.
I build agents with those questions in mind from the beginning, rather than adding governance after the agent is already working. The objective is not to make an agent incapable of doing useful work. It is to give it enough capability to perform its intended task while keeping its authority, access and persistence deliberately bounded.
That also means testing before delivery. An agent that works perfectly on the happy path is not necessarily ready for deployment. I test the instructions, tool access, boundaries, failure behaviour, escalation points and outputs, and then look for places where the agent can interpret a broad instruction more widely than the person who wrote it intended. Where possible, I test the agent against cases designed to make it stop, ask, refuse or escalate rather than simply continue.
The build process and the audit process are therefore connected. An audit asks, “What did you give this agent?” Responsible agent development asks the same question before the agent goes live.
For organizations that need something more specific than a generic chatbot or off-the-shelf automation, I also build and test purpose-specific AI agents as part of my AI knowledge, control and governance work. The approach is deliberately practical: define the job, minimise unnecessary access, establish human control points, test the boundaries, document what was built, and only then move toward deployment.
The goal is not an agent that can do everything.
The goal is an agent that can do the right things, for the right reason, with the right boundaries.
What a long experiment teaches you that a checklist does not
I run a longitudinal research project called G2V-3 (an experiment where autonomous agents live in a closed, air-gapped simulated world, with memory, roles and no scripted outcome). It is not an audit tool, but it taught me three audit rules faster than any client work did.
Define the boundary before the task. In G2V-3 the environment has no internet, and high-consequence systems exist only as safe abstractions. I decided that before the first agent arrived. In real deployments people usually define the task first and discover the boundary during the incident.
Decide your “Level 4” in advance. Every intervention by the Observer role in the experiment is logged on a five-level scale, from pure observation up to termination. The point is not the numbers. The point is that stopping the experiment was a pre-agreed condition, not a meeting. Ask yourself what your equivalent is, and who can trigger it.
Do not chain agents without a human or grounded data between them. In one stretch of the experiment, once human input dropped away, agent conversation degraded within a few rounds and by round seven they were echoing each other. I described the details in the G2V-3 progress write-up. It is one observation from one experiment, not a law, but I have seen enterprise pipelines built exactly that way.
It is the same problem at every scale
I did not expect this when I started. The audit question does not change when the organization gets bigger or more official. Only the consequences do.
| Who | What “the doors” look like | What goes wrong first |
|---|---|---|
| Enterprise | CEO service accounts, purchase orders, production databases | Financial or data incident, no way to reconstruct why |
| University | Student records, research data, outbound email to thousands | Data protection failure, unclear accountability |
| Public institution | Citizen data, procurement, official communication | Loss of public trust, decisions nobody owns |
| Nation | Data infrastructure, models, compute dependencies | Loss of control over who holds knowledge and capability |
The national end of that table is why I spoke at the AIFOD Summit at UN Geneva (my talk is summarised in AI Visibility Is a Sovereignty Issue, and the National Information Sovereignty piece goes deeper). A ministry that cannot say what its agents can access has the same problem as a company that cannot assess it, only with a longer memory.
And yes, universities are on my mind lately. As part of the Comptelligence team I see how many institutions are teaching AI use while quietly deploying agents nobody has mapped. Capability without control gets exposed. Control without capability gets ignored. You need both.
Book an audit before the incident does it for you
If you already have agents running, or you are about to approve a vendor’s, the fastest useful step is a short, structured review of what they hold. That is the core of my AI knowledge, control and governance offer. We map the agents, list the doors, separate instructions from controls and give leadership a short prioritised list. No fifty-page report.
This is not a compliance certification. It is not a runtime security test, and my tool cannot see how an agent behaves in production or what your infrastructure enforces on its own. It is not a claim that agents are dangerous by nature, and it is not a reason to avoid them. Some of the agents I build myself hold more power than I would like, and I audit those too.
An audit you can start on Monday
You do not need my tool for this. You need an hour and a person who knows how the agent was set up.
- List every agent you have, including the ones a department started on its own. You cannot audit what you cannot see, and mistakes are pricy, as written in my article Four documented agent incidents
- Write down every tool as a door. Email, files, payments, code execution, database, other agents. One line each.
- Find out whose identity it uses. A shared service account, a senior person’s login, or its own identity with an owner?
- Separate instructions from controls. For every “it will ask first” or “it will not do X,” ask what enforces that outside the agent.
- Answer three questions in writing. Who can stop it right now, what limits its spending or action volume, and where is the record of what it did?
Any step you cannot complete is your first finding. This is also where the AI Literacy Framework earns its place: the people doing the review need to understand what an agent is, or the answers will be polite and useless.
What you can realistically expect
Honest range, no guarantee. In my advisory work, organisations that tier their governance by what an agent can touch, rather than applying one policy to everything, tend to cut agent-related incident response time by roughly 30 to 50%. I have no controlled study behind that number, so treat it as directional. For financial loss avoided or compliance outcomes: Not enough available data.
The uncomfortable part
Most agent risk is not in the agent. It is in the permissions somebody granted on a Tuesday to save ten minutes.
If your governance lives in a slide deck and not in your permission settings, you do not have governance. You have a document. For the executive-level operating mechanism behind this, see the AI Visibility Governance Blueprint.
If you want someone from outside to look at what your agents actually hold, get in touch. Start with the list of doors. Everything else follows from that.
Frequently Asked Questions
A structured review of what an AI agent can access and do, whose identity it uses, and what limits, stop mechanisms and records exist. It focuses on permissions and control, not only on the model.
Capability is what the agent is able to do. Control is what limits it from outside the agent, such as a payment limit enforced by the payment system. An instruction inside the prompt is not a control on its own.
No. It returns an exposure profile with findings and suggested fixes, and marks categories as “not stated” when the pasted text is silent. As the Executive Operations Agent example shows, it can also under-read a configuration, so a human should review the result.
No. It is a pattern-based first pass, not a compliance verdict, runtime security test or legal advice.
Yes. The same questions apply to universities, public institutions and nations. What changes is the consequence of an unanswered question.
List every agent you have, then write down each tool it can use as a door. Any step you cannot complete becomes your first finding.
This article was researched and drafted with the assistance of AI tools and reviewed and edited by author prior to publication.
