Blog / Field guides

Enterprise AI agent governance: 5 controls before production

A working agent demo lacks the controls production needs. Five controls that define what an agent may see, do and change before it goes live.

Xagent team · 22 July 2026 · 5 min read

Enterprise AI agent governance decides who an agent acts for, what it may access, which actions need sign-off and how you reconstruct what happened. A demo skips all of this because one person supplies the credentials and watches every step. Production needs five controls: identity, action records, data isolation, guardrails and end-to-end visibility.

Why does a demo need different controls from production?

A demo runs inside a bubble of trust. One developer supplies the credentials, chooses the data and watches every action. Production removes those assumptions. Users have different permissions, tools can change external systems, several agents may run at once, and auditors may ask for evidence months later.

Governance therefore covers what agents do, not only what models say. The five layers below are a practical starting point.

1. Who is the agent acting for?

Every agent action needs an identity chain: the agent, the user or service that asked for it, and the authorisation used.

Do not give a general-purpose agent a shared administrator credential. Run tools with the requesting user’s own permissions where possible. Scope service identities to one workflow, environment and set of resources.

Use short-lived credentials for high-risk tools. Require sign-off before any increase in privilege, and let the raised permission expire automatically.

The test is simple. Can the team say who authorised every consequential action, which policy allowed it and which credential was used?

2. What should you record about each action?

Ordinary application logs are not enough when an agent chooses tools and takes several steps. Record what you need to reconstruct the run:

  • Agent, user and tenant identifiers.
  • Model, policy and workflow versions.
  • Tool name, parameters and result.
  • Data sources accessed and whether access was allowed.
  • Sign-off, rejection and override events.
  • Latency, token use and estimated cost.
  • External side effects and whether they were reversed.

Do not rely on storing a model’s hidden reasoning. It may be unavailable, unstable or sensitive. Store the observable tool traces, policy decisions, evaluation results and a short decision summary written for operators.

Apply retention limits, redaction and access control to these records. They can hold the same sensitive information the rest of your controls protect.

3. How do you keep data isolated?

Prompt instructions cannot enforce tenant isolation. The retrieval system and each tool must check authorisation before returning data.

Apply access control separately to document retrieval, database queries, tool responses and memory. Filter data before it enters the context window. Keep tenant identifiers attached through caches, vector indexes and background jobs.

Test the same question from users with different permissions. Each answer should reflect only the data that user may see. Red-team exercises should also test indirect prompt injection, cache reuse across tenants and tools that return more fields than asked for.

4. Which guardrails and stop controls do you need?

Classify tools by consequence. Reading a public knowledge base is different from deleting records, sending an external message or moving money.

For consequential actions, set limits on amount, data volume, destination, frequency and scope. Require a person to sign off when an action cannot be undone or crosses a threshold. Show the reviewer the proposed action and the evidence, not a bare “approve” button.

Add circuit breakers for repeated tool failures, unusual activity, unexpected destinations, cost spikes and policy breaches. Define a safe stopped state. Operators should be able to switch off one agent or tool without taking unrelated services down.

Where an action can be reversed, test the reversal. Where it cannot, raise the sign-off requirement.

5. What does end-to-end visibility look like?

Model metrics describe only part of an agent run. Operators need one trace that links the request, retrieval, model calls, tools, sign-offs and final side effects.

Monitor task success, policy denials, fallbacks, tool errors, latency, cost and how often people override the agent. Break these down by agent version, tool, tenant and risk level.

Behaviour can change after a model, prompt, tool or knowledge update. Version those dependencies and compare them during rollout. More successful tool calls is not an improvement if people are reversing more of them.

How do you set a risk tier for each agent?

Base the tier on data sensitivity, tool authority, reversibility, autonomy and external impact. A research assistant that reads approved documents should not share a release process with an agent that changes customer accounts.

A low-risk agent may launch with logging, access control and routine evaluation. A high-risk agent may need threat modelling, formal sign-off, a person confirming each action, closer monitoring and a tested incident plan.

The NIST AI Risk Management Framework organises AI risk work around govern, map, measure and manage. NIST’s Generative AI Profile applies it to generative AI. The OWASP Top 10 for Agentic Applications lists common agentic security risks, and AWS publishes guidance on governing agentic AI at scale. None of these is a certification. Use them as inputs to your own control design.

What does a four-week starting plan look like?

  1. Week 1, inventory. List every production and pilot agent, its owner, models, tools, data sources, credentials and side effects. Switch off anything with no owner.
  2. Week 2, identity and tool policy. Replace shared credentials, add authorisation at the tool level and define which actions need sign-off. Start with the least privilege the workflow needs.
  3. Week 3, traces and stop controls. Link model and tool events in one trace. Add rate limits, circuit breakers and a tested kill switch. Set retention and redaction rules.
  4. Week 4, failure and abuse testing. Try unauthorised retrieval, prompt injection, excessive spending, repeated actions and cross-tenant access. Run a failed-tool and reversal exercise. Record owners and due dates for what you find.

An agent ready for enterprise use has a bounded identity, limited tools, enforced data access, proportionate sign-offs, a visible execution path and a safe way to stop.

Where does Xagent fit?

Governance starts with your organisation deciding what each agent may do. A platform should make those decisions easy to see. In Xagent, every tool call and result is visible while a task runs, a running task can be paused, and run logs and traces are kept afterwards. There are admin and team workspaces, and sign-in is with Google (OIDC). Self-hosting is available on the Enterprise plan, and agents can run on models self-hosted through Xinference.

To see how a run looks step by step, read what happens after you type a request. If you are still deciding how much autonomy to give, start with AI agents vs assistants.

Questions

What is AI agent governance?

It is the set of controls that decide who an agent acts for, what it can access and change, which actions need sign-off, and how you reconstruct what it did.

Is the NIST AI Risk Management Framework a certification?

No. It is a voluntary framework for managing AI risk. Use it, with sources like the OWASP agentic list, as input to your own controls.

Should AI agents use a shared admin account?

No. Run tools with the requesting user’s own permissions where possible, and give service identities only the access one workflow needs.

How long does it take to put basic agent governance in place?

A focused team can cover inventory, identity, traces and failure testing in about four weeks. High-risk agents need more, such as threat modelling and a tested incident plan.

Try Xagent. See all use cases, or book a demo on your own workflow.

Stop repeating the same work.

Book a demo on your own workflow. We will hand the busywork to agents while you watch.

Book a demo

Discover more from Xagent

Subscribe now to keep reading and get access to the full archive.

Continue reading