AI Governance

Compliance for Autonomous Agents: MCP Servers, Tool Use, and What Auditors Will Ask

LowerPlane Team••11 min read

TL;DR

  • • Agents break the usual audit playbook — they act on behalf of users, hold credentials, invoke tools, and make non-deterministic decisions with real blast radius
  • • The threat model shifts from "can the model output something bad" to "can the model be tricked into doing something bad" — prompt injection graduates from an inconvenience to a security incident
  • • Five existing SOC 2 controls stretch to cover agents (CC6.1, CC6.3, CC6.7, CC7.2, CC8.1) if you describe them right — no separate framework needed today
  • • MCP servers introduce a new class of external attack surface — treat each MCP server like an API integration with its own risk assessment
  • • Draft an "Agent Use Policy" that legal, security, and product all sign — it becomes the artifact every auditor and buyer will ask for

Six months ago, agents were a demo. Today they're shipping in real products — closing tickets, drafting PRs, executing trades, moving money. Every one of those verbs is what auditors care about, because every one of them can go wrong. This is the compliance shape that's emerging.

Why Agents Break the Usual Audit Playbook

A traditional app has a fixed control flow: user takes action → your code validates → your code executes. An agent inserts a non-deterministic decision-maker in the middle: user takes action → the model decides what to do → the model calls tools that execute. Three things change:

Blast radius
A traditional bug affects the request that triggered it. An agent bug — a jailbreak, a tool-call confusion, a hallucinated argument — can affect any resource the agent has credentials to touch.
Authorization scope
Traditional auth: user X can do action Y. Agent auth: user X asked an agent to do action Y, and the agent chose to do actions Y, Z, and W. The audit trail needs to reconstruct the decision.
Non-determinism
The same input can produce different tool calls next week when the model version changes. Change management now has to include prompt and model version bumps.

The Threat Model Auditors Are Converging On

The concise version: an attacker who controls any input the model reads can influence what the model does. That includes not just the user's prompt, but retrieved documents, tool outputs, MCP server responses, or content fetched from the web. This is the "prompt injection → tool abuse → data exfil" chain that shows up in every serious agent post-mortem.

The chain, step by step

  1. Injection. Attacker plants text somewhere the agent will read — a PDF, an email, a webpage, an MCP server response.
  2. Instruction confusion. The model treats the attacker's text as an instruction from the user.
  3. Tool abuse. The model calls tools using the agent's (privileged) credentials, following the attacker's instructions.
  4. Data exfil or destructive action. The tool call reads sensitive data and writes it somewhere the attacker can reach, or performs a destructive action the user never requested.
  5. Ambiguous audit trail. The user "asked" for the action — the logs show a request-in-flight. Figuring out that it was really an attacker takes hours of forensics.

Auditors do not expect you to make this impossible. They expect you to have documented controls that reduce likelihood and blast radius, plus logging that lets you reconstruct what happened. That's the SOC 2 framing you can work with today.

Mapping Agents to Existing SOC 2 Controls

ControlAgent interpretationEvidence
CC6.1Agent runs with the least privileges required. Different agents have different credentials.IAM role definitions per agent, quarterly review of role scope.
CC6.3Agent tools are allow-listed. New tools require a review.Tool registry, PR history of tool additions, reviewer signoff.
CC6.7Data the agent reads and writes is protected in transit; MCP server connections are authenticated.TLS enforcement config, MCP server auth token rotation records.
CC7.2Anomalous agent behavior triggers an alert — unexpected tool sequences, credential fetches, high-cost actions.Alert rules, sample alerts, on-call runbook, kill-switch procedure.
CC8.1Changes to system prompts, tool definitions, or MCP server URLs follow change management.PR history, approver evidence, version-controlled prompt files.

MCP Server Hardening Checklist

Model Context Protocol servers moved fast from "interesting protocol" to "attack surface in production." If your agents talk to any MCP server — internal or third-party — treat each one as an integration with its own risk profile.

□Every MCP server has an owner, a documented purpose, and a data-class scope
□Third-party MCP servers go through the standard vendor risk assessment before enablement
□MCP server connections use short-lived auth tokens, not long-lived shared secrets
□Tool descriptions returned by the MCP server are treated as untrusted input — do not let them auto-expand your agent's tool allow-list
□Rate limit per user, per MCP server, per tool
□Log every tool call: user, agent, MCP server, tool, arguments (redacted for PII), response summary
□Alert on abnormal patterns: unusual tools called in sequence, high call volume, tool calls that touch sensitive data classes
□Kill switch: an ops-team-runnable command that disables a specific MCP server across all agents in under a minute
□Quarterly review of enabled MCP servers — anything unused for 90 days gets disabled

Draft: The Agent Use Policy

Buyers, auditors, and your own legal team all want the same artifact: a policy that says what your agents can and cannot do, and who signs off on changes. Cover these clauses at minimum.

Scope
Which agents this policy covers — internal (employee-facing) versus external (customer-facing) — and which are explicitly out of scope.
Authorization model
Agents act with a service identity, not the user's identity. Tool access is granted per agent, not inherited from the user session.
Human oversight
Above a defined impact threshold (e.g., moving money > $X, deleting > N records, sending external comms), the agent stops and asks a human.
Tool allow-list
Agents may only invoke tools registered in the tool registry. Adding a tool requires a documented review.
Data class handling
Agents may not send regulated data (PHI, PCI, secrets) to third-party MCP servers without an explicit exception approved by the security team.
Logging and retention
Every tool invocation is logged with user, agent, tool, arguments, and response. Retention: 90 days minimum, per applicable regulation.
Incident response
Suspected prompt injection or tool abuse triggers the standard incident response procedure with an agent-specific runbook.
Change management
System prompts, tool registry, MCP server URLs, and model versions are version-controlled and follow the change-management process.
Kill switch
The security team has the authority and the tooling to disable any agent or MCP server without engineering escalation.
Review cadence
This policy is reviewed by security, legal, and product leads every six months or after any material incident.

Evidence for Autonomous Decisions

When an agent made a decision that a user, auditor, or lawyer wants to review, what do you show? Build these three artifacts and you'll cover most reasonable asks.

  1. Decision transcript. The user prompt, the retrieved context, the agent's reasoning (if you log chain-of-thought), the tool calls made, and the final response. One record per user session.
  2. Tool-call log. A structured log of each tool invocation — timestamp, user, agent, tool, arguments (redacted), response summary. Queryable by user or by tool.
  3. Human-review record. When the agent hit an impact threshold and stopped, who reviewed, when, what they decided. Stored where any other approval audit trail lives.

How LowerPlane Handles Agent Compliance

  • →Agent inventory and MCP server register as first-class objects
  • →Tool-registry review workflow with security signoff
  • →Agent Use Policy template pre-mapped to SOC 2 CC controls
  • →Tool-call log connectors and drift alerting
  • →Shared controls between SOC 2, ISO 42001, and NIST AI RMF for the same evidence
See Agent Compliance in Action

Ship Agents Without Losing Sleep — or Deals

LowerPlane gives you the agent inventory, tool registry, and evidence log auditors and buyers now expect from AI-native products.

Book a Demo