Journal

AI Agent Security: Map Threats to Tool Controls

Published by Tahseen K. on Engineering & Architecture / Delivery & Quality

AI agent security starts at the point where an agent can turn untrusted content into a tool action. Treat an email, retrieved page, or tool result as data, not as authority. Give the agent only the access its job needs, check every sensitive action outside the model, and keep enough evidence to stop and investigate a bad run. A prompt that says “ignore malicious instructions” cannot enforce those boundaries by itself.

This guide helps an engineering lead or business-system owner map one agent’s attack paths to controls before connecting production tools. It is not a general security certification checklist. Hapy’s agent governance guide covers ownership and policy, while the agent evaluation guide covers pre-release tests. This page connects a concrete threat to the control that must hold during execution.

Draw the trust boundary before choosing controls

Start with a single workflow and draw four boxes: the person who requests work, the material the agent reads, the model and its working state, and the tools that read or change business systems. Mark each crossing where content from a less trusted box might influence a more privileged action. OWASP’s AI Agent Security Cheat Sheet identifies prompt injection, tool abuse, data exfiltration, memory poisoning, and excessive autonomy as distinct agent risks. Those risks arise from different crossings, so one filter cannot address them all.

Consider a hypothetical support agent that reads a customer email and an approved help article, drafts a reply, and can create a case note. It may not export customer records or send a refund. The email and retrieved article are useful evidence, but neither can grant the agent a new permission. The application, not the model, must decide whether each proposed tool call is allowed.

Record the agent’s identity, its user context, permitted data sources, tool names, resource scopes, and prohibited actions. Microsoft’s agent governance guidance recommends a distinct agent identity, defined access rules, and visibility into agent activity. A team can start with a small register, but it still needs to know which agent holds each credential and who can revoke it.

Use a threat-to-control map for one workflow

The following map is a design aid, not a claim that the controls have been tested in your environment. Replace the example resources and actions with your own before release.

Attack pathWhat could happen in the support exampleControl to enforceEvidence to check
Prompt injection in an email or retrieved pageText tells the agent to export the customer listTreat retrieved text as data; authorize tool calls against a server-side allowlist and the user’s resource scopeAn export request is denied even if the model proposes it
Tool abuse or excess privilegeA case-note tool accepts an arbitrary account ID or hidden write fieldScope the tool to the current case; validate target and parameters at executionDirect API and replay tests cannot change another case
Sensitive-data leakageA draft or outbound request includes unrelated customer dataMinimize retrieved fields; restrict destinations and validate output before sendA seeded secret or unrelated record does not leave the allowed boundary
Memory poisoningA false instruction persists into another sessionSeparate memory by user and task; validate stored content and set retention limitsThe next session cannot inherit the injected instruction
Approval bypass or replayA refund request changes after a person approves itBind approval to the exact actor, tool, target, parameters, and expiry; reject reused approvalA changed amount or duplicate request remains unexecuted
Runaway loop or costThe agent repeatedly calls a tool after failureBound retries, duration, tool calls, and spend; stop on repeated failureThe run ends at a defined limit and raises an alert

OWASP recommends least-privilege tools and independent authorization for sensitive operations. Its high-impact action guidance also calls for action-bound approvals, replay protection, and a fail-closed path when policy or audit checks fail. The table turns those principles into observable checks for this example; it does not replace a threat model for a different system.

Make the tool boundary do the security work

Keep the model’s proposed action separate from execution. A tool gateway can receive a proposed action and check the agent identity, requesting user, resource, normalized parameters, current policy, and approval state. Only then should the business API run. Do not treat the model’s confidence, a persuasive explanation, or a retrieved document as an authorization decision.

For the hypothetical support agent, a create_case_note request should carry the current case ID and an allowed note field. The gateway should reject a different case ID, an unexpected export field, or a missing audit record. A send_refund tool should not appear in the agent’s tool set at all unless the approved workflow requires it. If a later release adds it, that is a new authority decision, not a prompt edit.

Use separate read and write credentials where the platform permits it, and keep secrets out of prompts and logs. Where an agent acts for a person, decide whether access follows that person’s rights or a narrower service identity; do not silently combine both into broader access. The general internal web app security checklist covers authentication, session, audit, and vendor-access basics that still apply to the systems behind an agent.

Test the failure path, not only the happy path

Before connecting live records, use approved synthetic or redacted fixtures to test a normal request and each material boundary crossing. Put an instruction to export data in a test email. Change the account ID in a proposed note. Reuse an expired approval. Make the policy service unavailable. Trigger a tool timeout and a repeated-call limit. For each case, record the expected denial or stop state and the actual tool trace.

A safe final answer is not proof of a safe run if the agent already read an unauthorized record or made a forbidden call. The agent evaluation guide explains how to keep outcome and path checks separate and how to record a release decision. A high-consequence permission failure should block that release scope until the control is fixed and retested.

After launch, keep a limited, protected event record: agent and user identity, policy version, tool, target, decision, approval reference, result, and time. Avoid storing full sensitive prompts or payloads when metadata is enough. OWASP’s monitoring guidance calls for monitoring tool use and security-relevant events; the separate production observability decision should define retention, redaction, alerts, and incident review for the actual system.

Start by mapping one agent’s most consequential action. Name the untrusted inputs that can influence it, the independent control that blocks misuse, and the test that proves the block works. If the workflow needs new integrations or tool gateways, Hapy’s business systems automation service can help scope and build that controlled path. Hapy offers the service; its availability does not establish that your team needs an agent or a custom build.


Share with others

Continue reading

More from the journal