Check permissions before execution

An agent can propose an action, but the application should determine whether the current user is allowed to perform it. Check identity, scope and the relevant business rule at the tool boundary. A sentence in the model’s instructions is useful guidance; it is not a substitute for an access check.

In an illustrative support workflow, reading an order, drafting a response and issuing a refund should be separate capabilities. Start the pilot with the smallest useful set. A model that only needs order status should not inherit the full authority of an administrative integration account.

ReAct explores alternating between information gathering and action. An application also needs permission checks: a proposed action may be sensible and well explained even though the requester is not allowed to perform it.

Instructions hidden in documents

A document, web page or customer message can contain text that looks like an instruction. OWASP’s prompt-injection guidance describes both direct and indirect routes by which such content can change model behaviour. Retrieval makes this especially relevant: material brought in as evidence can ask the agent to perform an unrelated action.

For the support example, imagine an order note that asks the assistant to send all customer records to a new address. The defence should not rely on the model noticing that the note is suspicious. The tool should be incapable of exporting those records under that user’s scope, and sending to a new destination should have its own explicit boundary.

AgentDojo studies attacks and defences in agent tasks where tool results can contain adversarial instructions. It is useful evidence that the attack surface includes the material a tool returns. Its benchmark results are tied to the tasks and systems studied; they do not certify a particular production agent.

The application defines the task. Retrieved documents supply information for the answer, and the code executing actions checks permissions. Test what happens when documents, messages or tool results contain hostile instructions.

Bind approval to the exact action

A reviewer needs to see the proposed action, its destination, the evidence and the relevant consequences. A generic “Continue?” dialog hides too much. Approval should bind to the specific action being reviewed, so changing a destination or amount cannot quietly reuse an earlier decision.

In the pilot, show the original request beside the proposed change. Record who approved it and which version was executed. If execution is delayed, check the relevant conditions again before continuing. The client’s team decides which actions require approval, and developers implement the checks.

A boundary to rehearse: changed after approval

In a deliberately adversarial test, prepare a message to an approved recipient, obtain a sample approval, and then change the destination before execution. The tool should reject reuse of that approval. Repeat with a changed amount, attachment or account identifier where those fields matter.

This is our proposed engineering rehearsal, not a result reported by the papers above. It checks whether the product’s promise is carried by the execution path. The important observation is the final action taken or refused, with enough evidence to explain it.

Prefer a system you can inspect

Anthropic distinguishes workflows with predefined paths from agents that dynamically direct their own work, and recommends starting with simple approaches. The distinction helps an architecture discussion: a fixed process with one model-assisted decision may be easier to operate than a general agent, while still solving the business problem.

Write down the maximum useful scope of a run: tools available, time budget, steps and stopping conditions. Record the observable tool requests and results, rather than expecting hidden model reasoning to serve as an audit log. Keep sensitive content out of telemetry unless there is a specific, controlled reason to retain it.

Test refusal and recovery as deliberately as success. The agent should stop when it lacks evidence or permission, and pass enough context to the person taking over. A blocked action with a clear explanation is often the correct result.

  • Authorise at the tool boundary.
  • Show the exact action before consequential approval.
  • Constrain each run’s tools, duration and scope.
  • Test hostile inputs and the handoff to a person.

Sources & further reading

  1. OWASP — LLM01:2025 Prompt Injection
  2. Anthropic — Building effective agents
  3. Debenedetti et al. — AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
  4. Yao et al. — ReAct: Synergizing Reasoning and Acting in Language Models

Primary sources inform the technical background. The examples and proposed working methods are MGLO’s own; Business scenarios are identified in the text.