Prompt Injection in AI Agents: Controls for 2026
Prompt injection places an instruction inside content an AI system reads. With tool access, the effect can reach email, files or business records. Learn the controls that limit permissions, approval and data flow.
Content that an agent reads can contain instructions aimed at changing its behaviour. Controls must limit what the system can see, call and change.
Syntalith
Prompt injection places an instruction inside content an AI system reads. The content can be an email, document, web page, customer ticket, memory item or tool result. When the system only prepares a reply, the result may be a wrong answer. When an agent can call tools, the same weakness can reach email, files, payments or business records.
OWASP classifies prompt injection as LLM01:2025 and publishes a separate Top 10 for Agentic Applications for 2026. Treat the issue as a system design and operating-control question.
What the attack changes
There are two common paths:
- Direct injection: a person enters an instruction in a request, such as an attempt to override the system's operating rules.
- Indirect injection: content from another source contains an instruction that the agent reads while carrying out a task.
The model receives instructions, data and tool results in a shared working input. It can assign the wrong authority to a piece of content. A company therefore needs controls outside the model's wording.
How tool access increases impact
Four properties determine the possible damage:
| Property | Risk question |
|---|---|
| Content sources | Which emails, pages, files, tickets or records can influence a decision? |
| Memory | Can an instruction persist in a stored note, retrieval index or task history? |
| Tools | Can the agent read, write, send, pay, delete or execute? |
| Authority | Which person or rule approves an action before it reaches a live system? |
OWASP describes excessive agency as a risk when a system receives more functionality, permission or autonomy than the task needs. A useful design starts with the worst action the agent could take after a manipulated instruction and removes every permission that does not support the task.
What the documented cases show
Public security material gives two different lessons:
| Source | Documented point | Design lesson |
|---|---|---|
| OWASP prompt injection guidance | No foolproof prevention method is currently known; indirect instructions can enter through external content | Use layered controls and adversarial tests |
| AWS-2025-015 | AWS described malicious code inserted into an extension repository and distributed in version 1.84.0; a syntax error prevented execution | Protect code supply chains, releases and permissions as well as model input |
The cases point to a practical rule: the model is one component in a larger authority chain. Code review, account separation, release controls and action checks all matter.
Controls for a connected system
Build the controls around the task:
- Label content by trust. Keep instructions written by the system owner separate from text imported from customers, pages, documents or tools.
- Limit permissions. Give the agent the smallest read and write access needed for the chosen process. Separate draft creation from sending and reading from changing.
- Check proposed actions. Compare the action with the original request, allowed destination, record owner and amount before invoking a tool.
- Require approval. A person approves messages to external recipients, payments, deletion, production changes and other irreversible actions.
- Isolate sensitive work. Use separate accounts, environments and secrets for unrelated tasks. Keep a low-risk research task away from production credentials.
- Keep an audit record. Store the request, source, proposed action, approval, tool result, version and time so the owner can reconstruct an incident.
- Test hostile content. Include poisoned pages, attachments, customer fields, tool descriptions and stored notes in acceptance tests.
- Prepare recovery. Define how to revoke access, stop a run, restore a record and inform the owner when a control fails.
Model-based filters can support these controls. They cannot replace permissions, deterministic checks and an approval point.
Questions for a vendor
Ask for concrete answers before connecting company systems:
- Which sources are treated as untrusted content?
- Which records and tools can the agent reach on each task?
- What prevents an imported instruction from changing the allowed action?
- Which actions require a named person's approval?
- What exactly is recorded, for how long and where?
- How can the company stop a run and revoke credentials?
- How are new tools, prompts, models and source connectors tested?
- Can the supplier show a failure case and the recovery procedure?
An answer such as “the model is safe” leaves the authority chain undefined. The supplier should connect each control to a permission, record or decision point.
Article 50 is a separate duty
Article 50 of the EU AI Act covers transparency duties for certain AI systems. The European Commission's guidance on those duties explains the requirements and application context. The relevant transparency provisions apply from 2 August 2026.
Transparency tells people when they interact with certain AI systems. Prompt-injection controls protect system authority and data. A deployment plan should record both tracks, assess the company's exact role and obtain qualified review for its obligations.
Plan the first security test
Choose one proposed action and map its full path:
untrusted content -> model input -> proposed action -> check -> approval -> tool -> record
Prepare ordinary examples and hostile content. Check that the agent:
- ignores an instruction embedded in a source document,
- keeps the action within the requested destination and record,
- stops before an irreversible step,
- records the source, decision and approval,
- allows the owner to revoke access and recover.
Use the results to set the first production scope. A free process scan can map one proposed action, its permissions, the approval point and the records required. The AI agent implementation service can then turn that map into a scoped build. Include tests of permission decisions, required approvals and blocked target calls in the acceptance criteria.
Related articles
Frequently asked questions
- What is prompt injection?
- It is an attack in which content read by a model contains an instruction that changes the system's behaviour. The content can arrive through a message, document, web page, customer request, memory or tool result.
- Why is prompt injection serious for an AI agent?
- An agent can call tools after producing a decision. If the system has broad access, a manipulated instruction may lead to a message, record change, disclosure or other action. Permissions and approval determine the possible impact.
- Can a filter eliminate prompt injection?
- No single filter removes the risk. Use layers: separate untrusted content from privileged actions, grant only required permissions, check proposed actions, require approval for irreversible work and keep an audit record.
- What should a vendor explain about prompt injection?
- Ask which content is untrusted, which tools and records the agent can reach, how actions are checked, which actions require approval, what is recorded and how an incident can be reconstructed.
- Does Article 50 of the AI Act cover prompt injection?
- Article 50 concerns transparency for certain AI systems and has its own application date. Prompt-injection controls address system security and action authority. A company should assess both duties for its use case and obtain qualified review where required.
Free process scan
Start with a free process scan.
- A 30-minute call with the engineer who would lead the work.
- A review of the processes that cost you the most time and money.
- A written summary: a possible direction, missing information and the next step.
The scan chooses one process to assess, and within 2 business days you receive a recommendation, including when a simpler route is the better fit.
€0
30 minutes · written takeaway within 2 business days
Times are shown in your own time zone. We work with clients across time zones.
Describe the process in the form