An AI agent that reads a web page, an email or a support ticket will act on any instruction hidden in that text, so a planted sentence can make it leak a customer record or fire off a payment. Preventing that is the job of AI agent security, which in production depends less on the model than on limiting what the agent can do.
What is AI agent security, and why does it matter now?
AI agent security is the set of controls that limit what a model and its tools can be manipulated into doing in production. It matters now because agents no longer only answer, they act: they call APIs, query databases and start workflows, so one malicious instruction can turn useful automation into financial or data damage.
A reliable model is not the same as a safe one. Operating a model well, with latency targets and monitoring, is part of running LLMs in production, but none of it stops an attacker from steering the agent. Security has become the filter buyers apply before they discuss scope or price, and the failures are architectural rather than ordinary code bugs.
What new attack surface does an AI agent create?
An AI agent widens the attack surface in three ways ordinary software does not: it treats plain language as instructions, it acts through tools that hold real permissions, and it pulls in untrusted content from documents, pages and users. Each new capability is also a new way in, so an agent needs more than the checks a normal application receives.
Standard defenses still matter, so the usual rounds of application security testing stay on the list for the code around the model. What they miss is the model itself, where there is no clean line between data and command, and an instruction hidden in content carries the same weight as one from you.
The industry reference for these model-level risks is the OWASP LLM risk list, which names the categories most agents fail on. The table below maps the ones that matter most for a live agent to the control that contains each.
| OWASP risk | What goes wrong | Control that contains it |
|---|---|---|
| Prompt injection | Hidden text overrides your instructions | Treat all input as untrusted, isolate instructions |
| Excessive agency | Agent has more power than the task needs | Least privilege per tool, approval on risky actions |
| Sensitive information disclosure | Secrets surface in a reply | Keep secrets out of the prompt, filter output |
| Improper output handling | Model output runs downstream unchecked | Validate output before it is used or shown |
| System prompt leakage | Your hidden prompt is extracted | Never place a real secret or sole defense in it |
| Data and model poisoning | Bad data enters the knowledge base | Validate sources on ingestion, control access |
| Unbounded consumption | Runaway calls drive cost or downtime | Rate caps, spend limits, quotas per user |
What is prompt injection, and how do you defend against it?
Prompt injection is an attack where instructions hidden in the text an agent reads override the ones you gave it. It comes in two forms: direct, where a user types the override, and indirect, where the payload sits in a document or page the agent processes later. You cannot fully prevent it, so you contain what a success can reach.
Indirect injection is the dangerous one, because the attacker never touches your interface. An agent that summarizes incoming email can be told, by the email itself, to forward a customer record, and it complies unless something stops the action. Defense means treating every retrieved string as hostile, separating trusted instructions from untrusted data, and refusing to let any instruction raise the agent's permissions.
How do you stop an AI agent from doing damage?
You stop an AI agent from doing damage by giving it the least authority the task needs and putting a person in front of anything irreversible. Most harm traces back to excessive agency: an agent handed broad credentials and no approval step. Scope every tool, cap what each action reaches, and confirm operations that move money or delete data.
Least privilege means each tool gets its own scoped credential rather than one powerful key, and the agent sees only the tools a task requires. A read action should never ride the same permission as a write, and a destructive action sits behind an approval the model cannot grant itself. These limits turn a successful injection from a breach into a blocked attempt.
Add hard ceilings for the same reason. A cap on spend, a rate limit per user and a quota on actions per session keep a manipulated or looping agent from draining budget or hammering a downstream service.
How do you keep an LLM app from leaking sensitive data?
An LLM app leaks data when it repeats secrets from its context, its system prompt or its sources into a reply an attacker can reach. Defend it by keeping secrets out of the prompt entirely, filtering model output before it is shown or acted on, and never trusting the model to hold something back because the instructions asked it to.
Two failure modes deserve attention. System prompt leakage lets an attacker extract the hidden rules you depend on, so those rules must never be your only defense or hold a real secret. Improper output handling is subtler: when model output flows into a query, a page or a command unchecked, the model becomes a path for classic injection, so escape it like any untrusted input.
How do you secure a RAG pipeline against poisoning?
A RAG pipeline, short for retrieval-augmented generation, feeds an agent outside documents at query time, and you secure it by controlling what enters the knowledge base and who can reach each slice. Poisoned or mislabeled documents become instructions the agent later trusts, so validate every source on ingestion and enforce access control at retrieval, not only in the interface.
The risk is easy to underestimate because retrieval feels like a read. In practice the retrieved passage is injected straight into the model context, so whoever places a file in your index can plant instructions that fire for every future query. Isolate stored vectors per tenant so one customer data cannot surface in another answer, and record where each passage came from.
How do you test an AI agent before it goes live?
You test an AI agent by attacking it on purpose before anyone else does. Red-teaming an agent means running injection attempts, jailbreaks and permission-escalation probes against the full system, tools included, not the model alone. The goal is to confirm that even when the model is fooled, the controls hold and nothing irreversible happens without a person in the loop.
Make the test repeatable rather than a one-off. Build a suite of adversarial cases, run it on every change to prompts, tools or models, and block a release when a contained attack starts getting through. The checklist below is the minimum to clear before an agent handles real users and real data.
Pre-launch AI agent security checklist
- Every tool uses its own scoped credential, never one shared key.
- Irreversible actions require human approval the model cannot bypass.
- All retrieved and user content is treated as untrusted by default.
- Model output is validated before it hits a query, page or command.
- Spend, rate and action limits are enforced per user and session.
- An adversarial test suite runs on every prompt, tool or model change.
Who should own AI agent security in your build?
The awkward part of staffing this is that the people who build an agent fastest rarely think like an attacker, and judging whether a contractor does is hard when you cannot run the attacks yourself.
Vetted Outsource matches you with engineers whose security work on live LLM systems has been checked against the risks above. Describe what your agent may touch, and we connect you with people who have contained exactly that.











