Build layered defenses around user input, retrieved content, tools and agent workflows so malicious instructions cannot quietly turn model flexibility into business authority.
A conventional application can often separate executable code from ordinary user data. Language models operate differently. They process instructions, retrieved content, user messages and system context through the same general language interface.
That creates a new class of security problem: text that was intended to be treated as data can attempt to influence the model as if it were an instruction.
A malicious user might type an override directly. A document might contain hidden or visible instructions. A website could include text intended for an AI crawler or browsing agent. An email or support ticket could try to manipulate an automated workflow. A poisoned knowledge source could influence future users long after the original attacker is gone.
The defensive objective is therefore not “make the model perfectly resistant to manipulation.” The stronger objective is: design the surrounding system so manipulated model behaviour cannot cross critical trust boundaries.
No single detector or system prompt is enough. The production architecture should reduce exposure, isolate untrusted context, constrain tools and validate every sensitive consequence.
Give the model only the context, tools and data needed for the active task instead of broad standing access.
Distinguish system instructions, trusted policy, user content and retrieved external data in the application architecture.
Expose narrow, schema-bound actions rather than raw administrative APIs or unrestricted execution.
Check user, role, tenant, record and action permissions outside the model before execution.
Apply business rules, required fields, evidence checks and human confirmation where risk is high.
Log attempted boundary crossings, suspicious retrieval patterns, tool misuse and policy failures.
Prompt injection can originate from the user, but agentic and retrieval-enabled systems also consume content from external sources that attackers may be able to influence.
| Attack Type | How It Enters | Why It Matters |
|---|---|---|
| Direct injection | The user explicitly asks the model to ignore instructions, reveal protected context or perform an unauthorized task. | The attack is visible, but may still succeed if model behaviour is treated as an authorization decision. |
| Indirect injection | Malicious instructions are embedded in a document, email, web page, support ticket or other retrieved source. | The system may treat the content as trusted evidence even though it was authored by an attacker. |
| Stored injection | Malicious content is added to a persistent knowledge base, CRM note, document repository or other indexed source. | One compromise can influence many future users and automated workflows. |
| Cross-tool injection | Output from one tool or system includes instructions that influence how the model uses another tool. | Agent orchestration can unintentionally propagate attacker-controlled instructions across trust boundaries. |
| Multi-turn injection | The attacker gradually changes context or persuades the model across a longer conversation. | The malicious instruction may not resemble a single obvious jailbreak attempt. |
The severity of an attack depends on the authority connected to the model. A read-only assistant and a production agent with tool access do not have the same blast radius.
Attempt to make the model ignore system rules, application policy or task boundaries.
Coax the model into revealing hidden context, private records, system prompts or data belonging to another user.
Manipulate an agent into calling an API, modifying a record, sending a message or taking another sensitive action.
Reframe a request so the model attempts behaviour outside the intended business workflow.
Insert content into a persistent source so future retrieval repeatedly produces attacker-controlled instructions.
Cause the model to misclassify, misroute, skip steps or make incorrect recommendations inside an automated process.
The goal is not to reject normal language. It is to know which content is allowed to influence interpretation and which content is never allowed to grant authority.
Assume users may explicitly attempt to override policy, reveal hidden instructions or manipulate tool selection.
Treat uploaded and retrieved files as data even when they contain language that resembles instructions to the model.
Separate the business content of a message from any instructions that attempt to control the automation processing it.
Assume external websites may deliberately include text designed to manipulate browsing or research agents.
Do not assume API or connector responses are trusted simply because they arrived through a technical integration.
Prevent manipulated past content from silently becoming durable authority in future sessions or workflows.
System instructions are useful for defining task behavior, but they remain part of the model interaction. Critical access and execution boundaries should be implemented outside the prompt.
RAG systems need security controls around source trust, content ingestion, permissions and how retrieved material is presented to the model.
Distinguish approved internal sources, lower-trust external sources and user-generated content before retrieval.
Ensure the attacker cannot use injection to retrieve documents the authenticated user should never see.
Restrict who can add or modify high-authority knowledge and review changes to sensitive sources.
Represent retrieved passages as source material rather than blending them indistinguishably with trusted application instructions.
Keep track of which document, page or record contributed information to a model response or agent decision.
Constrain the number, type and sensitivity of sources brought into context for each workflow.
The safest agent architecture assumes the model may occasionally produce the wrong action request and makes the tool layer responsible for rejecting anything outside policy.
Expose specific business operations rather than broad API surfaces or administrative endpoints.
Require typed parameters, permitted values and required fields before the tool handler accepts a request.
Check user, role, tenant and action permissions at execution time rather than trusting the model's reasoning.
Apply deterministic logic for eligibility, limits, allowed transitions and other sensitive workflow requirements.
Require a person to confirm actions that are irreversible, unusually sensitive or outside normal thresholds.
Limit repeated calls, unexpected action sequences and abnormal tool invocation patterns.
The strongest protection against prompt-based data theft is enforcing data boundaries before context reaches the model.
Include only the records and fields needed for the active task rather than entire customer or enterprise datasets.
Prevent one tenant's content from becoming retrievable in another tenant's session or agent context.
Remove or mask sensitive attributes that the current workflow does not require.
Inspect responses for prohibited data classes or patterns before information is returned externally.
Keep credentials, keys and backend secrets out of model-visible context wherever possible.
Ensure debug traces and observability tools do not become alternate paths for sensitive data exposure.
Heuristics, classifiers and model-based detection can identify suspicious patterns, but false negatives are inevitable. A secure design still needs hard boundaries around data and actions.
Flag obvious jailbreak language, instruction overrides, prompt-leakage attempts and repeated policy probing.
Identify when a request appears unrelated to the allowed business workflow or attempts to change the operating role of the model.
Assign different trust levels to user-generated content, external web content and approved enterprise sources.
Flag unusual tool selection, parameter values, repeated retries or action sequences inconsistent with normal workflow behaviour.
Detect responses containing sensitive data, hidden-instruction references or other patterns the application should not expose.
Send suspicious requests into reduced-capability paths or human review instead of normal high-authority execution.
A useful security test asks whether an attacker can cross a real boundary: access protected data, trigger an unauthorized tool, alter a workflow or persist malicious instructions.
Attempt explicit overrides, hidden prompt requests, policy bypasses and role manipulation through normal user input.
Embed malicious instructions in documents, emails and retrieved content and verify that critical controls remain intact.
Attempt to retrieve or infer information belonging to another user, customer, tenant or department.
Try unauthorized parameters, out-of-sequence operations, duplicate actions and hidden requests to sensitive tools.
Attempt to plant malicious instructions in knowledge, memory, CRM notes or other content that may influence future sessions.
Confirm that model, provider or tool failures do not route the workflow into a less secure path.
Security controls should match the application's actual authority rather than applying the same generic guardrail stack to every use case.
| Application | Primary Injection Risk | Control Priority |
|---|---|---|
| Internal assistant | Unauthorized disclosure across departments or sensitive internal knowledge sources. | Identity, permission-aware retrieval, tenant or department boundaries and output controls. |
| Customer chatbot | Attempts to reveal hidden instructions, internal data or unsupported account information. | Public-vs-private source separation, authenticated record access and response filtering. |
| RAG application | Indirect instructions embedded in retrieved content or poisoned knowledge sources. | Source trust, ingestion control, provenance, context segmentation and least privilege. |
| AI agent | Manipulation of tool selection, parameters or multi-step workflow logic. | Tool-level authorization, deterministic validation, human approval and action logging. |
| Browser/research agent | Malicious instructions embedded in websites or online documents. | Untrusted-source treatment, isolated browsing, restricted tools and output validation. |
When an attack crosses a boundary or exposes an unsafe path, teams need enough observability and control to contain the issue quickly.
The security problem spans model behaviour, source governance, identity, tools and operational workflows, so responsibility cannot sit with prompt engineering alone.
Own workflow scope, prompts, tool contracts, user experience and application-level failure handling.
Own threat modelling, access policy, secrets, monitoring, testing and incident response standards.
Control who can add, modify and approve sources that influence retrieval-enabled systems.
Own APIs, tool permissions, validation and downstream system boundaries.
Own model selection, version changes, provider behaviour and evaluation of safety-relevant model changes.
Define which actions require human approval and how much automation risk the workflow can tolerate.
Peak Demand combines model guidance with deterministic middleware, permission-aware retrieval, narrow tools, validation and production monitoring so one manipulated response cannot become unrestricted business authority.
Map where untrusted language enters the system and what data or authority could be reached from each path.
Separate trusted instructions from user input and retrieved evidence, with explicit context boundaries.
Implement source trust, provenance, permissions, ingestion controls and indirect-injection defenses.
Build schema-bound tools, deterministic authorization, validation and human approval for sensitive actions.
Create attack suites covering direct, indirect, persistent, cross-tenant and tool-manipulation scenarios.
Instrument injection attempts, blocked actions, unusual tool use, retrieval anomalies and security regressions.
These related pages cover the identity, governance, retrieval and integration controls that limit the impact of malicious instructions.
Prompt injection is an attempt to manipulate an AI model through language that conflicts with the intended application instructions or workflow. The malicious instruction can come directly from a user or indirectly through content the AI retrieves.
Indirect prompt injection occurs when malicious instructions are embedded in external content such as documents, websites, email, tickets or knowledge sources that the AI later reads and treats as context.
A strong system prompt can improve behaviour, but it is not an enforceable security boundary. Sensitive data access, authorization and high-consequence actions should be controlled by software outside the model.
Use narrow tool contracts, explicit schemas, deterministic authorization, business-rule validation, workflow state checks and human approval for sensitive actions.
Classify source trust, control ingestion, enforce permissions before retrieval, preserve provenance, treat retrieved content as untrusted evidence and limit what downstream tools can do with the result.
It can create data-exposure risk if the model already has access to protected information. Strong systems enforce permissions before data enters context and minimize what information is exposed to each session.
Detection can be useful as one layer, but it should not be the only defense. Secure architecture assumes some attacks will evade detection and contains the impact through least privilege and deterministic controls.
Test direct overrides, indirect instructions in documents and web content, cross-tenant retrieval, data-exfiltration attempts, unauthorized tool requests, persistence and fallback behavior.
No. It can affect RAG systems, browsing agents, research agents, customer-service systems, internal assistants and any AI workflow that consumes untrusted language.
The risk becomes more consequential because the model can influence real actions. Agent systems therefore need stronger authorization, validation, tool scoping, monitoring and human-approval boundaries.
Yes. Peak Demand can assess input paths, retrieval architecture, tool permissions, model context, validation and production monitoring, then implement layered controls around the existing workflow.
Peak Demand can map prompt injection paths across your users, documents, retrieval layer and agents, then implement the permission, validation and tool controls that contain the impact.