AI and application security
Defending RAG applications against prompt injection
By the BackendDrills editorial team · Published and checked October 6, 2026 · 7-minute read
Before you start: Authorization, API contracts and failure-path testing. Find this reading in a study path →
A retrieved document can contain instructions aimed at changing an assistant's behavior. Treat retrieval results as untrusted data and enforce permissions outside the model. A stronger system prompt can help, but cannot replace tenant isolation or authorization at a tool boundary.
An original threat model to work through
Imagine a support assistant for a subscription service. It reads help articles and can propose an account adjustment. A newly uploaded article tells it to ignore the user, fetch another customer's account and apply a credit. The article is relevant to the search query, but relevance does not make its instructions authoritative.
Illustrative trust boundaries:
authenticated user -> authorized retrieval -> untrusted documents
-> model proposal
-> validated application action
Tenant and action permissions are derived from the authenticated session,
not from fields suggested by a document or generated by the model.
Put controls where they can enforce the policy
- Filter retrieval by the caller's authorized document scope before results reach the model. Test that a search cannot retrieve another tenant's private content.
- Label retrieved material as evidence. Keep its text separate from trusted application instructions.
- Give tools narrowly scoped operations. Validate identity, target, arguments and permission in application code on every call.
- Require an appropriate approval step for consequential changes. Reading an article must not grant permission to change an account.
In this example, a generated account ID is only a proposed argument. The adjustment service must derive the allowed tenant from the session and reject any target outside it. A secret required by the backend belongs in the backend, not in retrieved context. Logging the tool decision and its reason helps investigate a rejected proposal without copying private documents into telemetry.
Test the outcome, not only the wording
Build a test document collection containing an ordinary help article and an adversarial variant. Ask the same legitimate question against both. Record which documents were retrieved, whether an unauthorized tool call was proposed, and whether the application permitted it. A polite response is not a security result if an unwanted operation happened before the answer.
| Test case | Expected application result |
|---|---|
| Article requests a cross-tenant read | Retrieval and tool layers deny access |
| Article requests an account credit | No unapproved adjustment is committed |
| Attack is split across retrieved passages | Permissions remain enforced for every operation |
| A legitimate question uses the same tool | The allowed workflow still works |
Contain a detected incident
For the example service, quarantine the suspect document and suspend the affected write capability while investigating. Preserve a minimal audit trail, assess whether any account changed, and restore capability only after the permission tests pass. If credentials were exposed, follow the service's credential-rotation process.
There is no universal prompt that guarantees immunity. State what the test demonstrates: an attacker-controlled article cannot cross the application's tenant or action boundary in the tested workflow. Continue testing as retrieval sources, model behavior and tool capabilities change.
Check the underlying behavior
Original illustrative examples, prepared with AI assistance and checked against the linked primary documentation. No customer incident or vendor endorsement is claimed. Our editorial approach.
Practice the next decision
Try a complete free backend drill: inspect evidence, make three decisions and review the reasoning. No account or card required.
Try a free incident drill →Selected practice from the study paths
- Assistant tools trust a model-supplied owner ID · Full edition drill
- Evaluation scores hide dataset leakage · Full edition drill
Free readings need no account. Full edition drills require verified access; opening a paid link does not expose its answers.