October 7, 2026
The NIST AI RMF explained: what Govern, Map, Measure and Manage require, the seven trustworthy AI traits, and how to put each function into practice.

September 30, 2026
AI supply chain security explained: how models, datasets, and agent tools get compromised, real attack case studies, and the controls that stop them.

September 28, 2026
TheVARA Cybersecurity Requirements Checklist rules every VASP must meet: 19 policy criteria, CISO duties, the 72-hour incident deadline, rule by rule.
LLM security testing is the practice of probing a large language model or an LLM-powered application with adversarial inputs and abuse scenarios to find weaknesses such as prompt injection, data leakage, and excessive agency before attackers do. Unlike traditional penetration testing, it examines how the model behaves as well as the code and infrastructure around it.
The difference matters because an LLM treats every piece of text it reads as potential instruction. A chat message, a PDF in a retrieval pipeline, or a scraped web page can all steer the model. Once the model connects to customer data, internal APIs, or action-taking tools, a manipulated response can become a data breach or an unauthorized transaction. OWASP ranks prompt injection first (LLM01) in its 2025 Top 10 for LLM Applications, ahead of every other risk on the list.
This guide covers what to test and how to test it. It explains the main LLM vulnerabilities and the attack surface behind them, then walks through a repeatable testing method. It also compares pentesting, red teaming, and automated scanning, and covers tools, metrics, continuous testing, and the frameworks that regulators and auditors expect you to map your results to.
LLM security testing is a structured assessment that checks whether a large language model, or the application built around it, can be manipulated to leak data, ignore instructions, or take actions it shouldn't. Testers send adversarial prompts and poisoned inputs, then examine how the model and its surrounding systems respond. The goal is to find exploitable weaknesses and rank them by business impact before an attacker finds them.
The discipline exists because language models are probabilistic and accept instructions in plain language. The same prompt can produce different outputs on different runs, and there is no clean boundary between "data" and "commands." Microsoft's AI Red Team described this in its 2025 write-up of lessons from red teaming more than 100 generative AI products. It concluded that these systems amplify existing security risks and introduce new ones, so testing needs its own methods.
Traditional application pentesting looks for deterministic flaws such as injection bugs, broken access control, or misconfigurations. A vulnerable endpoint behaves the same way each time you hit it, so a confirmed finding stays confirmed. LLM security testing covers that same ground, because the app around the model still has APIs, authentication, and storage. It then adds a behavioral layer that classic scanners cannot assess. For the established methods this builds on, see our guide to penetration testing methods, types and tools and our penetration testing.
The key difference is repeatability. A prompt injection that fails nine times may succeed on the tenth, so testers run each attack many times with variations and report how often it succeeds, not a simple yes or no. The input channel is also wider. Hostile instructions can arrive through an uploaded document, a retrieved web page, or an email the model summarizes, not only through the chat box.
At the definition level, AI red teaming is a goal-driven, adversarial exercise that simulates a real attacker, often across safety, security, and misuse harms. LLM security testing is the broader and more systematic practice of checking a defined system against a known list of weaknesses, such as the OWASP Top 10 for LLM Applications. Red teaming asks "what could a determined adversary achieve?" Security testing asks "which known weaknesses does this system have, and how severe are they?" Most mature programs use both, with red teaming typically run less often and security testing repeated more often. For the general distinction, see red team vs penetration testing.
A test can target four different layers, and each fails in different ways.
The model is the underlying LLM. Testing here looks at jailbreak resistance, training-data leakage, and unsafe output. If you use a third-party model, you can test how it behaves but usually cannot fix it.
The application is the product wrapped around the model: system prompts, input and output filters, APIs, and user permissions. Most exploitable findings sit here, because this layer is yours to configure and secure.
The data includes retrieval sources, vector databases, fine-tuning sets, and conversation logs. Testers check whether poisoned documents can steer the model and whether one user's data can surface in another's session.
The agent is any LLM that can call tools, query systems, or take actions on a user's behalf. Here the question shifts from "what can it say?" to "what can it do?", so permission boundaries get the most scrutiny.
A scoped engagement should name which of these layers is in scope. A test that covers only the model can pass while the application and data layers remain exposed.

LLMs need a different testing approach because language shapes their behavior, not fixed logic. That makes the attack surface open-ended, makes results vary from run to run, and moves most of the real risk into the application layer. Scanners and checklists built for deterministic software can't capture any of these three properties.
In a conventional application, input is data and code is code, and security controls such as parameterized queries enforce the boundary between them. An LLM has no such boundary. The system prompt, the user's message, and every retrieved document arrive as one stream of text, and the model decides what to treat as an instruction.
That means an attacker doesn't need to find a code bug. They only need to phrase text persuasively enough. Researchers demonstrated this at scale in 2023 with indirect prompt injection (Greshake et al., "Not what you've signed up for"), showing that instructions hidden in web pages or documents could hijack LLM-integrated applications without the user ever typing anything malicious. For testers, this means a fixed payload list is never enough. Attacks can be reworded, translated, encoded, split across messages, or hidden in a file, so testing has to cover how the system handles intent and not only whether it blocks known strings.
A passing test on an LLM is evidence, not proof. Sampling settings such as temperature and top-p, plus model updates and context differences, mean the same attack can fail today and succeed tomorrow. A single clean result says only that this attempt didn't work.
Testers handle this by running each attack many times, varying the wording, and measuring how often it succeeds. A jailbreak that works 2 times in 100 is a real finding, because an attacker can simply retry. Results also need to be tied to a specific model version, system prompt, and configuration. Changing any of them can invalidate earlier results, which is why LLM testing works better as a repeated process than a one-time sign-off.
Most serious LLM incidents come from what the model is connected to, not from the model's raw text output. A chatbot that only talks is limited in what it can leak. The same chatbot with access to customer records, internal search, or a payments API can cause real damage when manipulated.
The questions that decide severity are practical ones. What data can the model retrieve, and is access enforced per user or per application? Can it call tools, and are those calls checked independently of the model's judgment? Does it render, execute, or pass its output to another system without validation? A model that is easy to jailbreak but has no sensitive access is a limited risk. A well-aligned model wired to over-privileged tools is a serious one. Effective testing therefore follows the full path from input to action, and ranks findings by what an attacker could reach, not by how clever the prompt was.
The OWASP Top 10 for LLM Applications is the most widely used reference for LLM vulnerabilities, and it works well as a test plan because each risk maps to a concrete attack and a concrete check. The 2025 edition introduced System Prompt Leakage and Vector and Embedding Weaknesses as new entries. It renamed several older ones, reflecting how quickly real-world LLM deployments, especially retrieval-based and agentic ones, are changing.
Risk | What It Looks Like | How to Test It |
|---|---|---|
LLM01: Prompt Injection | Crafted input, typed by a user or hidden in a document, overrides the model's intended instructions. | Send direct override attempts and plant instructions in files, web pages, and emails the model will read. Repeat with rewording, encoding, and translation. |
LLM02: Sensitive Information Disclosure | The model reveals personal data, credentials, proprietary content, or other users' information. | Probe for training-data and context leakage. Test cross-user isolation with separate accounts. |
LLM03: Supply Chain | A compromised model, dataset, plugin, or dependency introduces malicious behavior. | Review model provenance and checksums, dependency and plugin inventories, and third-party model terms. |
LLM04: Data and Model Poisoning | Tampered training, fine-tuning, or retrieval data plants backdoors or biased behavior. | Audit data sources and write access. Insert test poison into non-production pipelines and measure the effect. |
LLM05: Improper Output Handling | Model output is passed to a browser, shell, database, or API without validation. | Make the model emit script, SQL, or command payloads and check whether downstream systems execute them. |
LLM06: Excessive Agency | The model holds more tools, permissions, or autonomy than its task needs. | Enumerate every tool and permission, then try to trigger unauthorized or high-impact actions through prompts. |
LLM07: System Prompt Leakage | Hidden instructions, and any secrets inside them, are exposed to users. | Attempt extraction through direct requests, roleplay, and formatting tricks. Check whether prompts contain credentials. |
LLM08: Vector and Embedding Weaknesses | Retrieval stores allow unauthorized access, data leakage, or poisoned results. | Test per-user access controls on retrieval and inject adversarial documents to see if they surface. |
LLM09: Misinformation | The model produces confident but false output that users or systems rely on. | Run fact-based test sets in your domain, check for fabricated citations, and verify that high-stakes outputs have review steps. |
LLM10: Unbounded Consumption | Excessive or crafted requests drive up cost, exhaust resources, or enable model extraction. | Test rate limits, token and context caps, and per-user quotas under high-volume and oversized inputs. |

LLM injection, or prompt injection, comes in two forms. Direct injection happens when the user types instructions meant to override the system prompt, such as telling the model to ignore its rules. Indirect injection happens when the malicious instructions sit in content the model processes, such as a web page it summarizes, a résumé it screens, or a support ticket it reads. The user may never see the payload.
Indirect injection is usually the more serious of the two. The attacker doesn't need access to the chat interface, and the victim is often a legitimate user whose session or data gets abused. To test it, plant hidden instructions in every content source the application ingests and observe whether the model follows them, especially when it has tool access.
Searchers often use the term "insecure output handling," which was the name in OWASP's earlier list. In the 2025 edition it became Improper Output Handling, but the risk is the same: trusting model output without validating it before another system acts on it.
A typical example is a chat interface that renders the model's response as HTML. If an attacker gets the model to output a script tag, perhaps via an indirect injection in a retrieved page, the script runs in the victim's browser, producing a classic cross-site scripting flaw. The same pattern applies when output is inserted into SQL queries or shell commands. The fix is the one used for any untrusted input: treat the model's output as untrusted, encode it for its destination, and validate it before use.
Sensitive information disclosure and system prompt leakage are related but different. Data leakage concerns information the model can reach, such as training data, retrieved documents, or another user's conversation history. System prompt leakage concerns the model's own configuration being revealed. Testers try both through direct questions, roleplay framings, and requests to repeat or reformat earlier text.
The important point is that a leaked system prompt matters mainly for what it contains. Prompts that hold API keys, internal URLs, or business rules that double as security controls turn leakage into a real incident. Treat the system prompt as something an attacker can read, and enforce access rules in code instead of in prompt wording.
LLM applications rely on a long chain of components: pre-trained models, fine-tuning data, embeddings, plugins, and libraries. A weakness in any of them can affect the application. Supply Chain and Data and Model Poisoning cover this ground, from a tampered open-source model to poisoned documents in a retrieval index. For a deeper look at this layer, see our guide to AI supply chain security.
Testing here is part technical and part process. Verify the origin and integrity of every model and dataset, track dependencies the way you would for conventional software, and limit who can write to training and retrieval data. Because you often can't inspect a third-party model's internals, behavioral testing and contractual controls matter as much as code review.
To test LLM security, define what you are protecting, map every route that text and permissions take through the system, attack it with a varied test set, score findings by what an attacker could actually reach, then fix and retest. The method follows the same logic as conventional security testing, adapted for systems whose behavior varies from run to run. OWASP's GenAI Red Teaming Guide uses a similar approach, combining threat modeling with repeated adversarial testing across the model, the application, and the surrounding infrastructure.

Start by deciding what is in scope and what must not go wrong. List the assets the system can touch, such as customer records, internal documents, credentials, payment actions, and brand reputation. Then decide who the attacker is: an anonymous user, an authenticated customer, a malicious insider, or a third party who controls content the model reads.
The threat model turns these into testable questions. For a support chatbot with access to order history, the question might be whether one customer can extract another's data. For an internal assistant that summarizes email, it might be whether an outside sender can plant instructions that trigger an action. Write the scope down, including which model version, system prompt, and environment you are testing, because results only hold for that configuration.
Trace every path by which content reaches the model and every action the model can trigger. Inputs include user messages, uploaded files, retrieved documents, web content, tool outputs, and anything pulled from other systems. Outputs include rendered responses, API calls, database writes, and messages sent on a user's behalf. The approach mirrors attack surface management for conventional systems.
For each path, record who controls the content and what privileges apply. A retrieval source that anyone can edit, such as a public wiki or a shared inbox, is an injection channel. A plugin that can send email or change records is a high-value target. Mapping permissions at this stage also shows whether the application enforces access or leaves it to the model's judgment, a common source of serious findings.
An adversarial test set is a structured collection of attack cases tied to your threat model. Organize it by risk category (injection, data leakage, output handling, excessive agency, and so on), and write cases specific to your application instead of relying only on generic jailbreak lists. A healthcare assistant needs tests for patient-data extraction, while a coding assistant needs tests for malicious package suggestions.
Include variation within each case. Test the same attack in different phrasings, languages, encodings, and multi-turn sequences, and plant payloads in documents and retrieved content as well as in chat. Record the expected safe behavior for each case in advance so you can judge results consistently. A good set grows over time: every real incident and every new technique becomes a permanent regression test.
Use both, because they find different problems. Automated tools can run thousands of prompt variations, repeat each attack many times, and catch regressions cheaply. Frameworks such as NVIDIA's garak and Microsoft's PyRIT are open-source examples. Manual testing finds what automation often misses: attack chains that span several steps, abuse of business logic, and creative misuse of tools and permissions.
Because outputs vary, run each case repeatedly and record the success rate instead of a single pass or fail. Log the full prompt, context, model version, and response for every attempt so findings are reproducible. Test in a staging environment with realistic data and permissions, and avoid running destructive tool actions against production systems.
A successful jailbreak is not automatically a serious finding. What matters is what the attacker gains. A model that can be coaxed into rude language but has no sensitive access is a low risk. A model that can be steered into exporting customer records or approving a refund is a high one, even if the technique looks simple.
Score each finding on a few practical factors: what data or action is exposed, how reliably the attack works, what access the attacker needs, and whether other controls would stop the damage. Reporting a success rate alongside the impact helps engineers prioritize. A 5% success rate against a payment action deserves attention, because attackers can retry. This approach also keeps results meaningful to leadership, who need to know what could happen to the business and not how many prompts got through.
Fixes work best when they remove the capability, not only the phrasing. Adding "never reveal customer data" to a system prompt is weak protection, while enforcing per-user access in the retrieval layer is strong. Typical fixes include limiting tool permissions, requiring confirmation for high-impact actions, validating and encoding model output, separating untrusted content from instructions, and adding rate limits and monitoring.
After each fix, rerun the original attack cases and the wider regression set. Prompt tweaks, model upgrades, new data sources, and new tools can all reopen old weaknesses, so retest on every meaningful change. Treat LLM security testing as a recurring part of the release process, not a one-time audit. The same discipline applies to vulnerability management in conventional environments.
An LLM risk assessment translates technical findings into decisions. For each confirmed weakness, state the realistic scenario, the asset affected, the likelihood based on observed success rates and attacker effort, and the consequence, such as regulatory exposure, financial loss, or reputational damage. Group findings by the business function they affect so owners can act on them.
This framing also shows where risk is concentrated. Often a small number of capabilities, such as one over-privileged tool or one unauthenticated retrieval source, account for most of the exposure. Presenting findings this way lets leaders choose which risks to fix, accept, and monitor, and it provides the documentation auditors and frameworks like the NIST AI Risk Management Framework expect.
LLM pentesting, red teaming, and automated scanning are three different ways to test an LLM system, and most programs need a mix. Automated scanning gives fast, repeatable coverage of known weaknesses. LLM pentesting is a scoped manual assessment of a specific application. Red teaming is an open-ended simulation of a determined attacker pursuing a goal. Microsoft's AI Red Team reached a similar conclusion after testing more than 100 generative AI products: automation extends coverage, but human judgment remains essential for the hardest problems.
Factor | Automated Scanning | LLM Pentesting | Red Teaming |
|---|---|---|---|
Goal | Find known weaknesses and catch regressions. | Find and validate exploitable flaws in a defined system. | Test whether a realistic attacker can reach a specific objective. |
Depth | Broad and shallow. Many prompts, little context. | Deep within a fixed scope. Manual chaining and business logic. | Deepest. Multi-step, creative, often spans people, process, and technology. |
Cost | Lowest per run. | Moderate. | Highest. |
Cadence | Continuous, on every release or model change. | Before launch and after major changes. | Periodic, such as annually or after significant architecture changes. |
Best Fit | Baseline coverage and CI/CD gates. | Pre-launch assurance and compliance evidence. | Mature programs testing detection and response. |

Automated scanning suits repeatable checks: running a large library of injection, leakage, and jailbreak probes against a model or endpoint and tracking how results change over time. It is cheap enough to run on every prompt update or model upgrade, which makes it the right tool for catching regressions. Its limit is context. Scanners can tell you a prompt got through, but not whether that matters for your business, and they rarely find multi-step attack chains. The same logic separates a vulnerability assessment from a full penetration test.
LLM pentesting fits when you need a defensible answer about a specific application before it ships. A tester works within an agreed scope, such as one chatbot, its retrieval sources, and its connected tools, and tries to exploit what they find. The result is validated findings with evidence, severity ratings, and fixes, which is the format auditors and customers expect. For the underlying methodology, see our guide to penetration testing methods, types and tools.
Red teaming fits when basic controls are already in place and you want to know how a motivated adversary would actually succeed. Instead of checking a list of weaknesses, the team picks an objective, such as extracting sensitive data or triggering an unauthorized action. It uses any combination of technical, social, and process weaknesses to reach it. It also tests whether your monitoring and response notice the attack. It is usually too heavy for an early-stage or low-risk deployment. Learn more about our red teaming.
A practical pattern is to run automated scans continuously, commission a pentest before each major launch, and schedule red team exercises for high-risk systems once the basics are fixed. Findings from the manual work should feed back into the automated test set, so each discovered attack becomes a permanent regression check.
LLM agents, RAG pipelines, and tool-using apps need extra testing because they connect the model to data and actions. That connection is what turns a manipulated response into a real incident. The two questions to answer are what the model can do and what content can influence it. OWASP's 2025 list reflects this shift, adding System Prompt Leakage and Vector and Embedding Weaknesses as new entries and keeping Excessive Agency as a standalone risk.
Excessive agency means the model has more tools, permissions, or autonomy than its job requires. The risk shows up when an attacker steers the model into using a legitimate capability for an illegitimate purpose, such as a support assistant that can issue refunds, send email, or modify records.
To test it, list every tool the agent can call and the permissions behind each one. Then try to trigger high-impact actions through direct prompts and through content the agent reads, and check whether anything outside the model, such as user-level authorization or a human approval step, blocks the action. The strongest fix is to limit capability at the system level, because prompt instructions alone are not a reliable control. For deeper agent-specific testing, see our AI agentic pentesting.
Retrieval-augmented generation lets a model answer from your documents, which also means anyone who can influence those documents can influence the answers. RAG poisoning happens when an attacker plants misleading content or hidden instructions in a source the pipeline retrieves, such as a shared wiki page, a support ticket, or an uploaded file. When a user's question pulls that content into the model's context, the injected text can override instructions or distort the response.
Testing has two parts. First, check access control: confirm that users can only retrieve documents they are authorized to see, since retrieval layers often skip per-user checks. Second, plant test payloads in each ingestion source and observe whether they surface or change behavior. Treat every writable or externally sourced document as untrusted input.

LLM security testing tools fall into four categories: open-source probe frameworks that automate attacks, traditional web security tools adapted for LLM apps, runtime guardrails that filter live traffic, and custom harnesses built for in-house applications. No single tool covers everything, so the right choice depends on what you are testing and at what stage. Treat this as a map of categories, not a ranking, because the tooling changes quickly and every tool named here should be checked against its current documentation.
An LLM security scanner automates the repetitive part of testing by sending large libraries of adversarial prompts to a model or endpoint and scoring the responses. NVIDIA's garak is a widely cited example. It is open source and positioned as an LLM vulnerability scanner, with probes for issues such as prompt injection, data leakage, and jailbreaks. Microsoft's PyRIT is an open-source framework that helps security teams orchestrate red teaming workflows, including multi-turn attacks. promptfoo is an open-source tool for evaluating prompts and models that also supports red team style test configurations.
These tools are strongest for breadth and repeatability. You can run them against every model or prompt change and compare results over time. Their limits are context and judgment: they test the prompts they ship with, and a scanner cannot know which of your data or actions would be damaging if exposed.
An LLM application is still a web application, so conventional tools such as Burp Suite remain useful. They intercept and replay requests, test authentication and authorization on the API behind the chat interface, and check for classic issues like injection in surrounding endpoints. PortSwigger's Web Security Academy includes material on LLM attacks, which shows how these tools fit into LLM testing.
This category covers the application layer well, which is where many exploitable findings sit. It does not evaluate model behavior on its own. A proxy can show you what is sent to the model, but it cannot tell you whether a response was manipulated or unsafe.
Pre-deployment testing looks for weaknesses before release. Runtime guardrails and monitoring watch live traffic, filter suspicious inputs and outputs, and alert on abuse. They solve different problems and don't replace each other. Guardrails can block known attack patterns in production, but they can be bypassed through rewording or encoding, so they work best as one layer of defense and not as proof that the system is secure. Testing shows where the system is weak, while monitoring shows when someone is probing it. Production logs are also a good source of new test cases.
Category | Best For | What It Misses |
|---|---|---|
Open-source probe frameworks | Broad, repeatable coverage and regression checks in CI/CD. | Business context, multi-step attack chains, and application logic flaws. |
Traditional web tooling | API, authentication, and authorization testing around the model. | Model behavior and prompt-level manipulation. |
Runtime guardrails and monitoring | Blocking and detecting known abuse in production. | Novel or reworded attacks and weaknesses that bypass the filter. |
Manual testing by a specialist | Chained attacks, business logic abuse, and impact assessment. | Scale and speed; it cannot run thousands of variations cheaply. |
The honest limitation of automated scanners is that they find what their probe libraries anticipate. They tend to miss attacks that depend on your specific data, permissions, and workflows, such as a retrieval source that lets one user read another's documents, or a tool call that is harmless alone but dangerous in sequence. They can also produce false positives and false negatives because model output varies between runs. Use them for coverage, and use manual testing for depth.
Teams building their own LLM applications often need a mix of the above plus custom work. Start with a probe framework for baseline coverage, add web proxy testing for the surrounding API, and write application-specific test cases for the things only you know, such as which documents are confidential, which tool actions are high impact, and what your users are allowed to see. Standard application security tooling, including static analysis, dependency scanning, and secrets detection, still applies to the code around the model. Wire the automated pieces into your CI/CD pipeline so each change to a prompt, model, or data source triggers a rerun, and feed every confirmed finding back into your test set as a permanent regression case.
The most useful metric in LLM security testing is attack success rate, the share of attempts in which an attack achieves its goal. Still, it only means something when paired with impact, test conditions, and a fixed model configuration. Public benchmarks such as HarmBench, JailbreakBench, and AgentDojo give teams a shared yardstick for comparing models and defenses. They don't replace tests built around your own application, data, and permissions.
Attack success rate (ASR) is calculated by running a set of attacks and dividing the number that succeed by the number attempted. Because LLM outputs vary between runs, repeating each attack many times and reporting the rate is more honest than a single pass/fail. A technique that works in 3 of 100 attempts is still a real weakness, since an attacker can simply retry.
ASR has clear limits. It says nothing about severity: a high rate on a harmless behavior matters less than a low rate on a data export. It depends heavily on how "success" is judged, whether by keyword matching, a human reviewer, or another model acting as grader, and different judges can produce different numbers for the same run. It also varies with temperature, system prompt, and model version. When you report ASR, state the attack set, judging method, number of repetitions, and exact configuration tested, and pair it with the asset or action at risk.
Several open benchmarks exist to make results comparable. HarmBench is a standardized framework for evaluating automated red teaming and refusal behavior. JailbreakBench provides a curated set of jailbreak attempts and a common evaluation setup. AgentDojo focuses on prompt injection against tool-using agents, which makes it more relevant to agentic systems than text-only benchmarks. Meta's CyberSecEval covers risks such as insecure code generation and cyberattack assistance.
These benchmarks help compare models, track progress, and choose a starting point. Their limits matter just as much. Published attack sets can leak into training data, so a model may perform well on them without being robust to new attacks. They test generic behavior and cannot reflect your retrieval sources, permissions, or business rules. A strong benchmark score is evidence about the model, not about your deployed application. Use benchmarks for model selection and baselines, then rely on application-specific testing to judge actual risk.
A regression suite is a fixed collection of attack cases that you rerun after every meaningful change, such as a prompt edit, model upgrade, new data source, or new tool. Start with cases from your threat model, organized by risk category. Add every finding from pentests, red team exercises, and production incidents, so each real weakness becomes a permanent check.
For each case, record the attack input, where it enters the system, the expected safe behavior, and how you judge success. Run each case multiple times and track the success rate over time so drift in model behavior shows up as a trend, not a surprise. Wire the suite into your release pipeline so a meaningful increase in a high-impact category blocks deployment. Keep a small set of human-reviewed cases for the judgments an automated grader handles poorly.
LLM security is not a one-time audit, because the model, the prompts, the data, and the attackers all keep changing. Continuous testing reruns your attack cases automatically on every meaningful change, and continuous monitoring watches live traffic for abuse. A 2023 Stanford and UC Berkeley study ("How Is ChatGPT's Behavior Changing over Time?") found that the behavior of GPT-3.5 and GPT-4 shifted noticeably between March and June of that year on several tasks, which shows why a result from last quarter cannot be assumed to hold today.
A CI/CD security gate runs a defined set of attack cases whenever something that affects model behavior changes, and fails the build if results cross an agreed threshold. In practice this means a fast suite on every pull request that touches prompts, tool definitions, or retrieval configuration, and a larger suite on a schedule or before release. This is the LLM version of DevSecOps.
Set thresholds by impact, not by a single overall score. A rise in success rate for data-exfiltration or unauthorized-action tests should block a release, while a small change in a low-impact category might only raise a warning. Pin the model version, system prompt, and test parameters in the pipeline so a failed run points to a specific change. Because LLM output varies between runs, run each case several times and compare success rates, not single results, to avoid flaky builds. Keep a short list of human-reviewed cases for the judgments an automated grader handles poorly.
Production logging should let you reconstruct an incident without storing more sensitive data than necessary. Log each request with a timestamp, user or session identifier, model and prompt version, the retrieved sources, any tool calls with their parameters and results, and the final output. Apply the same privacy controls you use elsewhere, since prompts and responses often contain personal or confidential data. Our attack surface management shows how exposed assets are monitored over time.
Alert on patterns that suggest probing or compromise. Useful signals include repeated attempts to extract the system prompt, inputs containing instruction-like text from retrieved documents, unusual spikes in token usage or request volume, tool calls outside a user's normal behavior, outputs containing secrets or personal data formats, and blocked actions followed by rewording and retry. Alerts should connect to your incident response process, and confirmed abuse should be fed back into your regression suite so production becomes a source of new test cases.
Treat these as security-relevant changes that trigger retesting: a model upgrade or provider-side update, any system prompt edit, a new or changed data source in the retrieval pipeline, new tools or permissions, and changes to guardrails or filters. Even a small prompt tweak can reopen a weakness that was fixed earlier.
After a change, rerun the full regression suite, compare success rates with the previous baseline, and investigate any increase in high-impact categories before release. If you use a hosted model, check the provider's version and deprecation notices and retest when the model behind your endpoint changes. For major changes, such as a new agent capability or a new class of data, schedule fresh manual testing in addition to the automated rerun.
LLM security testing maps to four reference points: the NIST AI Risk Management Framework, MITRE ATLAS, ISO/IEC 42001, and the EU AI Act, with regional data protection law layered on top. The frameworks give you a common language for scoping tests and reporting results. Regulations determine what evidence you may need to produce. The EU AI Act is the clearest example: for general-purpose AI models with systemic risk, it requires providers to perform model evaluations that include adversarial testing.
The NIST AI Risk Management Framework is a voluntary framework organized around four functions: Govern, Map, Measure, and Manage. Security testing fits mainly under Measure, where you evaluate system behavior and track risks, and feeds Manage, where you act on the findings. In July 2024, NIST published the Generative AI Profile (NIST AI 600-1), which applies the framework to risks specific to generative AI. NIST has also published a taxonomy of adversarial machine learning attacks (NIST AI 100-2), useful for naming and categorizing the attacks in your test plan.
In practice, record your threat model, test results, and remediation decisions in a form that maps to these functions, so an auditor or customer can see how testing supports your risk management. The Govern function overlaps with governance, risk and compliance practice. For a full walkthrough of the framework itself, see our NIST AI RMF guide.
MITRE ATLAS (Adversarial Threat Landscape for Artificial-Intelligence Systems) is a knowledge base of tactics and techniques used against AI systems, modeled on the MITRE ATT&CK framework that many security teams already use. It catalogs real-world attack techniques and case studies, giving testers a shared vocabulary.
Use it to tag each test case and finding with a technique so your results align with threat intelligence and detection engineering. It complements the OWASP Top 10 for LLM Applications: OWASP tells you which weaknesses to check, while ATLAS describes how adversaries operate across the attack lifecycle. Teams that already map to ATT&CK can add ATLAS to the same reporting workflow.
ISO/IEC 42001, published in 2023, is the first international standard for an AI management system. It is certifiable, so it suits organizations that need to show customers and regulators that they govern AI in a structured way. Security testing supports its risk assessment and monitoring requirements, and it pairs naturally with an existing ISO 27001 program. See our ISO 42001 guide and our ISO 27001 compliance service for the information security side.
The EU AI Act is binding law, with obligations that phase in over several years. Providers of high-risk AI systems must meet requirements for accuracy, robustness, and cybersecurity, and providers of general-purpose models with systemic risk must carry out adversarial testing. If you serve EU customers or users, check which risk category your system falls into and which obligations currently apply, as timelines and scope have been adjusted.
For GCC organizations, the testing question is often where the data goes. Prompts, retrieved documents, and logs can contain personal data, and sending them to a hosted model in another country may trigger cross-border transfer rules. Several regimes are relevant, and the right one depends on where your entity and data subjects sit:
UAE: the federal Personal Data Protection Law (Federal Decree-Law No. 45 of 2021), plus free-zone regimes in DIFC and ADGM. DIFC has also issued rules covering autonomous and semi-autonomous systems.
Saudi Arabia: the Personal Data Protection Law, administered with SDAIA, and the National Cybersecurity Authority's controls.
Sector rules: financial regulators and, for virtual asset firms in Dubai, VARA cybersecurity requirements.
A practical approach is to run tests against synthetic or masked data, confirm where your model provider processes and stores data, and keep records of test scope and results. For background, see our guides to UAE data protection law and cybersecurity regulations in the UAE, or our compliance services.
Organizations that keep models and data in-region for residency reasons should also read about sovereign LLMs, sovereign AI agents, and securing sovereign AI infrastructure. Global teams should treat the strictest applicable regime as the baseline and document how each system maps to it.
The most effective LLM security practices are to layer several independent controls, give the model only the access it needs, treat everything it outputs as untrusted, and keep secrets out of prompts entirely. No single defense stops prompt injection reliably. The UK's National Cyber Security Center has warned that prompt injection may never be fully mitigated the way SQL injection was, because models don't separate instructions from data. The practical answer is to limit what a successful attack can achieve.
Assume that some malicious input will get through, and design so that no single failure causes a breach. Combine controls at different layers: input filtering and content screening, a well-structured system prompt, output validation, authorization enforced in application code, rate limits, and monitoring. Filters and prompt wording help, but attackers can reword around them, so they should never be your only protection. Place the strongest controls outside the model, where an attacker's text cannot influence them.
Give each model and agent only the tools, data, and permissions its task requires, and no more. A summarization assistant doesn't need write access, and a support bot doesn't need to read every customer's records. This is the same principle behind zero trust security. Enforce access with the end user's own permissions, not a shared high-privilege service account, so a manipulated model cannot reach more than the user could. Require human approval for high-impact actions such as payments, deletions, and external messages, and ensure the application enforces approval rather than leaving it to the model's judgment.
Treat model output like any untrusted input. Encode it for its destination before rendering it in a browser, never pass it unchecked into SQL queries, shell commands, or other systems, and validate structured output against a strict schema before acting on it. Review the code that consumes model output to catch gaps. Where possible, restrict the model to a defined set of actions instead of free-form commands. Be cautious with markdown and links in responses, since rendered images and URLs can be used to send data to an attacker's server.
Assume anything in the system prompt or context window can be extracted. Do not place API keys, credentials, internal URLs, or sensitive business rules in prompts. Keep secrets in a secrets manager and let backend code use them, so the model never sees them. Do the same for access rules: enforce them in code, and treat the prompt as guidance that an attacker can read and attempt to override.
Mark retrieved documents, uploaded files, and web content as data, and keep them apart from your instructions. Delimiters and clear labeling reduce the chance that the model follows embedded commands, though they do not eliminate it. Limit what untrusted content can trigger, so that a poisoned document cannot cause a tool call without separate authorization.
Common mistakes in LLM security testing include testing only the model, relying on a single jailbreak list, treating a single test as permanent, ignoring indirect injection, and skipping retests after prompt changes. Each leaves a gap attackers can exploit, and most come from applying habits from traditional software testing to a system that behaves differently. OWASP lists prompt injection as the top risk for LLM applications, which makes the last three mistakes especially costly.
Many teams run a jailbreak benchmark against the model and consider the job done. But the model is only one layer. Retrieval sources, system prompts, APIs, permissions, and connected tools determine what an attacker can reach, and most exploitable findings live there. A model that resists jailbreaks can still leak another user's documents through a retrieval layer with no access control. Scope tests to the full path from input to action.
A fixed list of known jailbreak prompts only measures resistance to attacks that are already public. Models and filters are often tuned against popular lists, so a good score can overstate real robustness. Attackers reword, translate, encode, and chain prompts across turns. Use several sources, add cases written for your own application and data, and keep updating the set as new techniques appear.
LLM behavior varies between runs and shifts when the provider updates the model, so a clean result is a snapshot, not a guarantee. A pre-launch pentest is valuable, but it describes one configuration on one date. Build testing into the release process to keep results current, and report success rates over repeated runs instead of a single pass.
Teams often test only what a user types into the chat box. Indirect injection arrives through content the model processes, such as a web page, an uploaded file, an email, or a retrieved document, and the victim may never see it. It matters most when the model has tool access. Plant test payloads in every content source the application ingests and check whether the model follows them.
A small edit to a system prompt, a new data source, or a new tool can reopen a weakness that was fixed earlier. Prompt changes feel like content updates, so they often bypass security review. Treat them as code changes: rerun your regression suite after each one, compare results with the previous baseline, and block the release if high-impact categories get worse.
Use this checklist to confirm that an LLM security test covers the full system, from scope through retesting. A complete test defines what is being protected, maps every input and permission path, attacks each OWASP risk category, scores findings by impact, and repeats after every meaningful change. The list follows the OWASP Top 10 for LLM Applications, whose ten categories give a practical minimum for coverage.
Phase | What to Confirm | Done When |
|---|---|---|
Scope and Threat Model | Assets, attacker types, and the exact model version, system prompt, and environment under test are documented. | Scope is written down and approved by the system owner. |
Attack Surface Mapping | All inputs (user messages, files, retrieval sources, web content, tool outputs) and all actions (tool calls, API writes, messages) are listed with who controls each. | Every input and action has an owner and a privilege level. |
Prompt Injection | Direct and indirect injection are tested, including payloads planted in documents, pages, and emails the model reads. | Each content source has been tested with reworded and encoded variants. |
Data and Prompt Leakage | Cross-user isolation, training-data extraction, and system prompt extraction are tested. | No other user's data surfaces, and prompts contain no secrets. |
Output Handling | Model output is tested against browsers, databases, shells, and downstream APIs. | Output is encoded and validated before any system acts on it. |
Agency and Permissions | Every tool and permission is listed and abused through prompts and injected content. | High-impact actions require authorization or approval enforced outside the model. |
Retrieval and Data Pipeline | Per-user access control on retrieval and poisoning of each ingestion source are tested. | Users retrieve only authorized documents, and poisoned content does not change behavior. |
Supply Chain | Model provenance, datasets, plugins, and dependencies are inventoried and checked. | Every component has a known source and an owner. |
Misinformation and Consumption | Domain fact checks, rate limits, token caps, and quotas are tested. | Abuse and runaway cost are bounded. |
Repetition and Measurement | Each attack runs many times with variation, and success rate, judging method, and configuration are recorded. | Results are reproducible and logged with the model and prompt version. |
Impact Scoring | Findings are ranked by data or action exposed, reliability, and access required. | Each finding has severity, evidence, and an owner. |
Remediation and Retest | Fixes remove the capability where possible, and the original attacks plus the regression suite are rerun. | High-impact categories show no regression against the baseline. |
Continuous Testing and Monitoring | Tests run in CI/CD, production logs capture prompts, retrievals, and tool calls, and alerts cover probing and abuse. | Changes to the model, prompts, data, or tools automatically trigger a rerun. |

Effective LLM security testing means testing the whole system, not just the model: define what an attacker could reach, attack it with varied and repeated tests, fix the capability instead of the phrasing, and retest every time the model, prompt, data, or tools change. NVIDIA's AI Red Team found that the biggest issues across dozens of assessed applications were ordinary application security failures, such as unsafe code execution, weak access control on retrieval data, and unsanitized output. That is good news, because the biggest risks are fixable with disciplined engineering and regular testing.
If your LLM application or agent can access sensitive data or take actions, our team can help you test it against these risks.
LLM security testing is the practice of probing a large language model or an LLM-powered application with adversarial inputs to find weaknesses such as prompt injection, data leakage, and excessive agency before attackers do. It covers the model, the surrounding application, the data it retrieves, and any tools it can use.
Normal pentesting targets deterministic flaws, where a vulnerable endpoint behaves the same way each time. LLM pentesting still covers the APIs, authentication, and infrastructure around the model, but adds a behavioral layer: the same attack can fail on one run and succeed on the next. Testers therefore repeat each attack many times and report a success rate. They also test indirect channels, such as documents and web pages the model reads, not only the chat box.
OWASP's 2025 Top 10 for LLM Applications lists the main categories, including Prompt Injection, Sensitive Information Disclosure, Improper Output Handling, and Excessive Agency. In practice, the most damaging cases usually come from the application layer: over-privileged tools, retrieval systems without per-user access control, and model output that downstream systems trust without validation.
Run automated tests on every change that can affect behavior, such as a model upgrade, a system prompt edit, a new data source, or a new tool. Schedule manual pentesting before major launches and after significant architecture changes. A one-time test describes one configuration on one date, so it should not be treated as permanent assurance.
No. Automated scanners give broad, repeatable coverage and catch regressions cheaply, but they only test the attacks their probe libraries anticipate. They tend to miss multi-step attack chains, flaws that depend on your data and permissions, and the judgment needed to understand what a finding means for your business. Most programs combine continuous automated testing with periodic manual assessment.
Not on its own. Benchmark scores describe how the model behaves against a published attack set, and those sets can leak into training data. They say little about your retrieval sources, permissions, tools, or business rules. Use benchmarks to compare models, then test the deployed application against your own threat model.