How AAISM Addresses LLM Security: Prompt Injection, Jailbreaks, and Defense Strategies Explained

  •   min.
  • Updated on: July 19, 2026

    • Expert review
    • Home
    • /
    • Resources
    • /
    • How AAISM Addresses LLM Security: Prompt Injection, Jailbreaks, and Defense Strategies Explained

    Prompt injection has been ranked by OWASP as the number one AI security risk for two consecutive years. Attack success rates in production environments run between 50 and 84 percent, depending on system configuration. Real CVEs with scores above 9.0 have been documented in Microsoft 365 Copilot, GitHub Copilot, AWS Q Developer, and Cursor IDE. These are not theoretical vulnerabilities in research environments. They are active exploits in the same enterprise LLM deployments that security leaders are responsible for governing right now.

    What makes this particularly important for AAISM is the nature of the problem. Prompt injection does not work like any attack addressed by traditional security frameworks. It operates at the semantic layer: an attacker manipulates natural language inputs to override an LLM's instructions, bypass its controls, or cause it to act against the interests of the organization that deployed it. No firewall stops it. No network segmentation blocks it. The defenses are architectural and governance-level, falling squarely in the domain of the security leader, not the developer.

    AAISM Domain 2 (AI Risk Management) and Domain 3 (AI Technologies and Controls) together address how security leaders should identify, assess, and respond to LLM-specific threats, including prompt injection and jailbreaks. The exam does not ask you to build detection systems or write prompt sanitization code. It asks whether you understand the governance decisions that determine whether your organization's LLM deployments are secure at a program level.

    This guide explains what these attacks are, how they work in practice, how the real-world cases connect to what AAISM tests, and what the governance defenses look like at the level the exam rewards.

    What Prompt Injection Actually Is

    Prompt injection is an attack that manipulates how an LLM processes and responds to inputs by inserting instructions that override the system's intended behavior. Unlike traditional injection attacks that target code or database queries, prompt injection targets the language model's instruction-following capability itself.

    The fundamental reason prompt injection exists is architectural: LLMs cannot reliably distinguish between trusted instructions in a system prompt and untrusted content in a user message or external data source. Both are processed as natural language tokens in the same context window. An attacker who understands this can craft inputs that compete with the system prompt for instruction authority.

    Two forms appear consistently in AAISM exam content:

    • Direct prompt injection occurs when a user submits input that directly overrides or bypasses the system prompt. The simplest example is a user message that tells the model to ignore its previous instructions and behave differently. More sophisticated versions use persona injection, roleplay scenarios, or encoded instructions that the model interprets as authoritative directives.
    • Indirect prompt injection is more dangerous in enterprise deployments because the attacker does not interact with the system directly. Instead, malicious instructions are embedded in external content that the LLM ingests during operation: a document it summarizes, a web page it retrieves, an email it processes, or data returned from an API call. When the model reads that content, it processes the embedded instructions as if they were legitimate directives.

    Jailbreaks vs Prompt Injection: The Distinction AAISM Tests

    The exam draws a distinction between prompt injection and jailbreaking that matters for governance responses, and conflating the two leads to control gaps.

    • Prompt injection targets the application layer. The attacker manipulates what the LLM does in a specific deployment context: overriding system instructions, accessing data it should not have access to, or taking actions outside its authorized scope. The attack succeeds because the deployment architecture does not adequately separate trusted instructions from untrusted inputs.
    • Jailbreaking targets the model's safety alignment. The attacker crafts inputs that cause the model to bypass the safety behaviors built into it during training, producing outputs it was designed to refuse. Jailbreaking exploits the model itself rather than its deployment context.

    This distinction matters for governance because the defenses are different. Prompt injection defenses focus on input validation, instruction hierarchy enforcement, output monitoring, and least-privilege architecture for LLM agents. Jailbreaking defenses focus on model selection, safety evaluation during procurement, ongoing monitoring for alignment drift, and escalation protocols when the model produces policy-violating outputs.

    A governance program that only addresses one of these has a visible gap. AAISM Domain 2 tests whether you can identify which attack type applies in a given scenario and which risk treatment approach is appropriate.

    Looking for some exam prep guidance and mentoring?


    Learn about our personal mentoring

    Image of Lou Hablas mentor - Destination Certification

    What Happens When LLM Security Fails: Real Cases

    The governance implications of prompt injection become concrete when you look at what actual production exploits have achieved.

    The EchoLeak vulnerability (CVE-2025-32711, CVSS 9.3) demonstrated zero-click prompt injection in Microsoft 365 Copilot. An attacker sent a crafted email containing hidden instructions. When the recipient asked Copilot to summarize their inbox, the AI silently exfiltrated sensitive documents to an external server, without any interaction from the target beyond asking a routine question. The attack chained multiple bypasses: evading Microsoft's prompt injection classifier, exploiting auto-fetched images, and abusing a Microsoft Teams proxy to move data out of the environment. Three steps, no magic, no malware.

    As The Register reported in August 2025, a similar indirect prompt injection vulnerability in AWS Q Developer allowed attackers to exfiltrate API keys from developer environment files. If a developer interacted with a malicious file and queried the LLM, Q Developer would read the file, invoke a bash command, and dump the contents, including API credentials, without the developer's knowledge or consent. AWS quietly patched the vulnerability after it was demonstrated.

    Both cases share the same governance failure pattern: LLM agents were granted access to sensitive data and action capabilities without adequate controls separating trusted instructions from untrusted content. The architectural decisions that enabled these attacks were governance decisions, not engineering mistakes.

    For a structured framework on how AI threats like these are identified and tracked in practice, the free AI Threat Hunting Playbook from Destination Certification maps the detection and response thinking that connects directly to what Domain 3 tests. Understanding that framework before studying Domain 3 formally makes the exam content significantly more concrete.

    Certification in 3 Days 


    Study everything you need to know for the AAISM exam in a 3-day bootcamp!

    How AAISM Maps These Threats Across Its Domains

    AAISM does not organize its exam content around attack types. It organizes it around governance functions. Understanding where prompt injection and jailbreaks appear across the domain structure changes how you prepare.

    Domain 2: AI Risk Management (31%)

    In Domain 2: AI Risk Management, prompt injection and jailbreaks appear as threat vectors in AI risk assessment and treatment scenarios. The exam tests whether you can identify these as risks during the threat modeling phase of an AI deployment, evaluate their likelihood and impact based on the deployment architecture, select appropriate risk treatment options (mitigation, acceptance, transfer, or avoidance), and assess third-party and supply chain risk when the LLM component comes from an external vendor.

    Domain 2 questions are governance decisions, not technical implementations. A typical question presents an organization deploying an LLM with access to internal documents and asks which risk treatment approach best addresses indirect prompt injection exposure. The right answer is the one that reflects a sound program-level decision: architectural mitigation that reduces the attack surface, not a technical control that addresses a single vector without addressing the underlying access model.

    Domain 3: AI Technologies and Controls (38%)

    In Domain 3: AI Technologies and Controls, the same attacks appear in the context of control design and evaluation. The exam tests whether you understand what effective controls for LLM security look like, how to evaluate whether those controls are working, and what governance obligations apply when they fail.

    The set of controls the exam expects you to know includes: instruction hierarchy enforcement (hard system rules take precedence over user inputs, enforced at the orchestration layer rather than by relying on the model's own judgment), input validation and output filtering before tool calls or responses are delivered, least-privilege architecture for LLM agents (limiting what data they can access and what actions they can take), provenance-based access control (tagging and isolating external content so it receives stricter scrutiny than trusted instructions), and continuous adversarial testing to detect alignment drift and new injection vectors as they emerge.

    Domain 3 questions test governance judgment about these controls, not configuration expertise. The exam asks whether a given control addresses the right layer of the problem, whether it creates false confidence while leaving a gap, or whether it is appropriate given the deployment context.

    Once you understand the governance mental model that AAISM uses to frame these controls, working through exam questions becomes significantly more efficient. The free Neutral Playbook from Destination Certification builds exactly that mental model: how to approach security management decisions in a way that maps to how AAISM frames its questions across both Domain 2 and Domain 3.

    The Defense Strategies AAISM Expects You to Know

    The exam does not expect you to know how to implement these defenses at the engineering level. It expects you to understand what they address, what gaps they leave, and what governance decisions determine whether they are adequate.

    1. Instruction hierarchy: The orchestration layer, not the model itself, must enforce the rule that system prompt instructions take precedence over user inputs and external content. Relying on the model to resist prompt injection is not a viable governance approach because no frontier model has demonstrated reliable resistance.
    2. Input validation and output filtering: Pre-inference scanning for known injection patterns reduces the attack surface but cannot fully prevent injection. Output filtering before tool calls or responses is delivered provides a second control layer. Neither is sufficient alone.
    3. Least privilege for LLM agents: An LLM agent that can access sensitive data, make API calls, execute code, and take actions on behalf of users has a much larger blast radius when compromised than one with narrowly scoped permissions. Least privilege architecture is a governance decision made at deployment, not a technical fix applied after compromise.
    4. Continuous adversarial testing: Static controls become less effective as attack techniques evolve. A governance program that does not include ongoing red-teaming and adversarial evaluation of LLM deployments will drift out of alignment with the actual threat environment over time.
    5. Incident response planning for AI-specific failures: When a prompt injection attack succeeds, the incident response process looks different from a traditional breach. The evidence is in model behavior and outputs rather than network logs, the scope of impact may be difficult to determine from standard forensic tools, and the remediation may require retraining or restricting the model rather than patching a vulnerability.

    For a broader view of how these topics connect to the full AAISM certification context, the ISACA AAISM guide maps all three domains and how they work together as a certification framework.

    What This Means for Your AAISM Preparation

    LLM security topics appear across both Domain 2 and Domain 3, which together account for 69% of the exam weight. Getting these domains right is not optional.

    The preparation mistake most people make in this area is studying the technical details of prompt injection attacks rather than the governance reasoning that the exam rewards. Knowing how EchoLeak worked at a technical level does not help you answer a Domain 2 question about risk treatment for indirect prompt injection. Knowing which governance decision would have prevented the architectural conditions that made EchoLeak possible does.

    The mental shift is from "what is this attack?" to "what governance decision addresses the risk this attack represents." Every LLM security question on AAISM asks the second question. The organizations that have suffered production prompt injection exploits did not fail because their developers did not know what prompt injection was. They failed because their governance decisions created deployment architectures where the attack was possible.

    For a complete picture of how all three AAISM domains connect and what the full certification requires, the AAISM certification guide maps the full scope before you finalize your preparation approach.

    Frequently Asked Questions 

    Is prompt injection tested directly on the AAISM exam?

    Yes. Prompt injection appears in both Domain 2 and Domain 3 as a key AI threat vector. Domain 2 tests your ability to identify and assess it during threat modeling and risk treatment scenarios. Domain 3 tests your understanding of the governance controls that address it. The exam does not test technical implementation. It tests governance-level decision-making about how to manage the risk in a deployment context.

    What is the difference between direct and indirect prompt injection?

    Direct prompt injection involves a user submitting input that overrides or bypasses the LLM's system instructions. Indirect prompt injection involves malicious instructions embedded in external content that the LLM ingests during operation, such as documents, web pages, emails, or API responses. Indirect injection is more dangerous in enterprise deployments because the attacker does not need direct access to the system, and the attack can be triggered by normal user behavior.

    Does AAISM treat jailbreaks as a separate topic from prompt injection?

    Yes, and the distinction matters for exam answers. Prompt injection targets the application layer and is addressed through architectural and deployment controls. Jailbreaking targets the model's safety alignment and is addressed through model selection, safety evaluation during procurement, and monitoring for alignment drift. AAISM tests whether you can identify which attack type is relevant in a given scenario and which governance response applies.

    What governance controls does AAISM expect you to know for LLM security?

    The key controls are instruction hierarchy enforcement at the orchestration layer, input validation and output filtering, least-privilege architecture for LLM agents, provenance-based access control for external content, and continuous adversarial testing. The exam tests governance judgment about these controls: whether a given control addresses the right layer of the problem and whether it is appropriate for the deployment context described in the scenario.

    How does Domain 3's weight affect how much time to spend on LLM security topics?

    Domain 3 carries 38% of the exam weight, the highest of the three domains. LLM security topics, including prompt injection, adversarial attacks, and AI-specific controls, are a significant portion of that domain. Combined with the risk management framing in Domain 2 (31%), these topics appear across 69% of the exam. Preparing thoroughly for LLM security governance is not optional for a passing score.

    Prompt Injection Is Already in Your Environment. AAISM Is How You Prove You Can Govern the Response

    You now understand what prompt injection and jailbreaks look like in production, why traditional defenses do not address them, and how AAISM frames the governance response across Domains 2 and 3. The organizations that have experienced production exploits did not lack technical expertise. They lacked governance programs that treated LLM security decisions with the same rigor applied to the rest of the security environment. AAISM is the credential that validates you can close that gap.

    If you want to move fast, the AAISM Bootcamp delivers all three domains in three intensive days of live online instruction, with expert-led sessions and real-time Q&A throughout. If your schedule requires more flexibility, the AAISM MasterClass gives you the same expert instruction at your own pace, with an adaptive learning system that identifies exactly what you still need to work on across all three domains.

    Before committing to a full program, the free AI Threat Hunting Playbook from Destination Certification is worth reading first. It maps the AI threat detection and response framework that connects directly to what Domain 3 tests, and shows you exactly what gaps formal preparation will need to fill.

    LLM threats are not going to wait for your governance program to catch up. AAISM is how you close that gap.

    Image of John Berti - Destination Certification

    John is a major force behind the Destination Certification CISSP program's success, with over 25 years of global cybersecurity experience. He simplifies complex topics, and he utilizes innovative teaching methods that contribute to the program's industry-high exam success rates. As a leading Information Security professional in Canada, John co-authored a bestselling CISSP exam preparation guide and helped develop official CISSP curriculum materials. You can reach out to John on LinkedIn.

    Image of John Berti - Destination Certification

    John is a major force behind the Destination Certification CISSP program's success, with over 25 years of global cybersecurity experience. He simplifies complex topics, and he utilizes innovative teaching methods that contribute to the program's industry-high exam success rates. As a leading Information Security professional in Canada, John co-authored a bestselling CISSP exam preparation guide and helped develop official CISSP curriculum materials. You can reach out to John on LinkedIn.

    Model Poisoning and Prompt Injection Are Not the Same Thing.

    This free class explains the difference clearly.

    • The specific difference between model poisoning and prompt injection, and why the AAISM exam treats them as distinct threats
    • A walkthrough of a real AI attack scenario so the distinction becomes concrete rather than theoretical
    • How to identify which type of attack is happening when the exam puts you in a scenario-based question
    • What the AAISM exam specifically expects you to know about each threat before you sit the exam

    The easiest way to get your AAISM Certification 


    Learn about our AAISM MasterClass

    Image of masterclass video - Destination Certification

    The fastest way to get AAISM Certified. Join our bootcamp


    Our bootcamp isn't just about getting you to pass—it's about developing the leadership skills security managers need.