Skip to main content

đź”” Notifications Hub: See compliance updates in one place

  • blog
  • AI Penetration Testing: What Security Teams Need to Know About LLMs, MCP, and AI Agents

AI Penetration Testing: What Security Teams Need to Know About LLMs, MCP, and AI Agents

  • September 24, 2026
Author

Sherif Koussa

CEO and Founder of Software Secured

Reviewer

Anna Fitzgerald

Senior Content Marketing Manager

This article is written and contributed by penetration testing company Software Secured, a proud Secureframe partner.

Most security teams still scope a penetration test the way they did five years ago: count domains, count forms, count API endpoints. That model breaks down the moment AI enters the picture, because an AI feature doesn't behave like a form that pings a database and returns a result. It reasons, holds context, and increasingly acts on your behalf across other systems.

I recently sat down with Marc Rubbinaccio, Head of Cybersecurity and Compliance at Secureframe, to talk through what actually changes when penetration testers start scoping AI functionality.

We discussed how the industry has seen four real technology shifts:

The internet changed everything.
Mobile changed everything again.
Cloud did it a third time.
AI is now the fourth.

You wouldn't test an on-prem server the way you test a cloud environment, and you can't test an AI feature the way you test a standard web form.

Here's what security teams should actually take from that conversation.

AI has changed what a penetration test needs to cover

In practice, AI functionality shows up in at least four distinct forms, and each one represents a different attack surface:

  • Direct LLM API calls: The application passes a prompt straight to a model like ChatGPT or Claude
  • MCP servers: A protocol layer that lets an AI system connect to and act on other tools and data sources
  • RAG layers: The model's prompt gets combined with retrieved data, such as a database or a set of internal documents
  • Agents: AI systems that take multi-step actions, often across several connected tools, with limited human oversight in the loop

Each needs a different testing approach because each changes what an attacker can reach and how. Software Secured built a dedicated AI test plan roughly a year ago, aligned to the OWASP Top 10 for LLMs, specifically because a generic web app methodology doesn't surface these risks.

Most engagements start as a gray box: pentesters get application access and credentials but not source code, which mirrors how most organizations actually want to validate resilience within a reasonable budget and timeline.

If you're scoping an AI pentest, the first question to answer is: Which of these four categories does our AI functionality fall into, and what does each one expose?

Recommended reading

Penetration Testing 101: A Guide to Testing Types, Processes, and Costs

Three AI security risks penetration testers are looking for

Three vulnerability classes come up repeatedly in AI-specific testing, and they map surprisingly closely to attack patterns security teams already understand from traditional web application testing.

Prompt injection and jailbreaking

This is the most well-known AI risk: convincing a model to abandon its instructions and follow the attacker's instead.

In one real engagement, a fintech company's AI assistant was manipulated through a series of prompts designed to make the model more permissive over the course of a conversation. The result was a critical vulnerability with a 9.3 CVSS score, and pentesters could access another user's account entirely through the chat interface.

No traditional exploit code was involved. The attack path was conversational.

Indirect prompt injection

This variant is harder to detect because the attacker never talks to the chatbot directly. Instead, malicious instructions get embedded in content the AI is asked to process: a file upload, a document, or an external source the model consumes as part of its task.

Development teams often build guardrails assuming malicious input will come from the user typing into the chat. Indirect prompt injection breaks that assumption by hiding the attack inside the data itself.

Excessive agency

The third pattern involves manipulating an AI system into performing actions they shouldn’t, often to consume far more resources than intended, driving up API costs and potentially making the system unavailable to everyone else. For any CISO tracking AI spend, this risk shows up on the finance side of the house before it shows up as a security incident.

Each of these has a direct analog in traditional web application security. Prompt injection resembles authorization bypass. Indirect prompt injection echoes the logic behind trusting user-supplied content that gets processed downstream. Excessive agency is a denial-of-service vector, just paid for in tokens instead of bandwidth.

The tactics look new, but the underlying reasoning an attacker uses hasn't changed.

Traditional web app risks mapped to their AI-specific equivalents: authorization bypass to prompt injection and jailbreaking, trusting user input processed downstream to indirect prompt injection, and denial of service to excessive agency and resource exhaustion

Recommended reading

Emerging Cyber Threats in 2026: What SaaS Companies Need to Do Now to Prepare

Why MCP Servers and AI agents create new security risks

MCP exists to let an AI system talk to other applications and data sources, connecting a model to a CRM to summarize customer emails, for example. The risk shows up when the MCP server or the agent runs with permissions broader than the user's own, often admin-level access, rather than being scoped to what that specific user is actually allowed to do.

Agents carry the same problem, frequently with sharper consequences. In one engagement, pentesters found an agent with access to internal APIs that no other system component could reach. Rather than reverse-engineering those APIs manually, which would typically take days or weeks, pentesters simply asked the agent what APIs it had access to and how to call them. It answered.

The more systems an AI agent can reach, the greater the impact if an attacker gets it to cooperate. And getting an AI system to cooperate has repeatedly proven easier than compromising a traditional API directly.

For security and compliance teams already thinking in terms of least privilege and access governance, this is exactly the kind of access control question that belongs in a governance review.

Guiding Your Organization's AI Strategy and Implementation

Follow this guidance and checklist of best practices to effectively implement AI while addressing concerns related to transparency, privacy, and security.

What should an AI penetration test actually test?

Based on the risk categories above, a thorough AI penetration test should cover:

  • LLM and API interactions
  • Prompt injection and jailbreak resistance
  • RAG pipelines and the integrity of retrieved content
  • MCP tool access and permission boundaries
  • Agent permissions relative to the user they're acting on behalf of
  • Authentication and authorization at every layer the AI touches
  • Input validation for anything the model consumes, including indirect sources
  • Output handling, since content coming back from the model needs the same scrutiny as content going in
  • Attack chains that combine multiple smaller AI-specific weaknesses into a critical path

Vulnerabilities consistently cluster around four remediation areas: input validation, authentication, authorization, and output handling.

For example, an application summarized user input and presented it to an admin. An attacker embedded a benign-looking URL in that input. When the admin clicked it, the admin's session cookie went straight to the attacker, who logged in with full admin access.

The AI didn't need to be "hacked" in any traditional sense. It just did exactly what it was designed to do: summarize and present content, without validating what that content might trigger downstream.

If your organization is evaluating an AI pentest scope, this AI security checklist can help translate these categories into a concrete list of questions to bring to a vendor.

AI isn't replacing pentesters

Given how much of this conversation centers on AI as an attack surface, it's worth addressing the inverse question directly: if AI can be used to attack systems, why do you still need a human penetration tester?

Software Secured's internal use of AI answers that question in practice. The firm runs its own private AI instance rather than sending client data to third-party model providers, specifically to avoid any risk of client data being used for retraining.

Within that controlled environment, AI supports pentesters in five specific ways:

  1. Reconnaissance and enumeration: Summarizing scan data quickly, and even inferring a target's tech stack from historical job postings
  2. Payload generation: Adapting attack payloads based on live server responses instead of manually maintaining a fixed list
  3. Report writing
  4. Standardizing CVSS risk scoring across pentesters so severity ratings stay consistent
  5. Mapping attack chains: Linking the sequences of smaller vulnerabilities that combine into a path to something critical like database access.

AI is not used to run attacks directly against client systems or make judgment calls about severity or exploitability.

Current research on AI-driven offensive security suggests models perform well against known, well-documented vulnerability classes on structured targets, but human judgment remains the deciding factor in real engagements, especially anything involving business logic or novel attack chains.

The stated philosophy is straightforward: AI multiplies the pentester's effort. It doesn't replace the pentester.

Recommended reading

Former NSA Chief Paul Nakasone: "There Are Likely Adversaries in Your Network Right Now"

What security teams should do next

If your organization is deploying AI in any customer-facing or internal capacity, a few questions are worth asking before your next pentest, and before your next compliance audit:

  • What AI systems are actually in production, and which of the four categories above does each one fall into?
  • What data can each system access, and is that access scoped to the individual user or elevated by default?
  • What tools, APIs, or third-party systems can the AI invoke on its own?
  • Can external or user-generated content influence how the AI behaves?
  • What happens to AI-generated output once it leaves the model, and who or what consumes it downstream?
  • Have AI-specific attack scenarios, not just standard web app tests, actually been included in your last penetration test?
  • Do you retest vulnerabilities after remediation the same way you would any other critical vulnerability?

AI is changing what "in scope" means for a penetration test, and it's changing faster than most testing programs have caught up. Closing that gap starts with asking the right questions before the next assessment, not after an incident forces the conversation.

Recommended reading

2026's Biggest Cybersecurity Threats: Analyzing Recent Attacks, Emerging Threats + How to Defend Against Them

How Secureframe + Software Secured can help

AI is expanding attack surfaces faster than most traditional cybersecurity programs were built to handle. Every new LLM integration, MCP server, RAG pipeline, or agent introduces access paths and data flows that security teams, customers, or auditors will eventually ask about. Keeping up means testing and governing those systems the right way and proving your controls hold up over time.

An AI-specific penetration test gives you evidence that your AI functionality has been assessed against the risks that matter, including prompt injection, excessive agency, and over-permissioned agents.

That evidence supports the penetration testing and vulnerability management requirements in frameworks like SOC 2 and ISO 27001, and it helps you demonstrate responsible AI risk management under emerging standards like ISO 42001.

Compliance automation makes it easier to act on what a pentest uncovers. Secureframe's platform helps you track remediation, map pentest results to the controls they satisfy, maintain access reviews that keep AI systems scoped to least privilege, and continuously monitor your environment so gaps don't reappear between assessments. And our team of in-house compliance experts can help you navigate new AI frameworks without the cost of outside consultants. Request a demo today.

You can also learn more about Software Secured's AI penetration testing service, which is built around the OWASP Top 10 for LLMs and scoped to the way your AI actually works.

FAQs

What is AI penetration testing?

AI penetration testing is a security assessment in which ethical hackers attempt to exploit the AI functionality in an application, including LLM integrations, MCP servers, RAG pipelines, and AI agents. It evaluates whether an attacker could manipulate the AI into bypassing its instructions, exposing data, or taking actions the user isn't authorized to take.

How is AI penetration testing different from traditional penetration testing?

Traditional penetration testing focuses on deterministic components like forms, APIs, and infrastructure, where the same input produces the same output. AI penetration testing adds attack paths that are conversational and context-dependent, such as prompt injection and manipulation of agents that act across connected systems. The underlying goals stay the same: finding authentication, authorization, input validation, and output handling weaknesses before an attacker does. Learn more about penetration testing and another key assessment method, vulnerability scanning.

Do compliance frameworks require AI penetration testing?

While most today don’t explicitly require an AI-specific penetration test, frameworks like SOC 2, ISO 27001, and PCI DSS expect you to identify and remediate vulnerabilities across in-scope systems, and AI features that touch customer data are increasingly in scope. AI governance standards like ISO 42001 and the NIST AI Risk Management Framework also expect organizations to assess and manage AI-specific risks, and pentest results are strong evidence for both.

How often should you perform an AI penetration test?

Most organizations should test AI functionality at least annually as part of their regular penetration testing cycle, and again after significant changes. For AI systems, significant changes include adopting a new model, connecting a new MCP server or tool, expanding an agent's permissions, or adding new data sources to a RAG pipeline. Critical findings should be retested after remediation to confirm the fix holds.

Sherif Koussa

CEO and Founder of Software Secured

Sherif Koussa is the CEO and Founder of Software Secured, penetration testing and augmented security services company. After a rewarding career as a software developer, Sherif realized the significance of secure coding practices and became enamored with the world of security. Sherif grew the OWASP Ottawa Chapter from a few members to a few thousand, helped launch two open source projects for OWASP (WebGoat & Cheat Sheets), and helped SANS launch two exams GSSP-Java and GSSP-NET.

Anna Fitzgerald

Senior Content Marketing Manager

Anna Fitzgerald is a digital and product marketing professional with nearly a decade of experience delivering high-quality content across highly regulated and technical industries, including healthcare, web development, and cybersecurity compliance. At Secureframe, she specializes in translating complex regulatory frameworks—such as CMMC, FedRAMP, NIST, and SOC 2—into practical resources that help organizations of all sizes and maturity levels meet evolving compliance requirements and improve their overall risk management strategy.