Agentic AI Security Testing: From Scripted Scanning to Autonomous Investigation
Agentic AI security testing is the use of autonomous AI agents that can form hypotheses, take actions through specialized penetration testing tools, interpret application responses, and decide what to test next, while remaining constrained by human-defined scope, permissions, and approval rules. This is not another round of noisy, signature-based scanning. It is a fundamental rethink of how we discover weaknesses in complex systems. The key takeaway: penetration testers are starting to manage teams of AI investigators instead of running every payload and request by hand. The shift is overdue. As technology advances, securing systems, networks, and applications has become more critical, and manual testing alone cannot keep up with sprawling digital estates and relentless threats. At the same time, AI vulnerability discovery is no longer speculative—models can already find vulnerabilities, so the pressing question is whether we can trust them against real targets. Agentic AI with strong governance is the answer emerging from frontline security work.

Burp AT: Agentic AI Inside a Mature Testing Workflow
PortSwigger has announced the public beta of Burp AT, a new addition to Burp Suite that brings agentic AI to professional penetration testing. This matters because it embeds AI vulnerability discovery into a tool stack that testers already trust. Burp AT enables penetration testers to delegate defined investigative tasks to AI agents that use Burp Suite’s tools, project context, and purpose-built pentesting capabilities. The design choice is opinionated: the model is allowed "room to reason," but Burp controls what it can actually do, executes work through existing tools, and preserves the evidence. Scope, tool access, and approval rules are enforced in Burp’s tooling layer, architecturally separate from the model. In practice, that means the agent cannot quietly redefine scope or bypass boundaries—a crucial safeguard when you are pointing AI at real production systems. Burp AT allows pentesters to choose how much work agents take on for each task and engagement, starting with tight supervision and increasing autonomy when performance and target sensitivity permit. Responsibility for scope definition, judgment, and validation stays where it belongs: with the human tester.
From One-Off Scans to Continuous, Intelligent Security Testing
The security industry has been moving toward automation for years, but most penetration testing tools still revolve around large, predefined test catalogs. Astra, for example, supports over 10,000 automated tests across different infrastructures, combining manual and automated penetration testing for applications, networks, APIs, and blockchain. Acunetix supports over 12,000 web application vulnerability tests, coupled with scheduled or recurring scans and severity-based classification. These platforms, along with classics like Metasploit, have helped organizations gauge how strong their security systems are and identify weaknesses or vulnerabilities. Yet they largely follow a scripted pattern: run scans, get findings, react. Agentic AI changes that rhythm. Models can now do more than run predefined checks; they can form hypotheses, act through tools, interpret how an application responds, and decide what to try next. Used alongside these penetration testing tools, agentic AI security testing reduces manual overhead and pushes teams toward continuous, intelligent investigation instead of periodic box-ticking. The focus shifts from chasing scan outputs to steering AI agents through evolving threat models.
| Aspect | Traditional Automation | Agentic AI in Burp AT |
|---|---|---|
| Testing style | Predefined checks and signatures | Hypothesis-driven, adaptive investigation |
| Tool use | General HTTP libraries or scanners | Native Burp Suite tools and project context |
| Control model | Scanner config and rules | Scope, permissions, and approval enforced outside the model |
| Evidence | Scan reports | Recorded agent requests and tool activity within the Burp project |

Governance, Audit Trails, and Integration: What Enterprises Must Demand
Enterprises tempted by agentic AI security testing cannot treat these platforms as black boxes. Governance is non-negotiable. Burp AT’s approach shows what good looks like: scope, tool access, and approval rules are enforced in the tooling layer rather than inside the model, so guardrails cannot be reinterpreted by AI. Agent requests and tool activity are recorded in the Burp project as testing progresses, giving pentesters a record they can inspect alongside the rest of the engagement instead of relying only on the model’s own account. This level of auditability should become the baseline. Penetration testing tools already emphasize compliance checks and reporting features to help organizations meet standards like SOC2, GDPR, ISO 27001, and HIPAA. Agentic platforms must plug into that reality with clear logs, exportable reports, and integration points for issue trackers and CI/CD chains. Pricing models are another practical filter: Astra’s DAST Scanner Lite starts at USD 69 (approx. RM322) per month, rising to USD 499 (approx. RM2,331) for the Scanner Agency tier, while its pentest plans start at USD 199 (approx. RM929) per month and go up to USD 5,999 (approx. RM28,014) per year. These numbers reflect a market where spending is expected to reach USD 6.98 billion (approx. RM32.9 billion) in value by 2032. Enterprises should demand governance that matches that scale of investment.

Human-Guided Agents and the Shift to Proactive Threat Modeling
The most important change is philosophical: modern penetration testing tools are moving from reactive patching to proactive threat modeling, and agentic AI can accelerate that shift. Pentests have always blended experienced security professionals with powerful tools. Burp AT extends that pattern by adding structured, task-specific pentesting skills developed with PortSwigger Research, which agents can apply during real tests as new techniques are validated. Users can begin with tighter supervision and increase autonomy where agent performance, target sensitivity, and engagement rules justify it. That mix of agentic automation with human oversight is how we avoid the worst failure modes of both manual and AI-first testing: missed edge cases on one side, uncontrolled actions on the other. Looking ahead, as PortSwigger Research develops and validates new techniques, those approaches can be translated into skills that agents can apply during real tests. If security teams embrace this model—with strong governance, clear audit trails, and careful integration into their broader penetration testing tools stack—they can move from chasing yesterday’s vulnerabilities to shaping tomorrow’s threat landscape on their own terms.






