Automated penetration testing: what it catches and what it misses
Date
Date
Updated
Updated
Reading time
Reading time
4 minutes
4 minutes
Most automated penetration testing tools do one of two things. They read your source code or they poke at a live URL. Neither one can see a mistake that undoes a lot of otherwise careful security work: a dev environment sitting wide open on the internet, quietly configured with the same credentials as production.
What automated penetration testing tools do today
Two categories cover most of the market.
Source-code-only tools extend static analysis with AI. Cursor's Bugbot reviews pull request diffs and flags bugs, security issues, and code quality problems in the changed lines. Vercel's deepsec runs a regex scan across a codebase, then sends agents to investigate flagged files, tracing data flows and enriching results with git metadata. Visa's open-source vulnerability agentic harness goes further than either one: by its own description, it "reads the target and does not build, run, or test it," with only a narrow, opt-in exception for verifying specific API endpoints on localhost. All three are explicit that they work from source, not from a live, deployed system.
Live-endpoint tools go the other direction. XBOW pitches itself in one line by saying, "point XBOW at a URL and it does the rest" ("the rest" meaning chaining vulnerabilities into working attacks against a deployed application). Network-level red-teaming tools like Decepticon reach further into the infrastructure, mapping networks and running full attack chains against a live environment rather than a single application.
Strix blends both categories in one agent: it can scan a local codebase or a GitHub repo, and also test a live URL directly, chaining findings into working proof-of-concept exploits. That's progress on the code-versus-live split, but nothing in its documentation claims it reasons over deployment configuration, network topology, or the relationship between two environments.
Where each approach succeeds and fails
Source-code tools: fast, run in the PR loop, catch real bugs before they ship. But can't see anything outside the repository, including where or how the code is deployed.
Live-endpoint tools: prove exploitability against a real running system, not just theoretical risk in a diff. But they only see what a probe or attack chain happens to touch, and most of a deployed system is invisible from the outside. Backend services sit behind proxies, firewalls, and private networks, with no public endpoint at which to aim a scanner, while staying reachable through narrow holes like a misconfigured allowlist, an SSRF in something public, an environment nobody locked down. An external probe can only find those holes by stumbling into them, and the tools that map networks directly still have to find their way in first.
Every tool above runs when you tell it to: on a pull request, on a schedule, when someone kicks off a scan. But an attacker doesn't work that way. A persistent actor probes continuously (with the same model assistance available to any defender), which puts far more total effort into finding one thin hole than a scheduled test puts into closing it.
Neither category is built to answer a different, arguably more important question: is this environment connected to something it shouldn't be able to reach?
Four patterns that slip past both approaches
The clearest example of this gap: a team spins up a dev or staging version of a service. It's convenient to point it at the same downstream integrations production uses, so someone reuses the same API credentials rather than provisioning separate ones. The dev environment doesn't get the same network protection production does, because it's "just dev." It ends up publicly reachable, with no IP allowlist, holding credentials that work against production-equivalent systems.
Read the source code of either environment and you won't see this. The code that talks to the downstream integration looks identical in both places, because it is identical. The problem only becomes visible when you compare what's actually configured in each deployed environment side by side, and know which one of them is sitting on the open internet.
An exposed dev environment hands an attacker a map of production's interior. Backend services that never touch the public internet in production are often reachable in dev, so an attacker can enumerate them, find vulnerabilities in them, and learn the internal architecture without going near production. Many of those services have no authentication of their own, because the network boundary was the control, which means reaching them is the whole attack. Then they aim what they learned at production's front door through an SSRF or a permissive CORS policy. Exploring a production system through a front-door vulnerability is hard. Exploiting a backend service you already know is there is much easier.
Dev environments also signal where a company is investing. Sustained change in one shows an attacker which parts of the product are getting engineering effort, and that's a good proxy for what the business cares about. A vulnerability found there while a feature is still being built tells them what to watch for, and when that code ships to production they already know where to look.
The same shape shows up in three other places:
A decommissioned dev or staging subdomain still trusted by production. When a dev environment is torn down, teams often forget to remove the DNS record pointing to it. That leaves the subdomain dangling, and anyone can claim the underlying resource and take it over. If production's own code still sends secrets, like OAuth credentials, to that subdomain, Microsoft's guidance on dangling DNS is clear that "this data might be exposed to third parties" once someone else controls it.
A shared account or IAM role across environments. AWS treats this as basic hygiene: its own Well-Architected Framework calls an account "a hard boundary" and recommends "account-level separation... for isolating production workloads from development and test workloads." A dev environment's compute role that can also reach production's S3 bucket or KMS key doesn't show up in that dev app's code, and it doesn't show up from probing the dev app alone.
A CI/CD pipeline holding credentials for more than one environment. A pipeline routinely holds more reach than any single deployed environment reveals, so neither reading one app's code nor probing one environment shows what a compromise there would cost. Codecov's 2021 breach is adjacent rather than identical evidence: an attacker extracted a credential through a flaw in its Docker image build process rather than a dev/prod mix-up.
Neither category catches any of these, for different reasons. Static analysis tools are, in OWASP's own words, "frequently unable to find configuration issues, since they are not represented in the code." Every pattern above is a fact about how two or more deployments relate to each other, not a bug sitting in any one repository.
Live-endpoint testing has the mirror problem: it has to discover network topology by probing from outside, which is slower and misses whatever a probe never touches. "This dev environment holds a production-equivalent credential" isn't a network-reachability finding to begin with, it's a configuration fact, and a red-teaming tool only finds it by accident if its attack chain happens to compromise that specific environment.
Why deployment context sits outside pentesting tools
These aren't the only two approaches anyone could have built. But they are the only two that don't require a company to hand a third-party vendor deep access to its infrastructure. A source-code tool needs read access to a repository. A live-endpoint tool needs a public URL, the same thing an actual attacker already has. Neither requires standing, privileged access to a company's cloud account (its IAM roles, its network configuration, its secrets manager, etc). A tool that sees deployment context needs exactly that kind of access, and handing a pentesting vendor that level of trust is a far bigger decision than granting read access to a repo or pointing it at a URL.
That's also why Microsoft's own capability lives inside Defender for Cloud (a cloud security posture management product) and not inside a pentesting tool. A company already grants a CSPM tool broad, standing visibility into its cloud account, because that's the whole premise of the product. A pentesting vendor asking for that same access, specifically to run a scan, is a much harder sell. As a result, deep account access and adversarial testing mostly live in separate product categories, sold by different vendors, bought by different teams inside the same company.
Even where the access exists, adversarial testing doesn't follow. Defender for Cloud's own feature list (attack path analysis, risk prioritization, security explorer risk hunting) describes graph-based inference over configuration data, not the kind of chained, working exploit XBOW or Strix produce. It's also owned by whichever team manages cloud security posture, not whoever runs pentesting or AppSec, so a capability that works doesn't automatically show up in anyone's pentesting workflow. And the code-to-cloud piece connects only through GitHub, Azure DevOps, or GitLab, so a team on Bitbucket or a self-hosted git server doesn't get it at all, regardless of budget or awareness.
That separation is starting to break down where the infrastructure provider builds the tool itself. AWS Continuum combines code vulnerability analysis that "reasons over your environment, confirms what is real" with on-demand pentesting through "tailored multi-step attack scenarios, complete with reproducible proof," built by the company that runs the underlying infrastructure rather than a third-party vendor asking for access to it. What AWS's docs don't say is whether that environmental reasoning covers the pattern this piece is about: comparing configuration across two environments, or catching a dev-side credential that also works in production. Several of its capabilities are still gated or in preview. Continuum is evidence that deep access plus adversarial testing is a real category forming across the industry, not just something Aptible claims. It isn't evidence that anyone, including AWS, has closed this specific gap.
Your options for closing the gap
None of this is a model problem. A better AI won't give a source-code scanner visibility into infrastructure it was never pointed at, and it won't give a live-endpoint scanner prior knowledge of a network it has to discover by probing. Deployment context (meaning the live configuration, network architecture, and which credentials are in use where) only exists in one place: whatever system actually deployed that environment.
From there, a few options:
Check it by hand before buying anything. List every credential your dev and staging environments hold and ask whether each one is distinct to that environment or a copy of what production uses. Then ask the harder question: what can dev reach? If a production allowlist admits your stack's egress IP, every application sharing that stack can reach whatever that allowlist was meant to protect, which is an argument for dedicated, walled-off networks over shared ones. Distinct credentials don't make an exposed dev environment safe either, because the exposure is the risk on its own.
Wire your existing systems together yourself. Whoever manages the cloud account, whoever wrote the CI/CD pipeline, and whoever provisioned the credentials each hold a piece of this. Someone has to connect them manually to get the full picture.
Pay for a posture-management add-on. Microsoft's own Defender for Cloud sells a capability called "code-to-cloud contextualization" built for exactly this: mapping code to the cloud resources it deploys to. It's a real, purchasable option today, gated behind a paid CSPM tier and, per the reasoning above, adoption friction that has nothing to do with whether it works.
Run continuous testing instead of an annual pentest. Continuous automated red teaming, or CART, is the practice of running offensive testing on an ongoing basis rather than once a year, with automation making that economically possible. The case for it is a timing gap: attackers now generate working exploits in hours to days for a growing class of high-value vulnerabilities, while patch development, QA, and deployment still take weeks (a shift security researcher Rob Fuller lays out in his manifesto "The Day-Zero Normal"). A program built around one test a year lives entirely inside that gap, regardless of what the test covers.
Use a deployment platform that already has the context. This option sidesteps the trust problem laid out above rather than solving it. Aptible already has deep access to a customer's infrastructure, not because anyone made a special decision to grant a security vendor that access, but because Aptible is their hosting platform. A scan against an Aptible-hosted app therefore has the network topology and configuration in scope by default. Credential reuse across environments is detectable without the secret values themselves reaching a model or an engineer's terminal, and having the code and the deployed topology in one place is what makes it possible to trace a chain from a minor issue in one service to real impact in another. That's not a reason to switch platforms on its own; it's just one place where the access and the testing already happen to sit with the same party.
See what a scan would find on your own deployments
The product isn't self-serve yet. We're still working toward general availability with several existing customers. But if you want a sense of what a scan would find on your own deployment, you can talk to the engineers building it. Depending on what you're running, they may be able to run a scan against your actual infrastructure directly. Talk to our engineering team →
Most automated penetration testing tools do one of two things. They read your source code or they poke at a live URL. Neither one can see a mistake that undoes a lot of otherwise careful security work: a dev environment sitting wide open on the internet, quietly configured with the same credentials as production.
What automated penetration testing tools do today
Two categories cover most of the market.
Source-code-only tools extend static analysis with AI. Cursor's Bugbot reviews pull request diffs and flags bugs, security issues, and code quality problems in the changed lines. Vercel's deepsec runs a regex scan across a codebase, then sends agents to investigate flagged files, tracing data flows and enriching results with git metadata. Visa's open-source vulnerability agentic harness goes further than either one: by its own description, it "reads the target and does not build, run, or test it," with only a narrow, opt-in exception for verifying specific API endpoints on localhost. All three are explicit that they work from source, not from a live, deployed system.
Live-endpoint tools go the other direction. XBOW pitches itself in one line by saying, "point XBOW at a URL and it does the rest" ("the rest" meaning chaining vulnerabilities into working attacks against a deployed application). Network-level red-teaming tools like Decepticon reach further into the infrastructure, mapping networks and running full attack chains against a live environment rather than a single application.
Strix blends both categories in one agent: it can scan a local codebase or a GitHub repo, and also test a live URL directly, chaining findings into working proof-of-concept exploits. That's progress on the code-versus-live split, but nothing in its documentation claims it reasons over deployment configuration, network topology, or the relationship between two environments.
Where each approach succeeds and fails
Source-code tools: fast, run in the PR loop, catch real bugs before they ship. But can't see anything outside the repository, including where or how the code is deployed.
Live-endpoint tools: prove exploitability against a real running system, not just theoretical risk in a diff. But they only see what a probe or attack chain happens to touch, and most of a deployed system is invisible from the outside. Backend services sit behind proxies, firewalls, and private networks, with no public endpoint at which to aim a scanner, while staying reachable through narrow holes like a misconfigured allowlist, an SSRF in something public, an environment nobody locked down. An external probe can only find those holes by stumbling into them, and the tools that map networks directly still have to find their way in first.
Every tool above runs when you tell it to: on a pull request, on a schedule, when someone kicks off a scan. But an attacker doesn't work that way. A persistent actor probes continuously (with the same model assistance available to any defender), which puts far more total effort into finding one thin hole than a scheduled test puts into closing it.
Neither category is built to answer a different, arguably more important question: is this environment connected to something it shouldn't be able to reach?
Four patterns that slip past both approaches
The clearest example of this gap: a team spins up a dev or staging version of a service. It's convenient to point it at the same downstream integrations production uses, so someone reuses the same API credentials rather than provisioning separate ones. The dev environment doesn't get the same network protection production does, because it's "just dev." It ends up publicly reachable, with no IP allowlist, holding credentials that work against production-equivalent systems.
Read the source code of either environment and you won't see this. The code that talks to the downstream integration looks identical in both places, because it is identical. The problem only becomes visible when you compare what's actually configured in each deployed environment side by side, and know which one of them is sitting on the open internet.
An exposed dev environment hands an attacker a map of production's interior. Backend services that never touch the public internet in production are often reachable in dev, so an attacker can enumerate them, find vulnerabilities in them, and learn the internal architecture without going near production. Many of those services have no authentication of their own, because the network boundary was the control, which means reaching them is the whole attack. Then they aim what they learned at production's front door through an SSRF or a permissive CORS policy. Exploring a production system through a front-door vulnerability is hard. Exploiting a backend service you already know is there is much easier.
Dev environments also signal where a company is investing. Sustained change in one shows an attacker which parts of the product are getting engineering effort, and that's a good proxy for what the business cares about. A vulnerability found there while a feature is still being built tells them what to watch for, and when that code ships to production they already know where to look.
The same shape shows up in three other places:
A decommissioned dev or staging subdomain still trusted by production. When a dev environment is torn down, teams often forget to remove the DNS record pointing to it. That leaves the subdomain dangling, and anyone can claim the underlying resource and take it over. If production's own code still sends secrets, like OAuth credentials, to that subdomain, Microsoft's guidance on dangling DNS is clear that "this data might be exposed to third parties" once someone else controls it.
A shared account or IAM role across environments. AWS treats this as basic hygiene: its own Well-Architected Framework calls an account "a hard boundary" and recommends "account-level separation... for isolating production workloads from development and test workloads." A dev environment's compute role that can also reach production's S3 bucket or KMS key doesn't show up in that dev app's code, and it doesn't show up from probing the dev app alone.
A CI/CD pipeline holding credentials for more than one environment. A pipeline routinely holds more reach than any single deployed environment reveals, so neither reading one app's code nor probing one environment shows what a compromise there would cost. Codecov's 2021 breach is adjacent rather than identical evidence: an attacker extracted a credential through a flaw in its Docker image build process rather than a dev/prod mix-up.
Neither category catches any of these, for different reasons. Static analysis tools are, in OWASP's own words, "frequently unable to find configuration issues, since they are not represented in the code." Every pattern above is a fact about how two or more deployments relate to each other, not a bug sitting in any one repository.
Live-endpoint testing has the mirror problem: it has to discover network topology by probing from outside, which is slower and misses whatever a probe never touches. "This dev environment holds a production-equivalent credential" isn't a network-reachability finding to begin with, it's a configuration fact, and a red-teaming tool only finds it by accident if its attack chain happens to compromise that specific environment.
Why deployment context sits outside pentesting tools
These aren't the only two approaches anyone could have built. But they are the only two that don't require a company to hand a third-party vendor deep access to its infrastructure. A source-code tool needs read access to a repository. A live-endpoint tool needs a public URL, the same thing an actual attacker already has. Neither requires standing, privileged access to a company's cloud account (its IAM roles, its network configuration, its secrets manager, etc). A tool that sees deployment context needs exactly that kind of access, and handing a pentesting vendor that level of trust is a far bigger decision than granting read access to a repo or pointing it at a URL.
That's also why Microsoft's own capability lives inside Defender for Cloud (a cloud security posture management product) and not inside a pentesting tool. A company already grants a CSPM tool broad, standing visibility into its cloud account, because that's the whole premise of the product. A pentesting vendor asking for that same access, specifically to run a scan, is a much harder sell. As a result, deep account access and adversarial testing mostly live in separate product categories, sold by different vendors, bought by different teams inside the same company.
Even where the access exists, adversarial testing doesn't follow. Defender for Cloud's own feature list (attack path analysis, risk prioritization, security explorer risk hunting) describes graph-based inference over configuration data, not the kind of chained, working exploit XBOW or Strix produce. It's also owned by whichever team manages cloud security posture, not whoever runs pentesting or AppSec, so a capability that works doesn't automatically show up in anyone's pentesting workflow. And the code-to-cloud piece connects only through GitHub, Azure DevOps, or GitLab, so a team on Bitbucket or a self-hosted git server doesn't get it at all, regardless of budget or awareness.
That separation is starting to break down where the infrastructure provider builds the tool itself. AWS Continuum combines code vulnerability analysis that "reasons over your environment, confirms what is real" with on-demand pentesting through "tailored multi-step attack scenarios, complete with reproducible proof," built by the company that runs the underlying infrastructure rather than a third-party vendor asking for access to it. What AWS's docs don't say is whether that environmental reasoning covers the pattern this piece is about: comparing configuration across two environments, or catching a dev-side credential that also works in production. Several of its capabilities are still gated or in preview. Continuum is evidence that deep access plus adversarial testing is a real category forming across the industry, not just something Aptible claims. It isn't evidence that anyone, including AWS, has closed this specific gap.
Your options for closing the gap
None of this is a model problem. A better AI won't give a source-code scanner visibility into infrastructure it was never pointed at, and it won't give a live-endpoint scanner prior knowledge of a network it has to discover by probing. Deployment context (meaning the live configuration, network architecture, and which credentials are in use where) only exists in one place: whatever system actually deployed that environment.
From there, a few options:
Check it by hand before buying anything. List every credential your dev and staging environments hold and ask whether each one is distinct to that environment or a copy of what production uses. Then ask the harder question: what can dev reach? If a production allowlist admits your stack's egress IP, every application sharing that stack can reach whatever that allowlist was meant to protect, which is an argument for dedicated, walled-off networks over shared ones. Distinct credentials don't make an exposed dev environment safe either, because the exposure is the risk on its own.
Wire your existing systems together yourself. Whoever manages the cloud account, whoever wrote the CI/CD pipeline, and whoever provisioned the credentials each hold a piece of this. Someone has to connect them manually to get the full picture.
Pay for a posture-management add-on. Microsoft's own Defender for Cloud sells a capability called "code-to-cloud contextualization" built for exactly this: mapping code to the cloud resources it deploys to. It's a real, purchasable option today, gated behind a paid CSPM tier and, per the reasoning above, adoption friction that has nothing to do with whether it works.
Run continuous testing instead of an annual pentest. Continuous automated red teaming, or CART, is the practice of running offensive testing on an ongoing basis rather than once a year, with automation making that economically possible. The case for it is a timing gap: attackers now generate working exploits in hours to days for a growing class of high-value vulnerabilities, while patch development, QA, and deployment still take weeks (a shift security researcher Rob Fuller lays out in his manifesto "The Day-Zero Normal"). A program built around one test a year lives entirely inside that gap, regardless of what the test covers.
Use a deployment platform that already has the context. This option sidesteps the trust problem laid out above rather than solving it. Aptible already has deep access to a customer's infrastructure, not because anyone made a special decision to grant a security vendor that access, but because Aptible is their hosting platform. A scan against an Aptible-hosted app therefore has the network topology and configuration in scope by default. Credential reuse across environments is detectable without the secret values themselves reaching a model or an engineer's terminal, and having the code and the deployed topology in one place is what makes it possible to trace a chain from a minor issue in one service to real impact in another. That's not a reason to switch platforms on its own; it's just one place where the access and the testing already happen to sit with the same party.
See what a scan would find on your own deployments
The product isn't self-serve yet. We're still working toward general availability with several existing customers. But if you want a sense of what a scan would find on your own deployment, you can talk to the engineers building it. Depending on what you're running, they may be able to run a scan against your actual infrastructure directly. Talk to our engineering team →
Related posts
More on this topic from the Aptible blog.
More on this topic from the Aptible blog.
548 Market St #75826 San Francisco, CA 94104
© 2026. All rights reserved. Privacy Policy
548 Market St #75826 San Francisco, CA 94104
© 2026. All rights reserved. Privacy Policy
548 Market St #75826 San Francisco, CA 94104
© 2026. All rights reserved. Privacy Policy



