Penetration Testing Frequency: Continuous vs Annual

CT
CyberOrbit Team
25 min read
Share

The time attackers need to exploit a new vulnerability collapsed from 63 days to five inside five years, according to Mandiant's time-to-exploit analysis published by Google Cloud Threat Intelligence. Set that against a testing calendar and the penetration testing frequency question answers itself in an unexpected way: the problem is not annual versus continuous, it is that most organisations have never written down what event, other than a date on a calendar, causes them to test.

5 days
Mandiant's measured time to exploit, 2023
Down from 63 days across 2018 to 2019, and 32 days across 2021 to 2022, in the same series. The exploitation window is now shorter than most change-approval cycles.

Do the Arithmetic on Your Last Pentest

Pull up your most recent engagement letter and count the testing days. Ten is typical for a mid-market web application and external perimeter scope. Now do the division: 10 ÷ 365 = 0.027, or 2.7% of the year. That is the fraction of calendar time during which a qualified human was actively looking for a way into your estate.

This is arithmetic, not an argument. Two ten-day engagements a year get you to 5.5%. Even a heroic quarterly programme of ten-day tests reaches 11%. Nobody, at any budget, tests their way to full calendar coverage with scheduled human engagements, and no vendor claims otherwise.

⚠️Do the arithmetic on your own engagement
A ten-day engagement covers 10 ÷ 365 = 2.7% of the calendar year. A five-day engagement covers 1.4%. Run the sum for your last engagement and write the number down, because that is the proportion of your annual security posture you have direct evidence about.

The coverage gap only matters if something changes inside it. For most of the history of penetration testing, the gap was tolerable because the exploitation timeline was longer than the testing interval: vulnerabilities were disclosed, patches shipped, and weaponisation followed weeks or months later. A test in March plausibly held its value into the northern summer.

That relationship has inverted, and it inverted quickly:

2018 to 2019: 63 days
Mandiant's measured time to exploit. Defenders had roughly two months to patch before attacks began.
2021 to 2022: 32 days
The window halved as proof-of-concept publication and automated tooling matured.
2023: 5 days
The steepest proportional fall in Mandiant's series. Of the 138 vulnerabilities Mandiant recorded as exploited that year, 70% were exploited as zero days, before a patch existed.
2025: 28% inside a day
VulnCheck's Q1 2025 exploitation trends found 28.3% of newly catalogued known-exploited vulnerabilities had evidence of exploitation within one day of the CVE being published. Patch prioritisation became a daily operational task.
2026: reported negative
Mandiant's M-Trends 2026 reporting puts estimated mean time to exploit at roughly negative seven days, meaning exploitation increasingly precedes patch availability. Treat the direction as established and the exact figure as a single vendor's estimate.

When exploitation arrives within days of disclosure, the interval between your test and the next one is not a quiet period. It is the window in which your estate changes and the internet's collective scanning infrastructure notices. CISA's Known Exploited Vulnerabilities catalog is the public record of which of those vulnerabilities crossed from theoretical to actively used, and it grows every week. We covered how quickly that plays out in practice in the 48-hour exploit window for mid-market companies.

What a point-in-time test does and does not tell you

A point-in-time test tells you, with high confidence and evidentiary weight, that on a specific set of dates, against a specific declared scope, an independent tester found these findings and not others. That is a genuinely valuable statement. It is signed, it is scoped, and a third party can rely on it.

What it does not tell you is anything about the state of your estate on any other day. It is not a claim about your security posture. It is a claim about a moment. The confusion between those two things is where cadence arguments usually go wrong, on both sides, and it is also why so many teams reach for the compliance rulebook to settle the question, only to find the rulebook says something different from what they remembered.

The Standards Never Said "Annual"

Here is where most cadence conversations start from a false premise. Ask a security leader why they test annually and the answer is almost always "compliance requires it." Read the requirement and you find something more demanding than an annual test, and something almost nobody fully implements.

PCI DSS v4.0.1 is the clearest case. Requirement 11.4.2 (internal) and 11.4.3 (external) both require penetration testing performed at least once every twelve months AND after any significant infrastructure or application upgrade or change, as published in the PCI Security Standards Council document library. That conjunction is the whole point.

"
External penetration testing is performed: per the entity's defined testing methodology, at least once every 12 months, and after any significant infrastructure or application upgrade or change.
PCI DSS v4.0.1Requirement 11.4.3

The twelve-month clock is a floor, and the change trigger runs in parallel with it. Requirement 11.4.1 further requires that the testing methodology itself be defined and documented, and 11.4.4 requires that exploitable vulnerabilities found are corrected and that testing is repeated to verify the correction. A test with no retest does not satisfy 11.4.4. We break the full requirement down in our guide to PCI DSS 4.0 pentest requirements.

The other frameworks say markedly different things, and the differences matter more than the similarities:

Requirements 11.4.1–11.4.4 require a documented methodology, internal and external penetration testing at least every 12 months and after any significant change, correction of exploitable vulnerabilities, and retest to verify correction. The "and after any significant change" clause is the half most organisations skip.

SOC 2 is looser than most people assume. The Trust Services Criteria never name penetration testing at a specified interval. CC4.1 requires the entity to select, develop, and perform ongoing and/or separate evaluations to ascertain whether the components of internal control are present and functioning. It is risk-based. The annual pentest is a convention auditors and readiness platforms have converged on, not a written rule, which is why the defensible answer is a documented rationale rather than a number you inherited.

DORA is more prescriptive but narrower than the marketing suggests. Article 24(6) requires in-scope financial entities to perform, at least yearly, appropriate tests on all ICT systems and applications supporting critical or important functions. Threat-led penetration testing under Article 26(1) applies only to entities designated by competent authorities, and runs on an at-least-every-three-years cycle. Most in-scope firms owe the Article 24 baseline, not TLPT, a distinction covered in our DORA penetration testing requirements guide.

The Essential Eight mandates no penetration testing cadence at all. It is a mitigation maturity model, and testing is how you evidence the mitigations, not a control in itself.

Put those together and the conclusion is uncomfortable: a large share of organisations satisfy the twelve-month half of a two-part requirement and have never defined, in writing, what a significant change means for their estate. They are not over-testing or under-testing. They are missing a clause.

What counts as a "significant change" and why you must define it in writing

No standard defines it for you, deliberately, because it depends on your architecture. That means the definition is yours to write and yours to defend. A workable definition names concrete, observable events. Start from this list and cut what does not apply to you:

New internet-facing service, endpoint, or subdomain added
Authentication or authorisation flow changed
Major framework, runtime, or dependency upgrade (e.g. Node.js major version, Rails upgrade)
New third-party integration added (API, OAuth provider, payment processor)
Cloud account, region, or provider added
Network segmentation or firewall ruleset changed
M&A asset absorbed into the estate
New payment card processing path added or modified
MFA system or SSO provider changed
VPN, remote access gateway, or zero-trust policy changed

Write six to ten of those, get them approved, and attach them to your testing policy. Your assessor then has an artefact to evaluate instead of an opinion to challenge, and your engineering team has a trigger that fires without waiting for a calendar. The definition is worth more than the cadence number, because the cadence number without it is only half a control.

With the trigger defined, the next question is what you actually buy to cover the gaps between triggers, and that is where the vocabulary gets slippery.

Three Different Products Are Sold as "Continuous"

The word is doing too much work. When three vendors say continuous, they frequently mean three different products at three different price points with three different outputs. Before you can decide whether you need continuous testing, you have to decide which of these you are actually being sold.

Model What it is Who runs it Typical annual cost
Continuous scanning Automated, recurring vulnerability and exposure scanning against your estate You, self-service Free tiers to $20,000
Continuous PTaaS with rotating researchers A platform plus human testers cycling against your scope on a variable timeline Vendor $15,000 to $60,000, enterprise scopes $100,000+
Scheduled repeat testing The same declared scope tested on a fixed interval (monthly, quarterly), with a deliverable each time Vendor or platform Varies with interval and scope

Continuous scanning is cheap, fast, and genuinely useful. It catches configuration drift, newly exposed services, expired certificates, missing headers, and known CVEs on identified software versions. It runs unattended and it will tell you within hours that a developer exposed a staging environment.

Continuous PTaaS adds human researchers who work your scope over time rather than in a single block. Pricing commonly sits between $15,000 and $60,000 per year, with enterprise scopes exceeding $100,000, per FireCompass's 2026 PTaaS pricing breakdown. The value is human creativity applied more than once a year. The variable is that "continuous" describes the subscription, not necessarily the tester's attention, so ask how many researcher hours per quarter you are actually buying.

Pros
  • Change-triggered coverage: a new exposure is detected in the next scan cycle, not at next year's engagement
  • Shorter exposure window on average, especially for high-velocity estates
  • Retest built into the model, so remediation is verified continuously rather than as a change order
Cons
  • Higher annual spend: PTaaS models run $15,000 to $60,000 per year, and past $100,000 for enterprise scopes
  • Output is a queue, not an artefact, so a third party who was not present cannot rely on it for attestation
  • Scope drift: continuous scans often expand beyond the agreed boundary, creating alert fatigue
  • Alert fatigue: high signal volume requires internal triage capacity that smaller teams may not have

Scheduled repeat testing is the least glamorous and often the most honest fit: the same scope, tested on a defined interval, producing a comparable deliverable each time so you can see whether findings recur.

That third row is where CyberOrbit sits. We run scheduled testing against your declared scope, and the deliverable is reviewed and signed by a certified security professional independent of your team, with real request and response captures, reproduction steps and proof hashes behind each finding. You do not install a scanner, define the methodology, or run the test against yourself.

Be precise about why that last point matters, because the independence argument gets overstated in vendor copy. Requirements 11.4.2 and 11.4.3 both call for organisational independence of the tester while explicitly permitting a qualified internal resource, and they note the tester is not required to be a QSA or ASV. SOC 2 CC4.1 names no independence requirement for penetration testing at all. No standard forbids testing yourself. What independence actually buys is evidentiary weight with the party who was not in the room: the assessor, the QSA, and the enterprise procurement team who cannot verify your internal separation and will not attempt to. It is a question of who the report has to convince, not of what the rulebook permits.

What each model actually detects

Scanning detects the known: published CVEs, misconfiguration, exposure, drift. What it does not reliably detect is broken authorisation logic, business logic flaws, or chained access paths, because those require understanding what your application is for. Authenticated multi-role scanning catches some access-control issues, but the ones that matter are usually the ones that need a tester who understands the product. We work through where that boundary sits in when automated pentesting is enough.

The practical read is a three-way split rather than a two-way one. Scanning covers the calendar. A scoped assessment covers the systematic classes of exposure that can be tested methodically. Novel business logic, unique to how your product works, still needs human creativity aimed at it. We are not going to put a percentage on that split, because any vendor who quotes you one is estimating from their own book rather than measuring your estate. Cadence decisions that collapse those three into one line item buy the wrong thing.

Neither of them, however, settles the dial that cadence decisions get wrong most often.

See what your external surface exposes, mapped to the controls it touches.

Run a free External Security Check →

Frequency Is Not Depth

Before the mechanics, one thing to take off the table. Continuous output is a trend line and a work queue, and both are operationally valuable, but the questionnaire asks for a third-party pentest in the last twelve months and a dashboard does not answer that question. That is a distinction about evidence rather than cadence, and it is treated in full in What Is CTEM?. For cadence purposes, take the conclusion and move on: whatever frequency you land on, designate one engagement a year as the one that carries the signature your compliance clock needs.

The limitation that actually distorts cadence budgets is a different one. Continuous programmes are optimised for breadth and recurrence, so they re-find the same class of issue efficiently and go deep rarely. Depth is a function of scope definition and tester hours, not of frequency. Testing something twelve times shallowly is not equivalent to testing it once thoroughly, and buying a higher frequency will not fix a scope drawn too narrowly in the first place.

Frequency and depth are separate dials. The programmes that cost the most and find the least are the ones that turned the first dial when the problem was the second, which is why the framework below scores your estate rather than your calendar.

The Four Inputs That Set Your Cadence

Here is the framework. Four inputs, each scored 1 to 5, with anchors written at 1, 3 and 5 so you can place yourself between them. Total them and you get a number between 4 and 20 that maps to a cadence profile. It takes about ten minutes and it produces something you can defend line by line in a budget meeting or an audit interview.

1
Change velocity. Score 1–5: How many production changes ship between tests? Score 1 if you deploy monthly or less. Score 3 if weekly (roughly 50 production changes per year between annual tests). Score 5 if daily or on-demand. This is the single highest-weight input.
2
Exposure. Score 1–5: How large is your internet-facing attack surface? Score 1 for a single marketing site with no authenticated paths. Score 3 for a multi-tenant SaaS with an API. Score 5 for a financial platform with public APIs, partner integrations, and a subdomain estate you do not fully control.
3
Regulatory floor. Score 1–5: What does your highest-obligation framework require? Score 1 for no current regulatory obligation. Score 3 for SOC 2 or Essential Eight. Score 5 for PCI DSS v4.0.1 or DORA.
4
Blast radius. Score 1–5: What is the cost of a breach? Score 1 if no regulated data is stored and no customer contracts reference security commitments. Score 3 if you hold PII and have enterprise customer contracts. Score 5 if you handle payment data, health records, or financial transactions for regulated entities.

Change velocity is the input most cadence discussions skip, and it is the one that most directly determines how stale a test becomes. A team shipping weekly puts roughly 52 production changes into the estate between two annual tests, and the State of DevOps research from the DevOps Research and Assessment group (confusingly also abbreviated DORA, and unrelated to the EU regulation above) puts elite performers well above that. On exposure, count honestly: production domains and subdomains, public API endpoints, third-party integrations that hold credentials to your environment, and any subdomain nobody has claimed ownership of. On the regulatory floor, note that a score of 5 also imports the change-trigger clause, not just the interval.

Scoring your own estate

Add the four numbers. The total lands between 4 and 20. Two rules keep it honest. First, no single input overrides the total: a regulatory 5 alone does not put you in the top profile, because compliance sets a floor and not a ceiling. Second, rescore after any structural change to the business (an acquisition, a platform migration, a new enterprise contract with security terms), because those move two or three inputs at once.

Input 2 is the one you cannot score from memory. The free check runs live against a domain you control and grades your external surface: HTTP security headers, TLS, DNS and attack surface. It grades that surface only, not your whole control set, and it cannot see inside your application. No signup required.
Run a free external security check

Four Cadence Profiles, and What Each One Should Actually Do

Every profile carries the same change trigger. The score sets the baseline interval; the written definition of significant change governs everything that happens between intervals. A profile without that definition is still only half a control, whatever the number says.

Profile Score Baseline cadence Signed report
1. Static estate 4 to 7 Annual Annual
2. Moderate velocity 8 to 11 Semi-annual Annual
3. High velocity, regulated 12 to 15 Quarterly Annual, plus retest evidence
4. Continuous deployment 16 to 20 Scheduled testing against a declared scope Annual, on the compliance clock
Who you are: Low change velocity (monthly or less), limited internet-facing exposure, no current regulatory trigger. Recommended cadence: Annual pentest plus defined change triggers. What to buy: One scoped, signed engagement per year. Define "significant change" in a written policy and book a targeted retest when one occurs. What to skip: Do not buy a continuous PTaaS subscription. The additional spend does not match the rate at which your estate changes. A second annual test is more cost-effective than a continuous subscription if you want more coverage.
🎯Key Takeaway
Cadence is set by change velocity and blast radius, not by what your framework's floor happens to be. The floor is a minimum, not a target.

What Cadence Costs, Honestly

Traditional mid-market engagements land between $15,000 and $30,000 for a scoped web application and external perimeter test, with a change order for retest that typically adds 15% to 30%. That retest line is the one that quietly breaks quarterly plans, because four tests plus four retests is not four line items.

PTaaS subscriptions run roughly $15,000 to $60,000 per year, with enterprise scopes above $100,000. Compare on researcher hours per quarter and on whether retest is included, not on the headline subscription price.

CyberOrbit assessments run $7,999 to $14,999, reviewed and signed by a certified security professional who is independent of your team, with a 90-day re-test included. Subscription tiers exist for teams buying a cadence rather than a single engagement. The practical effect for a mid-market team is arithmetic: a budget line that previously bought one annual test plus a retest change order buys two or three signed tests a year with retest included, which moves most Profile 2 and Profile 3 organisations onto their correct cadence without a budget increase. Full pricing detail is on our pricing page, and the market-wide numbers are broken down in our penetration testing cost guide.

The number to take into your budget meeting is not the price of one test. It is the annual cost of the cadence your four-input score returned, including retest. A quote that covers one engagement and prices retest separately is a quote for half a programme, so ask for the retest line before you compare two numbers.

If the framework returned Profile 2 or higher, a subscription tier turns a single-annual-test budget into two or three signed tests a year. See what that costs.
View CyberOrbit pricing and subscription tiers

Frequently Asked Questions

How often should you perform a penetration test?
At minimum annually, plus after every significant change to your infrastructure or applications. The right interval depends on four inputs: change velocity, internet-facing exposure, your regulatory floor, and blast radius. Score each 1 to 5 and the total maps to an annual, semi-annual, quarterly, or scheduled-continuous cadence.
How often does PCI DSS require penetration testing?
PCI DSS v4.0.1 requirements 11.4.2 and 11.4.3 require internal and external penetration testing at least once every twelve months and after any significant infrastructure or application upgrade or change. Requirement 11.4.4 additionally requires that exploitable vulnerabilities are corrected and that testing is repeated to verify the correction.
What counts as a significant change requiring a new penetration test?
No standard defines it for you, so you must define it in writing for your own estate. Common triggers: a new internet-facing service or subdomain, a change to authentication or authorisation, a new or materially changed public API, a production hosting migration, a change to network segmentation, or an acquisition touching production data.
Does SOC 2 require an annual penetration test?
No. The Trust Services Criteria never name penetration testing at a fixed interval. CC4.1 requires risk-based evaluations to confirm controls are present and functioning. Annual testing is a widely adopted convention that auditors accept, but what you owe is a documented rationale for your chosen cadence, not a specific number.
What is the difference between continuous penetration testing and an annual pentest?
An annual pentest is a scoped, signed engagement covering a defined set of dates. Continuous testing produces an ongoing trend and remediation queue across the calendar. The gap is mostly one of coverage: a ten-day engagement accounts for 2.7% of the year, so the other 97.3% is whatever else you have running. Most mature programmes run both, with one engagement a year designated to carry the signature. See What Is CTEM? for why the signed report is not replaceable.
Do you need a new penetration test after every deployment?
No. The trigger is significant change, not any change, and you define what significant means for your own architecture in writing. A routine feature deploy behind existing authentication is not a trigger. A new internet-facing endpoint, a change to an authentication or authorisation flow, a new payment path, a new cloud account, or an acquired asset absorbed into the estate is. Write six to ten such events, get them approved, and attach them to your testing policy so the trigger fires without a debate each time.
How much does continuous penetration testing cost per year?
Continuous scanning ranges from free open-source tooling to roughly $20,000 a year for commercial platforms. PTaaS with rotating human researchers typically runs $15,000 to $60,000 per year, with large enterprise scopes exceeding $100,000. Compare offers on researcher hours per quarter and whether retest is included, rather than on the subscription headline.
Is quarterly penetration testing worth it for a mid-market company?
It is worth it if your four-input score reaches 12 or above, which typically means weekly or faster deployment plus regulated data or contractual security commitments. Below that, quarterly testing tends to re-confirm known findings. A static, low-velocity estate is better served by annual testing plus written change triggers.

Setting Your Cadence This Quarter

Four steps, in order, none of which require a purchase decision to start.

First, establish what is currently exposed. You cannot score exposure from an architecture diagram, because diagrams document intent rather than reality. Run a free external security check against a domain you control for a graded view of your external surface today, and score Input 2 from that rather than from the diagram.

Second, define significant change in writing. Six to ten concrete, observable events specific to your architecture. Get it approved and attach it to your testing policy. This is the single highest-value hour in the whole process, because it converts half a compliance requirement into a complete one.

Third, book the next test against the profile the framework returned, not against last year's invoice. If the score moved you from Profile 1 to Profile 3, that is the number you defend to the board, with the four inputs as your reasoning.

Fourth, confirm which of those tests carries the signature your compliance clock needs, and put that date in the calendar first. Everything else in the programme is operational and can flex. That one cannot.

The cadence number is the easy half. The clause almost nobody writes down is the one that decides whether your programme catches the change that matters, so if you take one thing from this post into next week, make it the second step. Six to ten events, approved, attached to the policy.

If you need the exposure picture to start from, the free external security check maps what is reachable on a domain you control to the specific Essential Eight, ISO 27001 and SOC 2 controls each finding touches. It is an external check rather than an audit, so treat the output as the opening of the conversation with your assessor rather than a verdict on your control set.

The security writing, weekly

New posts as they land: findings from real assessments, what the regulatory changes actually mean, and the occasional teardown.

Privacy