The scoping call is twenty minutes old when the tester asks the only question that really matters: “Is there anything on this list we shouldn’t touch?” Everyone looks at everyone else. Nobody actually knows whether the Windows box running the warehouse label printer will survive a port scan, because nobody has scanned it since it was installed.
That silence is where pen test outages come from. Not from a tester going rogue, and rarely from anything exotic. It comes from a test firing at a system nobody warned them about, at a time nobody agreed, with no agreed way to make it stop.
So the honest answer to how organisations run a pen test without real side effects: they don’t rely on the testers being careful. They rely on three things. A written rules of engagement document signed by someone with actual authority. Stop conditions both sides accept before anyone touches a keyboard. And a contact path that puts a human on the phone with the test lead in under a minute, at 3am, on a Sunday. Paperwork matters more than skill here, because paperwork is what still works when your checkout has been returning 500s for eleven minutes.
Why pen tests break things in the first place
The failure modes are boringly predictable, which means you can write most of them out of the engagement in advance.
- Fragile legacy systems. Old embedded interfaces, unsupported appliances, the Java app that came with an acquisition. Some fall over on a malformed request a modern stack shrugs off. A scan that is trivial for your main platform is a hard restart for them.
- Scanners meeting stateful endpoints. A crawler does not know that a link cancels an order, voids an invoice or deletes a user. It knows the link exists. If a state-changing action is reachable by a GET request, a scanner will find it and use it repeatedly.
- Metered calls and outbound floods. Every automated form submission that fires an SMS, an email, a credit check or an AI API call costs money, and password resets and “notify me” endpoints turn a scanner into a spam cannon pointed at your own users. Ten thousand submissions overnight is ten thousand billable events plus a wrecked sender reputation.
- Account lockout. Credential testing against a directory with a low lockout threshold locks out real staff, including the people who would have fixed it.
- Services that fail under load by design. Anything single-threaded, anything opening a database connection per request, anything behind a licence-limited pool.
None of it requires an exploit. It is ordinary traffic arriving faster than expected, at the wrong endpoint, at the wrong hour. Which is also why “we’ll just run a scan instead” is not the safer option people assume: nobody is watching the target degrade.
The rules of engagement document is the actual control
Rules of engagement are the written agreement defining how a penetration test may be carried out: which assets are in scope, which techniques are permitted, when testing may run, what testers may do with data they encounter, and the conditions under which they must stop. The Penetration Testing Execution Standard puts it neatly: scope defines what will be tested, the rules of engagement define how. Both belong in the contract.
- Scope as an enumerated list. Specific IP ranges, hostnames, application URLs, cloud accounts by ID. “The corporate network” is not a scope.
- An explicit out-of-scope list. The half everyone skips. Name the warehouse printer, the legacy CRM, the OT segment, the shared hosting, the SaaS tools, the parent company’s IP space.
- Permitted and prohibited techniques. Permitted is whatever you are actually buying: unauthenticated and authenticated application testing, internal testing from a supplied jump host, configuration review. Prohibited is usually denial of service, anything destructive, modifying production data, persistence outliving the engagement, brute forcing above an agreed threshold, phishing staff unless separately scoped, and pivoting outside scope.
- Testing windows. Real dates and clock times with a named time zone. Some activities should be business-hours-only, so someone is awake to notice; others out-of-hours-only, so mistakes hit fewer customers. Those pull in opposite directions; choose per activity.
- Rate limits. A ceiling on requests per second per host, concurrent connections and login attempts per account per hour, plus an agreement to throttle further on request, immediately, without renegotiating anything.
- Stop conditions. Written triggers that halt testing: the target degrading, live customer or patient data somewhere unexpected, signs of a prior compromise by someone else, a finding severe enough that continuing adds risk rather than information, or a message from your side saying stop.
NIST SP 800-115, the 2008 technical guide most methodologies still lean on, carries a rules of engagement template in its appendices and recommends the two dullest risk reducers available: test duplicates of production where you can, and test outside business hours where you cannot.
Who is actually allowed to sign this
Authorisation has to come from someone who can lawfully give it: an officer or director of the entity owning the systems, or someone with a documented delegation. A security manager’s verbal nod is not a defence for the tester, and it does not protect you if the systems turn out to belong to a subsidiary or a customer. PTES calls the permission-to-test document one of the most important artefacts in the engagement and is blunt that testing does not begin until the customer signs it. Unauthorised access to a computer system is a criminal matter in most jurisdictions, and the only thing separating a pen tester from a defendant is a signature from someone entitled to give it.
The part that catches people out is infrastructure you do not own. Your consent does not cover your hosting provider, your SaaS vendors, your payment processor or your MSP’s network. PTES makes the point directly: the client may have granted permission, but they do not speak for their third-party providers. Check the contract, then the provider’s published policy, then get it in writing.
What the three big cloud providers actually require
AWS lets customers assess their own resources on a published service list without prior approval, currently covering EC2, load balancers, WAF, RDS and Aurora, CloudFront, API Gateway, Lambda, Elastic Beanstalk, ECS and Fargate and others named there. You may not assess AWS infrastructure or the AWS services themselves. Prohibited outright: denial of service and DDoS whether real or simulated, port and request flooding, DNS zone walking or hijacking via Route 53, and S3 bucket or subdomain takeover. Another set needs a Simulated Events form at least two weeks ahead: DDoS simulation, network stress and iPerf testing, red, blue or purple team exercises using command and control, simulated phishing, and malware testing.
Microsoft dropped Azure pre-approval in June 2017 and does not require notification either. It does require compliance with the Microsoft Cloud Unified Penetration Testing Rules of Engagement, which prohibits denial of service testing in all circumstances, network-intensive fuzzing generating excessive traffic, using credentials that are not yours, touching tenants or storage accounts you do not own, and phishing or social engineering against Microsoft. DDoS resilience testing goes through approved simulation partners. Third parties testing for you need your explicit written authorisation, and Microsoft’s advice is to keep that document to hand, because Azure’s abuse detection will flag your testing and you will be asked to explain it.
Google Cloud takes the lightest position. Its published FAQ says that if you plan to evaluate the security of your Cloud Platform infrastructure with penetration testing, you are not required to contact them, provided you stay inside the Acceptable Use Policy and Terms of Service and affect only your own projects. Vulnerabilities in Google’s own products go to the Vulnerability Reward Program, not support.
All three let you test what you own, none let you test them, and all three treat denial of service as a separate category needing a specific process or a flat no.
The stop button, and who you tell
Most emergency procedures fail on the phone number. The ROE lists a contact, the contact is the account manager, the account manager is asleep, and your network team spends forty minutes deciding whether to just block the traffic. Build it properly. Name a test lead and a deputy with two forms of round-the-clock contact each, which is what PTES recommends and it is right. Name the same on your side, and collect the source addresses testing will come from so you can tell test traffic from the real thing. Open a shared channel and agree that a message saying “stand down on 10.4.x” is binding immediately, written confirmation to follow. Ring both numbers at kickoff to prove they work. Set the default: when in doubt, stop first and discuss after. Then agree who can pull the plug. If the only person authorised to halt a test is the CISO and the CISO is on a plane, you have not built a stop button, you have built a suggestion.
Then decide who knows. Deconfliction is how your defenders find out, quickly, whether what they are investigating is the authorised test. Without it your SOC does the right thing and treats it as a real intrusion: isolating hosts, disabling accounts, calling the incident response retainer, maybe the insurer and the regulator. Three options:
- Announced. Everyone knows. Fastest, cheapest, safest, and correct when what you want is a list of vulnerabilities. NIST notes the benefit of overt testing: staff who know can steer testers away from fragile systems.
- Partially announced. Two or three named people know, typically the SOC manager and infrastructure lead, while the analysts do not. You get a real read on detection while keeping someone who can say “that one is us” before containment runs. For most organisations this is the sweet spot.
- Unannounced. Nobody operational knows. This tests response end to end, and it is the version that generates the 2am call to your CEO. Buy it once measuring the response is the point, give a few executives an out-of-band deconfliction number, and accept that you are paying for realism with disruption.
Staging or production, and the tradeoff nobody likes
Testing in staging is safer and the results are worth less. That is the whole argument, and pretending otherwise helps nobody. Staging drifts: different WAF rules, different IAM roles, a sandbox payment provider, stubbed integrations, a thousandth of the data so timing and enumeration behave differently, and last quarter’s production hardening never merged back. A clean staging report says the code is probably fine. It says little about the environment, where a large share of real findings live.
So test the application in production under a tight ROE, with rate limits, a working stop button and dedicated test accounts, and push the destructive categories into staging or a bench rig. If everything must happen in staging, say plainly in the report that environment-level findings were not assessed.
Where “we’ll be careful” is not good enough
Denial of service. A separate exercise with its own contract, provider permission and window. AWS and Microsoft both prohibit it under ordinary testing terms, which settles it for most people.
Social engineering of employees. The governance here is as much HR as security. Agree themes in advance and rule out the cruel ones: fake redundancy notices, fake payroll changes, fake health results, anything impersonating a colleague in distress. Report results as aggregate rates, not a list of names for managers to act on, and involve HR plus works councils or unions where they exist. A test that teaches staff reporting gets them punished has made you less safe.
Physical testing. Testers carry a signed authorisation letter naming them, listing sites and dates, with a 24/7 number a guard or a police officer can ring and get an immediate answer from a real person. Someone senior in facilities holds that number and expects the call.
OT, ICS and medical devices. Usually the answer is don’t. Active testing against live plant, safety instrumented systems or connected clinical devices turns an IT incident into a physical safety incident, and much of this equipment behaves unpredictably under traffic it was never designed to see. Do architecture review, configuration review and passive analysis on the live environment. Save active testing for a bench rig or vendor lab, or a maintenance window with the vendor present and engineering sign-off in writing.
Evidence, data handling and the clauses worth paying a lawyer for
Decide before the test what counts as proof. The principle that keeps you out of trouble: testers demonstrate access without removing the data. A screenshot of a record count, one redacted record, a file hash, a query returning the schema rather than the contents. Not a dumped table “for the report”.
Put the rest in the contract: where evidence is stored and in which country, encryption at rest, who can access it, and a retention period tied to the retest, followed by secure deletion and written confirmation. Keep personal, card and health data out of the report body, and handle the report itself as confidential, because it is a ready-made target list.
On liability, the clauses that earn their keep: professional indemnity and cyber liability cover at a limit bearing some relation to what an outage would cost you, evidenced by a current certificate before work starts; mutual confidentiality; approval of subcontractors; the vetting standard applied to individual testers; and a realistic split of risk, where the tester carries negligence and ROE breaches while you accept that testing carries inherent risk of instability. Any provider promising zero risk of disruption is either not going to test hard or is not being straight with you. If something does break, restore service first and preserve evidence second: good firms keep timestamped request logs precisely so causation can be settled later.
After the test, and when testing never stops
Debrief within a few days, with engineers in the room and not just managers. Agree severities jointly rather than accepting the tool’s rating. Get the retest included in the original price, so remediation does not become a second procurement. Then review the test itself: what nearly went wrong, which asset was missing from the inventory, whether the stop button was ever exercised. The UK’s NCSC is right that a penetration test validates only that you are not vulnerable to known issues on the day of the test. A sample, not a guarantee.
Which is why many organisations add continuous automated testing or a bug bounty. Both change the safety model. With continuous scanning nobody approves each run, so the ROE becomes standing configuration: permanent rate limits, an out-of-scope asset list someone owns and updates, new assets excluded by default until reviewed, and alerting when the scanner starts hitting something new. With a bounty you also lose control of who tests and when, so the published policy becomes your entire control surface: a safe harbour statement, explicit scope and out-of-scope lists, a ban on social engineering and denial of service, registered test accounts and an identifying header so researcher traffic is separable, and a triage team that answers in hours. Without that scaffolding, a bounty produces exactly the incidents this article is about, with no contract and no phone number to call.
The through line is the same in all three. You are not asking anyone to be careful. You are writing down what may happen, who agreed to it, and how to make it stop.
If you are scoping a VAPT engagement and want the rules of engagement, authorisation chain and stop conditions written properly first, or you need SOC coverage that can tell authorised test traffic from a real intrusion, AB7 Solutions works on both sides of that line. Call +1 321 341 7733, email ab@ab7solutions.com or director@ab7solutions.com, or start at www.ab7solutions.com and tell us what the test is meant to prove. That settles scope faster than a list of IP ranges does.
Sources: AWS Customer Support Policy for Penetration Testing; Microsoft Azure penetration testing guidance and the Microsoft Cloud Unified Penetration Testing Rules of Engagement; Google Cloud Security FAQ; NIST SP 800-115, Technical Guide to Information Security Testing and Assessment (2008); Penetration Testing Execution Standard, Pre-engagement Interactions; UK NCSC guidance on penetration testing.