IAM protects whatever has an identity. But the visitor crawling AlpinaShop's catalogue at a thousand requests per second to copy the prices has no identity. Nor does the one trying ten thousand passwords against the sign-in form, nor the one typing ' OR 1=1-- into the search box to see what happens. That traffic arrives at alpinashop-lb-ip like any legitimate customer and, with nothing in front, ends up on the MIG instances.

Cloud Armor is Google Cloud's application firewall (WAF) and denial-of-service protection system. Like Cloud CDN, it is not a product deployed separately: it is a policy attached to the backend service of the load balancer you built in 03-02, and it is applied at the edge of Google's network, at the same points of presence where the load balancing and the caching happen. Malicious traffic is dropped thousands of kilometres from europe-west1, without consuming a single cycle of your instances.

This lesson has a concrete case: AlpinaShop's autumn campaign. Marta spots two simultaneous problems. A competitor is crawling the entire catalogue every night, which triggers the MIG's autoscaling and sinks the CDN hit ratio. And the sign-in form is receiving brute-force attempts from several hundred addresses. We are going to build the pol-catalogo-web policy that solves both, and above all we are going to roll it out the right way, which consists of blocking nothing until we are sure.

Important warning. A WAF is no substitute for secure code, and a badly calibrated security policy blocks legitimate customers, which is a commercial incident just as real as the one you were trying to avoid. Everything that follows is a valid teaching model, but before applying blocking rules in a real production environment, the design must be reviewed by a security professional, and any geography-based filtering decision must also be validated by the legal and commercial lead.

Contents

  1. What a WAF is, what it covers and what it does not
  2. Where Cloud Armor acts: a request's journey
  3. Policies, rules, priorities and the default rule
  4. Rules by IP address
  5. Rules by geography and their implications
  6. Preconfigured OWASP Top 10 rules
  7. False positives: a WAF's real problem
  8. The expression language: bespoke rules for AlpinaShop
  9. Preview mode: the only correct way to deploy
  10. Rate limiting and rate-based-ban
  11. Adaptive protection against layer 7 DDoS
  12. reCAPTCHA Enterprise for telling people from bots
  13. The Standard and Enterprise tiers, and cost
  14. The complete autumn campaign policy

  1. What a WAF is, what it covers and what it does not

A web application firewall inspects the content of HTTP requests — path, headers, parameters, body — and decides whether to let them through. It is different from the VPC firewall you configured in 03-01: that one works with IP addresses and ports and has no idea what is travelling inside; this one understands HTTP.

Layer Tool Decides using Example decision
Network (3/4) VPC firewall (03-01) Source IP, port, protocol "Port 5432 only from sa-catalogo-web"
Application (7) Cloud Armor Path, headers, parameters, body, geography, rate "This request contains a SQL injection pattern"
Identity IAM / IAP (03-04) Who you are "Only gcp-datos@ sees the reporting dashboard"

What Cloud Armor covers well:

  • Injection attacks in the parameters: SQLi, XSS, local and remote file inclusion, remote command execution.
  • Volumetric denial of service (layers 3 and 4), automatically and with no configuration, simply by virtue of sitting behind the global load balancer.
  • Application denial of service (layer 7): mass crawling, brute force, abuse of expensive endpoints.
  • Filtering by origin: addresses, ranges, countries, IP reputation.
  • Basic bots and automated scanners.

What it does not cover, which is just as important to know:

  • Business logic flaws. If your endpoint lets someone request another customer's order by changing an identifier in the URL, every request is perfectly legitimate in the WAF's eyes. That flaw is fixed in the code.
  • Stolen credentials. A sign-in with the correct password is a correct sign-in.
  • Vulnerabilities in your dependencies. It can mitigate the exploitation of some known ones, and it buys you time while you patch, but it patches nothing.
  • Permission misconfigurations. That is IAM.
  • Attacks from inside or that do not pass through the load balancer. If a VM has a public IP and can be reached directly, the WAF never knows.

A WAF is a mitigation layer, not an absolution. It is what gives you room to fix things, not what fixes them.

  1. Where Cloud Armor acts: a request's journey

flowchart TB
    A([Malicious request<br/>GET /buscar?q=' OR 1=1--])
    B["Nearest Google PoP<br/>Madrid, Lisbon, Frankfurt…"]
    C{"Cloud Armor<br/>pol-catalogo-web"}
    D["Cloud CDN<br/>cache lookup"]
    E["URL map<br/>alpinashop-url-map"]
    F["bs-catalogo-web → MIG"]
    G(["403 Forbidden<br/>from the edge"])

    A --> B --> C
    C -->|"matches rule 1000<br/>action: deny-403"| G
    C -->|"no match: allow"| D
    D -->|hit| A
    D -->|miss| E --> F

The four facts that matter in this diagram:

  1. Cloud Armor is evaluated at the edge, at the same point of presence the client enters through. The malicious packet never travels to europe-west1.
  2. It is evaluated before the cache. A blocked request neither consults the cache nor consumes fill; and conversely, enabling the CDN does not create a path by which traffic can bypass the WAF.
  3. It is evaluated before IAP. The bot attacking the reporting dashboard is dropped before it even reaches the sign-in screen.
  4. The blocking response is generated by the edge. Your instances never see the request, do not record it in their logs and burn no CPU.

Policies are attached to a backend service:

gcloud compute backend-services update bs-catalogo-web --global \
  --security-policy=pol-catalogo-web

For backend buckets like bb-catalogo-imagenes there is a variant called an edge security policy (--edge-security-policy), which supports fewer rule types — filtering by IP, geography and some basic expressions — but is also applied to content served from cache. It is the way of protecting the images without giving up the CDN.

  1. Policies, rules, priorities and the default rule

A security policy is an ordered container of rules. Each rule has:

  • A priority (an integer). They are evaluated from lowest to highest and the first one that matches wins; the rest are not even looked at.
  • A condition: either a list of IP ranges, or an expression.
  • An action: allow, deny-403, deny-404, deny-502, throttle, rate-based-ban, redirect.
gcloud config set project alpinashop-prod

gcloud compute security-policies create pol-catalogo-web \
  --description="Protection for the AlpinaShop public catalogue"

When you create it, Google automatically adds the default rule, with priority 2147483647 (the largest signed 32-bit integer) and the allow action. It is the one that catches everything that matched nothing earlier.

That rule raises the lesson's most important design decision:

Model Default rule Advantage Drawback
Blocklist (default allow) Allow Nothing breaks; you block what you identify It only protects against what you have foreseen
Allowlist (default deny) Deny Maximum security Unworkable on a public website

For AlpinaShop's public catalogue the model is necessarily a blocklist: anyone in the world must be able to view a backpack. For an internal API or an admin dashboard, an allowlist is viable and is the right thing.

Always leave a gap in the priorities. Number in hundreds (1000, 1100, 1200…) so you can insert rules in between without renumbering everything:

gcloud compute security-policies rules list --security-policy=pol-catalogo-web \
  --format="table(priority, action, preview, description)"

  1. Rules by IP address

The simplest rule and the one that must always go first: AlpinaShop's office and the load testing provider must never be blocked by anything, not even by a badly calibrated OWASP rule.

# Priority 100: the office ALWAYS gets through. Before everything else.
gcloud compute security-policies rules create 100 \
  --security-policy=pol-catalogo-web \
  --src-ip-ranges=203.0.113.24/29 \
  --action=allow \
  --description="AlpinaShop office: exempt from all rules"

# Priority 200: permanent block of confirmed abusive addresses
gcloud compute security-policies rules create 200 \
  --security-policy=pol-catalogo-web \
  --src-ip-ranges=198.51.100.0/24,192.0.2.77/32 \
  --action=deny-403 \
  --description="Ranges with confirmed abuse, reviewed 2026-08"

Practical details:

  • Up to several thousand ranges per policy are supported, but maintaining huge lists by hand is unsustainable. That is what the Enterprise tier's threat intelligence is for (section 13).
  • Blocking by IP has a shelf life. Home addresses are dynamic: today's attacker address is tomorrow's customer. Document the review date in the description, as in the example, and review those lists.
  • Behind a proxy or another CDN, origin.ip is the proxy's. If that is your case, the correct expression uses the IP from the trusted X-Forwarded-For; configure it carefully, because that header can be forged by the client if it is not validated properly.

  1. Rules by geography and their implications

Cloud Armor knows the country each request comes from and exposes it as origin.region_code, with two-letter ISO codes.

# Block countries AlpinaShop neither sells to NOR receives legitimate visits from
gcloud compute security-policies rules create 300 \
  --security-policy=pol-catalogo-web \
  --expression="origin.region_code == 'XX' || origin.region_code == 'YY'" \
  --action=deny-403 \
  --description="Countries outside the market; review with the commercial lead"

And a far more reasonable variant: instead of blocking the whole catalogue, restrict only the admin dashboard, which does have a perfectly well-defined legitimate origin.

gcloud compute security-policies rules create 400 \
  --security-policy=pol-catalogo-web \
  --expression="request.path.startsWith('/admin') && origin.region_code != 'ES'" \
  --action=deny-404 \
  --description="Admin dashboard only from Spain"

Note deny-404 instead of deny-403: we do not confirm to the attacker that an admin dashboard exists. It is a small difference that reduces the information you give away.

The three warnings about geographic blocking, and none of them is technical:

  1. It is trivial to bypass. A three-euro-a-month VPN nullifies it. It stops automated noise, not a determined attacker.
  2. It blocks real customers. A Spanish customer on holiday, an expatriate, somebody behind a corporate VPN. In a shop, every block is a lost sale and a call to customer service.
  3. It can have legal and commercial implications. Blocking countries is a business decision, and in some contexts a regulatory compliance one. The technical team must not take it alone.

The recommended stance: use geography to restrict sensitive paths (/admin, /api/interna) and to score risk in combination with other signals, not to close the shop to a continent.

  1. Preconfigured OWASP Top 10 rules

Here is the main value of a managed WAF. Google maintains rule sets based on the ModSecurity Core Rule Set, updated by its security team, invoked with the evaluatePreconfiguredWaf() function.

Set Protects against False positive risk
sqli-v33-stable SQL injection High: search boxes and free-text fields trigger many rules
xss-v33-stable Cross-site scripting High: any field accepting HTML or quotes
lfi-v33-stable Local file inclusion (../../etc/passwd) Medium
rfi-v33-stable Remote file inclusion Low
rce-v33-stable Remote command execution Medium
scannerdetection-v33-stable Automated scanners (Nikto, sqlmap, Nessus) Low
protocolattack-v33-stable Request smuggling, HTTP response splitting Low
methodenforcement-v33-stable Unexpected HTTP methods Low
sessionfixation-v33-stable Session fixation Low
php-v33-stable, nodejs-v33-stable, java-v33-stable Attacks specific to those stacks Low, and unnecessary on a Python stack
cve-canary Recent critical vulnerabilities (Log4Shell and the like) Low. Always enable it

Creating the rules, with sensitivity and always in preview the first time:

# SQLi at low sensitivity: only the most obvious patterns
gcloud compute security-policies rules create 1000 \
  --security-policy=pol-catalogo-web \
  --expression="evaluatePreconfiguredWaf('sqli-v33-stable', {'sensitivity': 1})" \
  --action=deny-403 \
  --preview \
  --description="OWASP: SQL injection, sensitivity 1"

gcloud compute security-policies rules create 1100 \
  --security-policy=pol-catalogo-web \
  --expression="evaluatePreconfiguredWaf('xss-v33-stable', {'sensitivity': 1})" \
  --action=deny-403 --preview \
  --description="OWASP: XSS, sensitivity 1"

gcloud compute security-policies rules create 1200 \
  --security-policy=pol-catalogo-web \
  --expression="evaluatePreconfiguredWaf('lfi-v33-stable', {'sensitivity': 1})" \
  --action=deny-403 --preview \
  --description="OWASP: local file inclusion"

gcloud compute security-policies rules create 1300 \
  --security-policy=pol-catalogo-web \
  --expression="evaluatePreconfiguredWaf('rce-v33-stable', {'sensitivity': 1})" \
  --action=deny-403 --preview \
  --description="OWASP: remote command execution"

gcloud compute security-policies rules create 1400 \
  --security-policy=pol-catalogo-web \
  --expression="evaluatePreconfiguredWaf('scannerdetection-v33-stable', {'sensitivity': 1})" \
  --action=deny-403 --preview \
  --description="OWASP: scanner detection"

# Recent critical vulnerabilities: this one goes straight to blocking
gcloud compute security-policies rules create 900 \
  --security-policy=pol-catalogo-web \
  --expression="evaluatePreconfiguredWaf('cve-canary', {'sensitivity': 1})" \
  --action=deny-403 \
  --description="Recent critical CVEs"

Sensitivity runs from 1 to 4 and works as a cumulative threshold:

Level What it detects False positives Recommendation
1 Only the clearest attack patterns Few Start here. Always.
2 Adds probable patterns Noticeable Only after debugging level 1
3 Adds suspicious patterns Many High-risk sites, with dedicated effort
4 Everything, including the doubtful Unworkable in production Forensic analysis

A warning that saves a lot of time: if your application is Python with Flask and PostgreSQL, do not enable php-v33-stable, nodejs-v33-stable or java-v33-stable. They add no protection and every active rule costs money and adds false positive probability.

  1. False positives: a WAF's real problem

This is the part the brochures do not tell you. The OWASP rules work by looking for patterns, and many legitimate patterns look like an attack. Real cases that occur in a shop like AlpinaShop:

Perfectly legitimate request Rule that fires Why
Searching for mochila 40l "trekking" sqli The double quotes
Searching for pantalón O'Neill sqli The apostrophe
A product description with <b>Ligera</b> xss HTML tags
A review containing 1=1 en cuanto al peso sqli The 1=1 pattern
Uploading an image named ../foto.jpg lfi The ../ sequence
A JSON with select as a field name sqli A SQL keyword

There are three ways of solving it, in order of preference:

1. Disable specific rules from the set, instead of lowering the global sensitivity:

gcloud compute security-policies rules update 1000 \
  --security-policy=pol-catalogo-web \
  --expression="evaluatePreconfiguredWaf('sqli-v33-stable', {'sensitivity': 1, 'opt_out_rule_ids': ['owasp-crs-v030301-id942432-sqli', 'owasp-crs-v030301-id942431-sqli']})"

The rule identifiers come from the log: every block records exactly which rule in the set fired. This is surgery, as opposed to the axe of lowering the sensitivity.

2. Exclude specific paths from the analysis, when you know an endpoint receives free text by design:

gcloud compute security-policies rules create 950 \
  --security-policy=pol-catalogo-web \
  --expression="request.path.startsWith('/api/resenas') && request.method == 'POST'" \
  --action=allow \
  --description="Reviews: free text, exempt from the WAF. Validated in the app."

With one very serious condition: if you exclude a path from the WAF, that path has to be impeccably validated in the code. You are giving up the safety net exactly where arbitrary text comes in.

3. Lower the sensitivity. It is the fastest and the least precise. Valid as a temporary measure during an incident.

And the advice that covers all three: the correct procedure is not to react to false positives in production, but to discover them beforehand with preview mode.

  1. The expression language: bespoke rules for AlpinaShop

Cloud Armor uses CEL (Common Expression Language), the same language as the IAM conditions you saw in 03-04. The available attributes:

Attribute Example
origin.ip inIpRange(origin.ip, '203.0.113.0/24')
origin.region_code origin.region_code == 'ES'
request.path request.path.startsWith('/admin')
request.method request.method == 'POST'
request.query request.query.contains('debug=1')
request.headers['name'] request.headers['user-agent'].contains('curl')
request.headers['host'] request.headers['host'] == 'alpinashop.example'
has(request.headers['x']) Checking for the presence of a header
request.path.matches('regex') Regular expression (RE2)

Case 1: the bot hammering the catalogue. The competitor crawls with a client that identifies itself and does not run JavaScript. It is recognised by the combination of user agent and absence of real browser headers:

gcloud compute security-policies rules create 2000 \
  --security-policy=pol-catalogo-web \
  --expression="request.path.startsWith('/producto/') && (
      request.headers['user-agent'].contains('python-requests') ||
      request.headers['user-agent'].contains('Scrapy') ||
      request.headers['user-agent'].contains('HeadlessChrome') ||
      !has(request.headers['accept-language'])
    )" \
  --action=deny-403 --preview \
  --description="Automated crawling of the catalogue"

The !has(request.headers['accept-language']) condition is the most useful of the four: practically every real browser sends that header and practically no simple script bothers to set it. And it is precisely the one that should spend the longest in preview, because some legitimate clients omit it too.

Case 2: protecting an expensive endpoint. The CSV catalogue export takes seconds and hits the database. It should only be used from the office:

gcloud compute security-policies rules create 2100 \
  --security-policy=pol-catalogo-web \
  --expression="request.path.startsWith('/exportar') && !inIpRange(origin.ip, '203.0.113.24/29')" \
  --action=deny-403 \
  --description="CSV export: from the office only"

Case 3: rejecting methods the application does not use. The catalogue only serves GET, HEAD and POST:

gcloud compute security-policies rules create 2200 \
  --security-policy=pol-catalogo-web \
  --expression="request.method == 'TRACE' || request.method == 'TRACK' || request.method == 'CONNECT'" \
  --action=deny-403 \
  --description="HTTP methods not used by the application"

Case 4: preventing access by direct IP. Every legitimate request arrives with the correct Host; those arriving with the IP in the Host are indiscriminate scanning:

gcloud compute security-policies rules create 2300 \
  --security-policy=pol-catalogo-web \
  --expression="request.headers['host'] != 'alpinashop.example' && request.headers['host'] != 'www.alpinashop.example' && request.headers['host'] != 'imagenes.alpinashop.example'" \
  --action=deny-404 \
  --description="Requests with no valid Host: scanning"

  1. Preview mode: the only correct way to deploy

Every rule accepts --preview. In that mode, the rule is evaluated and logged, but not enforced. The request carries on as normal.

This is not a convenience: it is the procedure. An OWASP rule with a badly chosen sensitivity can block your shop's search box on the biggest sales day of the year, and the worst part is that you will not find out through an alert, but through a drop in conversion that nobody connects with the WAF.

The deployment procedure, step by step:

# STEP 1 - Create ALL the new rules in preview
gcloud compute security-policies rules create 1000 \
  --security-policy=pol-catalogo-web \
  --expression="evaluatePreconfiguredWaf('sqli-v33-stable', {'sensitivity': 1})" \
  --action=deny-403 --preview

# STEP 2 - Enable detailed logging on the policy
gcloud compute security-policies update pol-catalogo-web \
  --log-level=VERBOSE

# STEP 3 - Wait. At least one week of real traffic,
#          including a weekend and a campaign day.

# STEP 4 - Analyse WHAT it would have blocked
gcloud logging read '
  resource.type="http_load_balancer"
  AND jsonPayload.previewSecurityPolicy.outcome="DENY"
' --project=alpinashop-prod --freshness=7d --limit=100 \
  --format="table(
    jsonPayload.previewSecurityPolicy.priority,
    httpRequest.requestUrl,
    httpRequest.remoteIp,
    jsonPayload.previewSecurityPolicy.preconfiguredExprIds
  )"

The previewSecurityPolicy.preconfiguredExprIds field is the one that gives the specific rule identifiers from the OWASP set, which are what you need for the opt_out_rule_ids in section 7.

An aggregate count, so you can decide with numbers instead of impressions:

gcloud logging read '
  resource.type="http_load_balancer"
  AND jsonPayload.previewSecurityPolicy.outcome="DENY"
' --project=alpinashop-prod --freshness=7d --limit=1000 \
  --format="value(jsonPayload.previewSecurityPolicy.priority)" | sort | uniq -c | sort -rn
    847 1000     ← SQLi: nearly all searches with apostrophes. False positives.
     93 2000     ← Crawling: review case by case.
      6 1200     ← LFI: real attacks.
      1 900      ← CVE: a real attack.

That count reads like this: rule 1000 as it stands would block 847 legitimate searches in a week. It is not enabled: it is debugged with opt_out_rule_ids or the search path is excluded, and it goes back into preview for another week.

# STEP 5 - Enable for real, rule by rule, starting with the clean ones
gcloud compute security-policies rules update 1200 \
  --security-policy=pol-catalogo-web --no-preview

# STEP 6 - Watch the traffic actually blocked for 24-48 h
gcloud logging read '
  resource.type="http_load_balancer"
  AND jsonPayload.enforcedSecurityPolicy.outcome="DENY"
' --project=alpinashop-prod --freshness=1d --limit=50 \
  --format="table(timestamp, jsonPayload.enforcedSecurityPolicy.priority, httpRequest.requestUrl)"

The distinction between the two log fields is the key to this whole section:

Field Meaning
jsonPayload.previewSecurityPolicy What would have happened. The rule is in preview
jsonPayload.enforcedSecurityPolicy What has happened. The rule is active

And an operational golden rule: never enable new rules the day before a campaign. AlpinaShop's autumn campaign starts on 15 September; the rules go into preview on 1 August and are enabled by 1 September at the latest.

  1. Rate limiting and rate-based-ban

Marta's second problem is brute force against /acceso. Here there is no malicious pattern to detect: each individual request is legitimate. What is anomalous is the frequency.

Cloud Armor offers two actions:

  • throttle: above the threshold, the excess requests are rejected, but as soon as the rate drops, service resumes. It is a continuous limiter.
  • rate-based-ban: above the threshold, that client is banned for a fixed period, even if it stops trying. It is the punishment.
# Sign-in form: 5 attempts per minute per IP; if exceeded, a 10-minute ban
gcloud compute security-policies rules create 3000 \
  --security-policy=pol-catalogo-web \
  --expression="request.path.startsWith('/acceso') && request.method == 'POST'" \
  --action=rate-based-ban \
  --rate-limit-threshold-count=5 \
  --rate-limit-threshold-interval-sec=60 \
  --ban-duration-sec=600 \
  --conform-action=allow \
  --exceed-action=deny-429 \
  --enforce-on-key=IP \
  --description="Brute force against the sign-in form"

# Search box: expensive on the database. 30 searches per minute per IP, continuous throttling
gcloud compute security-policies rules create 3100 \
  --security-policy=pol-catalogo-web \
  --expression="request.path.startsWith('/buscar')" \
  --action=throttle \
  --rate-limit-threshold-count=30 \
  --rate-limit-threshold-interval-sec=60 \
  --conform-action=allow \
  --exceed-action=deny-429 \
  --enforce-on-key=IP \
  --description="Search limit per IP"

# A very generous site-wide floor: 600 requests per minute
gcloud compute security-policies rules create 3200 \
  --security-policy=pol-catalogo-web \
  --expression="true" \
  --action=throttle \
  --rate-limit-threshold-count=600 \
  --rate-limit-threshold-interval-sec=60 \
  --conform-action=allow \
  --exceed-action=deny-429 \
  --enforce-on-key=IP \
  --description="General limit per IP"

The grouping key (--enforce-on-key) decides who the counting is done for, and choosing it badly invalidates the protection:

Key Counts by When to use it Risk
IP Source address The general case A school or a company shares one IP: everyone is penalised
ALL All traffic together Protecting a fragile endpoint with a global ceiling It does not discriminate: one attacker shuts everyone out
HTTP_HEADER The value of a header APIs with a client key The client controls the header and can rotate it
HTTP_COOKIE The value of a cookie Sessions Bypassed by deleting the cookie
XFF_IP The first IP in X-Forwarded-For Behind another proxy Forgeable if the proxy is not validated
REGION_CODE Country Geographically concentrated attacks Very coarse
HTTP_PATH Path Protecting each path separately —

Two calibration tips worth more than any rule:

  • Measure before choosing a number. The 99th percentile of requests per minute per IP for your real traffic is in the load balancer log. Set the threshold at two or three times that value, not at a round number picked by eye.
  • deny-429 and not deny-403. The 429 code (Too Many Requests) is the semantically correct one, legitimate clients know to retry with a wait, and search engines read it as "come back later" instead of "this is forbidden", which protects your ranking.

  1. Adaptive protection against layer 7 DDoS

The previous rules depend on somebody having foreseen the attack. Adaptive protection does not: it trains models on your application's normal traffic and detects anomalous deviations, proposing concrete rules.

gcloud compute security-policies update pol-catalogo-web \
  --enable-layer7-ddos-defense \
  --layer7-ddos-defense-rule-visibility=STANDARD

How it works in practice:

  1. It needs between several days and a week of traffic to establish a baseline. Enable it long before you need it.
  2. When it detects an anomaly, it generates an alert with the attack's signature: user agents involved, source ranges, paths attacked, and a confidence score.
  3. It suggests a Cloud Armor rule ready to apply, with an estimated impact on legitimate traffic.
  4. The decision to apply it is still yours. It can be applied automatically, but in a shop human review is preferable.

It is worth being clear about what it protects and what it does not: volumetric layer 3 and 4 attacks (SYN floods, UDP amplification) are absorbed automatically by Google's infrastructure simply by virtue of sitting behind the global load balancer, with nothing to configure. Adaptive protection is for layer 7 attacks, the ones consisting of perfectly formed HTTP requests in anomalous quantity and pattern, which are far harder to tell apart from real traffic.

  1. reCAPTCHA Enterprise for telling people from bots

When the bot is sophisticated — it runs JavaScript, rotates IP addresses, sends real browser headers — no static rule tells it apart. The answer is reCAPTCHA Enterprise, natively integrated with Cloud Armor.

The flow is interesting because it is not the classic "select the traffic lights" CAPTCHA:

  1. The application includes the reCAPTCHA script, which assesses the user's behaviour in the background and issues a token with a score from 0.0 (almost certainly a bot) to 1.0 (almost certainly a person).
  2. Cloud Armor reads that token at the edge and acts according to the score.
  3. Only if the score is doubtful is a visible challenge shown. Most users never see anything.
# 1) Associate the reCAPTCHA key with the policy
gcloud compute security-policies update pol-catalogo-web \
  --recaptcha-redirect-site-key="projects/alpinashop-prod/keys/CLAVE_RECAPTCHA"

# 2) Low score during checkout: a challenge, not a block
gcloud compute security-policies rules create 4000 \
  --security-policy=pol-catalogo-web \
  --expression="request.path.startsWith('/pago') && token.recaptcha_session.score < 0.5" \
  --action=redirect \
  --redirect-type=GOOGLE_RECAPTCHA \
  --description="Suspected bot at payment: reCAPTCHA challenge"

The design decision here is redirect and not deny: when in doubt, in a shop you never block, you ask. A false positive with deny is a lost sale; with redirect, a second of friction. That criterion — block the obvious, challenge the doubtful — is what separates a usable security policy from one that loses money.

  1. The Standard and Enterprise tiers, and cost

Cloud Armor Standard Cloud Armor Enterprise
Model Pay as you go Subscription (monthly or annual)
Rules by IP, geography and expressions Yes Yes
Preconfigured OWASP rules Yes Yes
Rate limiting Yes Yes
Layer 3/4 DDoS protection Yes, automatic Yes
Adaptive protection (layer 7) Limited Complete
Threat intelligence No Yes (evaluateThreatIntelligence)
Bill protection against DDoS No Yes
Specialist support during an attack No Yes

Threat intelligence deserves a mention because it solves section 4's problem: instead of maintaining IP lists by hand, Google maintains categorised, up-to-date lists:

# Cloud Armor Enterprise only
gcloud compute security-policies rules create 500 \
  --security-policy=pol-catalogo-web \
  --expression="evaluateThreatIntelligence('iplist-known-malicious-ips')" \
  --action=deny-403 --preview \
  --description="Known malicious addresses"

Other available lists include Tor exit nodes, anonymous proxies, search scanners and the ranges of the major cloud providers (useful: hardly any legitimate shop customer browses from a data centre).

Indicative Standard cost, as an order of magnitude:

Item Approximate
Security policy ~$5/month each
Rule ~$1/month each
Requests evaluated ~$0.75 per million

With this lesson's policy (about fifteen rules) and around 10 million requests a month, that comes to something like $25-30/month. Compared with what an incident costs, or simply with what the extra instances the autoscaler spins up during a mass crawl cost, it is one of the platform's best-value investments. Always check the current rates in the official documentation, because they change and vary by region.

Enterprise's bill protection against DDoS deserves a thought: in a large volumetric attack, the cost is not in the damage, it is in the egress and compute bill generated by absorbing it. That insurance is the main reason a company with real exposure buys Enterprise.

  1. The complete autumn campaign policy

This is the final result for AlpinaShop, ordered by priority. Read it as a design document:

Priority Rule Action State
100 Office 203.0.113.24/29 allow Active
200 Ranges with confirmed abuse deny-403 Active
400 /admin outside Spain deny-404 Active
500 Threat intelligence (Enterprise) deny-403 Preview
900 cve-canary deny-403 Active
950 Exemption for /api/resenas allow Active
1000-1400 OWASP: SQLi, XSS, LFI, RCE, scanners deny-403 Preview → phased activation
2000 Automated crawling of the catalogue deny-403 Preview
2100 /exportar from the office only deny-403 Active
2200 Unused HTTP methods deny-403 Active
2300 Invalid Host deny-404 Active
3000 Brute force on /acceso rate-based-ban Active
3100 Search box limit throttle Active
3200 General limit per IP throttle Active
4000 reCAPTCHA on /pago redirect Active
2147483647 Default allow Active
# Apply the policy to the backend service
gcloud compute backend-services update bs-catalogo-web --global \
  --security-policy=pol-catalogo-web

# Edge policy for the images, compatible with the CDN
gcloud compute security-policies create pol-borde-imagenes \
  --type=CLOUD_ARMOR_EDGE \
  --description="Basic edge filtering for the image catalogue"

gcloud compute backend-buckets update bb-catalogo-imagenes \
  --edge-security-policy=pol-borde-imagenes

# Verification
gcloud compute backend-services describe bs-catalogo-web --global \
  --format="value(securityPolicy)"

And the check that it works, which must always be done from outside the office, because rule 100 exempts you from everything:

# Should return 403 (with rule 1000 already active)
curl -s -o /dev/null -w "%{http_code}\n" \
  "https://alpinashop.example/buscar?q=%27%20OR%201%3D1--"

# Should return 200
curl -s -o /dev/null -w "%{http_code}\n" \
  "https://alpinashop.example/buscar?q=mochila"

# Brute force: the first 5 get through, the rest 429
for i in $(seq 1 10); do
  curl -s -o /dev/null -w "%{http_code} " \
    -X POST -d "usuario=test&clave=test" https://alpinashop.example/acceso
done; echo

Common Mistakes and Tips

  • Enabling rules without preview. It is this lesson's serious mistake. A badly calibrated SQLi rule blocks the search box and nobody connects the drop in sales with the WAF.
  • Starting at sensitivity 3 or 4. Sensitivity 1, and you only go up if the log analysis justifies it.
  • Not numbering with gaps. Consecutive priorities force you to renumber to insert a rule.
  • Forgetting the office exemption rule at the lowest priority. If you block yourself during an incident, debugging becomes hell.
  • Testing from the office. Rule 100 always lets you through; it will look as though nothing works. Test from an external network.
  • Blocking countries without consulting anyone. It is a commercial and sometimes legal decision, not a technical one.
  • Confusing throttle with rate-based-ban. The first limits while the excess lasts; the second bans for a fixed period. For brute force, ban.
  • Rate thresholds chosen by eye. Measure the real 99th percentile in the load balancer log and multiply by two or three.
  • enforce-on-key=IP without thinking about shared IPs. A school or a company goes out through a single address; a low threshold blocks thirty legitimate people.
  • Returning deny-403 from rate limiters. The correct code is 429.
  • Enabling rule sets for technologies you do not use. php-v33-stable on a Python application: it costs money and adds false positives while contributing nothing.
  • Believing the WAF replaces validation in the code. Parameterised queries, output escaping and input validation are still mandatory. The WAF is the second line.
  • Not enabling VERBOSE logging during preview. Without the specific rule identifiers you cannot use opt_out_rule_ids and all you are left with is the axe of lowering the sensitivity.
  • Tip: manage the policy as code (06-07) and export its state before every change: gcloud compute security-policies describe pol-catalogo-web --format=yaml > politica-$(date +%F).yaml.
  • Tip: set an alert on the number of denied requests. A spike in blocks may be an attack, but it may also be a deployment of your own application that has started sending something the WAF does not recognise.

Exercises

Exercise 1 — Interpreting a week of preview

After seven days with all the rules in preview, the count by priority is:

   1204 1000   SQLi
    412 1100   XSS
    156 2000   Automated crawling
     18 1400   Scanners
      4 1200   LFI
      2 900    CVE

A sample of rule 1000's requests shows that 90 % are of the form /buscar?q=camiseta+t%C3%A9cnica+O%27Neill and /buscar?q=%22gore-tex%22. Those from 1100 are mostly POST /api/resenas with text containing <3 and >>.

Decide, rule by rule, what to enable, what to debug and what to leave in preview, and write the commands.

Exercise 2 — Designing protection for a new endpoint

AlpinaShop launches a public API for partner shops: POST /api/socios/pedidos, authenticated with an X-Alpina-Api-Key header. Requirements:

  • A maximum of 100 requests per minute per API key (not per IP: several shops share an operator).
  • Requests with no header are rejected at the edge, without reaching the application.
  • Only POST is accepted.
  • The partner shops are in Spain, Portugal and France.
  • A partner who exceeds the limit three times in a row must be banned for 15 minutes.

Write the rules with their priorities and justify the order.

Exercise 3 — Responding to an incident in progress

It is 22:00 on a Friday. The shop is extremely slow, the MIG is at 10 instances (its maximum) and Marta sees in the load balancer log that 80 % of the traffic is GET /buscar?q=<random> requests from around 4,000 different IP addresses spread across the world, with plausible browser user agents and no Referer header.

Describe the immediate response, the next 24 hours' response and the structural one. Why is a rule by IP no use here?


Solutions

Solution 1

Rule 1000 (SQLi) — debug, stay in preview. The 1204 blocks are nearly all false positives: apostrophes in surnames and quotes in brand searches. Enabling it would leave customers with no search box. Two combined corrections:

# a) Exclude the search path from the WAF; it already uses parameterised queries
gcloud compute security-policies rules create 990 \
  --security-policy=pol-catalogo-web \
  --expression="request.path.startsWith('/buscar') && request.method == 'GET'" \
  --action=allow \
  --description="Search box: parameterised queries verified in the code"

# b) And also disable the specific identifiers that come from the log
gcloud compute security-policies rules update 1000 \
  --security-policy=pol-catalogo-web \
  --expression="evaluatePreconfiguredWaf('sqli-v33-stable', {'sensitivity': 1, 'opt_out_rule_ids': ['owasp-crs-v030301-id942432-sqli']})"

The path exemption requires first verifying with Dani that the search box really does use parameterised queries. If it did not, the priority would be fixing the code, not relaxing the WAF.

Rule 1100 (XSS) — debug, stay in preview. The problem is localised in /api/resenas, an endpoint that receives free text by design. The exemption from section 7 (priority 950) is created and it is checked that reviews are escaped when rendering. With that exemption in place, another week of preview.

Rule 2000 (crawling) — preview, manual analysis. 156 in a week is little traffic: you have to look at where they come from. If they are a few sustained addresses, it is the competitor and it can be enabled. If they are scattered, there are probably legitimate clients among them (price aggregators contracted by AlpinaShop itself, for instance) and the expression needs refining first.

Rules 1400 (scanners), 1200 (LFI) and 900 (CVE) — enable now.

for P in 900 1200 1400; do
  gcloud compute security-policies rules update $P \
    --security-policy=pol-catalogo-web --no-preview
done

Low volumes and patterns with no legitimate explanation whatsoever: they are real attacks. They are enabled with no further discussion.

Solution 2

# 4900 - POST only on the API path
gcloud compute security-policies rules create 4900 \
  --security-policy=pol-catalogo-web \
  --expression="request.path.startsWith('/api/socios/') && request.method != 'POST'" \
  --action=deny-405 \
  --description="Partner API: POST only"

# 4910 - With no API key, you do not get past the edge
gcloud compute security-policies rules create 4910 \
  --security-policy=pol-catalogo-web \
  --expression="request.path.startsWith('/api/socios/') && !has(request.headers['x-alpina-api-key'])" \
  --action=deny-401 \
  --description="Partner API: key header missing"

# 4920 - Geographic restriction
gcloud compute security-policies rules create 4920 \
  --security-policy=pol-catalogo-web \
  --expression="request.path.startsWith('/api/socios/') && !(origin.region_code in ['ES','PT','FR'])" \
  --action=deny-403 --preview \
  --description="Partner API: ES, PT and FR only"

# 4930 - Limit per KEY, not per IP, with a ban
gcloud compute security-policies rules create 4930 \
  --security-policy=pol-catalogo-web \
  --expression="request.path.startsWith('/api/socios/')" \
  --action=rate-based-ban \
  --rate-limit-threshold-count=100 \
  --rate-limit-threshold-interval-sec=60 \
  --ban-threshold-count=300 \
  --ban-threshold-interval-sec=180 \
  --ban-duration-sec=900 \
  --conform-action=allow \
  --exceed-action=deny-429 \
  --enforce-on-key=HTTP_HEADER \
  --enforce-on-key-name=x-alpina-api-key \
  --description="Partner API: 100 rpm per key, 15-min ban"

Justification of the order. Rules are evaluated from lowest to highest priority and the first match wins, so they go from the cheapest and most categorical to the most costly:

  1. Wrong method (4900) and missing header (4910) are trivial checks that discard junk traffic without evaluating anything else. The sooner the better.
  2. Geography (4920) in preview, because a French partner using a proxy in Belgium would be left out and that has to be confirmed with the commercial team before enabling it.
  3. Rate limiting (4930) last: it is the most expensive rule to evaluate and it only makes sense to apply it to requests that have already passed the previous filters.

The key point is --enforce-on-key=HTTP_HEADER with --enforce-on-key-name. Counting by IP would fail in both directions: it would unfairly penalise several shops sharing an operator and it would not stop an abusive partner rotating addresses. With the corresponding warning: the header is controlled by the client, so the real validation of the key is still done in the application; Cloud Armor only uses it to group the count.

Solution 3

Why a rule by IP is no use. There are 4,000 distributed addresses, probably a botnet or a residential proxy service. Blocking addresses one by one is chasing a target that moves faster than you can type.

Immediate response (minutes). What they have in common is not the origin, it is the behaviour: /buscar paths with random terms and no Referer.

# 1) Stop the bleeding: an aggressive, temporary limit on the search box
gcloud compute security-policies rules create 3050 \
  --security-policy=pol-catalogo-web \
  --expression="request.path.startsWith('/buscar') && !has(request.headers['referer'])" \
  --action=throttle \
  --rate-limit-threshold-count=5 \
  --rate-limit-threshold-interval-sec=60 \
  --conform-action=allow --exceed-action=deny-429 \
  --enforce-on-key=IP \
  --description="INCIDENT 2026-XX: searches with no Referer"

# 2) If that is not enough, a reCAPTCHA challenge on the search box: it does not block people
gcloud compute security-policies rules create 3060 \
  --security-policy=pol-catalogo-web \
  --expression="request.path.startsWith('/buscar') && token.recaptcha_session.score < 0.7" \
  --action=redirect --redirect-type=GOOGLE_RECAPTCHA \
  --description="INCIDENT 2026-XX: challenge on the search box"

The redirect to reCAPTCHA is the best tool during a bot incident in a shop: people carry on buying, bots fall over.

The next 24 hours.

  • Enable adaptive protection if it was not already on (although it will take a while to have a baseline, and that is the lesson: it is enabled beforehand, not during).
  • Analyse the log grouping by ranges, user agent and autonomous system, looking for a more precise signature that allows the emergency rule to be refined and its impact on real customers reduced.
  • Review the collateral impact: how many legitimate searches were dropped with 429 while rule 3050 was in force.
  • Temporarily raise the MIG autoscaler's maximum if there is still pressure.

Structural.

  • Cache frequent searches in the CDN (03-03): a search served from the edge never touches the database.
  • Look into why the search box is so expensive. A rate limit that saves the shop is a patch over an endpoint that should cope with more; perhaps an index is missing or the search should move to a dedicated service.
  • Leave the emergency rules documented and with a review date, not permanent with a crisis threshold.
  • Evaluate Cloud Armor Enterprise: in an attack like this, threat intelligence would have recognised a good proportion of those origins and bill protection would have covered the extra cost.

Conclusion

AlpinaShop no longer takes everything that comes its way. You know what a WAF is and, above all, what it does not solve: it does not fix business logic flaws, it does not detect stolen credentials and it does not patch vulnerable dependencies. It is a mitigation layer that buys time, not an absolution, and that is why parameterised queries and input validation are still mandatory in Dani's code.

You understand where Cloud Armor acts: at the edge, at the same point of presence the client enters through, before the cache and before IAP, attached as a policy to the bs-catalogo-web backend service and as an edge policy to the bb-catalogo-imagenes backend bucket. You have mastered the model of policies, rules and priorities, with the default allow rule at the far end, the office exemption at priority 100 and the habit of numbering in hundreds so you can insert.

You have built the complete pol-catalogo-web: rules by IP and by geography — with the warning that blocking countries is a commercial and legal decision, not a technical one — the preconfigured OWASP sets at sensitivity 1 and their real problem, which is the false positives from the search box and the reviews, and the three ways of dealing with them in order of precision: opt_out_rule_ids, path exemption and lowering the sensitivity. You have written your own CEL rules for the competitor's crawler, the expensive export endpoint, useless HTTP methods and requests with an invalid Host. And you have learned the procedure that makes all of this safe: --preview first, VERBOSE logging, a week of real traffic, a count by priority, debugging, phased activation and never on the eve of a campaign, reading the difference between previewSecurityPolicy and enforcedSecurityPolicy in the log.

You have protected the sign-in form with rate-based-ban and the search box with throttle, choosing the grouping key with judgement and returning 429 instead of 403. You know about layer 7 adaptive protection and why it has to be enabled before you need it, the use of reCAPTCHA Enterprise with redirect instead of deny — when in doubt, in a shop you ask, you do not block — and the difference between Standard and Enterprise, with threat intelligence and bill protection as the weighty arguments.

There remains, however, an uncomfortable matter. The password of the app_catalogo user of alpinashop-pedidos is still travelling in an environment variable of the MIG template's startup-script: in clear text, visible to anyone who can read an instance's metadata, impossible to rotate without redeploying and present in every copy of the template. We have protected the edge meticulously while the key to the database is under the doormat. In the next lesson, 03-06, Secrets and Encryption: Secret Manager and Cloud KMS, we take that password out of there and move it to a store with versions, per-secret permissions and rotation; and along the way we understand how Google Cloud encrypts your data at rest, what customer-managed encryption keys (CMEK) are and when it is worth taking on the responsibility of managing them yourself.

Google Cloud Platform (GCP) Course

Module 1: Introduction to Google Cloud Platform

Module 2: Core GCP Services

Module 3: Networking and Security

Module 4: Data and Analytics

Module 5: Machine Learning and AI

Module 6: DevOps and Monitoring

Module 7: Advanced GCP Topics

Module 8: Final Project

© Copyright 2026. All rights reserved