AlpinaShop's load balancer is working: a global anycast IP, HTTPS, the MIG behind it and the images served straight from the bucket. But the inefficiency we closed the previous lesson with is still there. Every time a customer in Lisbon opens a backpack's product page, the same bytes leave europe-west1 again and cross the peninsula. With a 60 GB catalogue, hundreds of thousands of image requests a month and customers spread across Europe, that is latency the customer notices and egress the bill accumulates.

Cloud CDN solves precisely that problem, and it does so without you having to change a line of the application or move a single object. It is not a separate product with its own console, its own IP and its own domain: it is a checkbox enabled on the backend service or the backend bucket you already built in 03-02. That design detail is what makes this lesson short on commands and long on judgement: enabling it costs one --enable-cdn; doing it well consists of understanding what gets cached, with which key, for how long and how it is measured.

In this lesson you are going to enable caching on bb-catalogo-imagenes and bs-catalogo-web, choose the right cache mode for each, stop the marketing campaigns' utm_* parameters wrecking the hit ratio, decide a TTL policy consistent with the immutable names you already adopted in 02-02, learn why invalidation should not be your everyday tool, and measure the result in milliseconds and in euros.

Contents

  1. What a CDN is and exactly what problem it solves
  2. Google's network of points of presence
  3. Where Cloud CDN lives: on the load balancer, not beside it
  4. Enabling it on the backend bucket and on the backend service
  5. Cache modes: CACHE_ALL_STATIC, USE_ORIGIN_HEADERS, FORCE_CACHE_ALL
  6. The cache key and the utm_* parameter disaster
  7. TTL: who decides how long an object lives at the edge
  8. Immutable names: the strategy that holds all of the above together
  9. Cache invalidation: the emergency button
  10. Private content in the cache: signed URLs and signed cookies
  11. Negative caching, stale content and other useful options
  12. Measurement: hit ratio, the Age and Via headers, and debugging with curl
  13. The real saving: latency and euros
  14. When Cloud CDN is not enough: Media CDN

  1. What a CDN is and exactly what problem it solves

A content delivery network (CDN) is a set of cache servers spread out geographically that keep copies of your content close to the users. When a client asks for an object, the nearest node serves it instead of the origin server.

It sounds simple, but it is worth being precise about the three distinct problems it solves, because they are not the same and they are measured differently:

Problem Without a CDN With a CDN
Latency The object travels from europe-west1 to Lisbon on every request It is served from the Lisbon point of presence: tens of milliseconds less
Egress cost Every byte served is billed as internet egress from the bucket Cache hits are billed at the CDN rate, which is cheaper
Load on the origin The MIG and the bucket serve every request The origin only sees the cache misses: fewer instances, fewer operations

There is a fourth effect, less obvious and very valuable: spike absorption. When AlpinaShop launches the autumn campaign and the home page receives ten times the usual traffic, if the home page and its images are cached, the MIG barely notices. The autoscaling you configured in 02-01 will still be there as a safety net, but it will trigger far less often.

And one limitation to be clear about from the start: a CDN only speeds up what can be cached. A specific customer's shopping cart page, the result of a payment or Lucía's dashboard are never cached. A CDN does not make your application faster; it makes most of the traffic never reach your application.

  1. Google's network of points of presence

Google maintains more than two hundred cache locations around the world, on top of dozens of regions and network points of presence. Many of those caches are inside the operators' own networks (what Google calls Google Global Cache), that is, physically at the operator the client is connected to.

What matters for your design is not the number, but the mechanics:

flowchart LR
    C1([Lisbon customer])
    C2([Helsinki customer])
    P1["Lisbon PoP<br/>cache"]
    P2["Helsinki PoP<br/>cache"]
    LB["Global load balancer<br/>alpinashop-lb-ip"]
    O1["Bucket alpinashop-catalogo<br/>europe-west1"]
    O2["MIG alpinashop-web-mig<br/>europe-west1"]

    C1 -->|1. GET /imagenes/...| P1
    C2 -->|1. GET /imagenes/...| P2
    P1 -->|2. cache miss| LB
    P2 -->|2. hit: goes no further| P2
    LB --> O1
    LB --> O2
    P1 -.->|3. cached response| C1
  1. The client resolves alpinashop.example and reaches the anycast IP. That IP is announced from every point of presence, so the client enters through the nearest one.
  2. At that point of presence the cache is consulted. If there is a hit, the response leaves from there and never reaches europe-west1.
  3. If there is a miss, the request travels over Google's private network to the origin — not over the public internet — the response is stored at the edge and delivered to the client.

Two practical consequences:

  • The cache is not one, there are many. An object can be warm in Madrid and cold in Warsaw. That is why the hit ratio is never 100 %, and nor should it be.
  • Google collapses concurrent requests for the same object (request coalescing, on by default): if a thousand clients ask at once for an image that is not cached, the origin receives one request, not a thousand. This avoids the stampede effect at launches.

  1. Where Cloud CDN lives: on the load balancer, not beside it

This is the idea that causes the most confusion for people arriving from other platforms. With many providers the CDN is a separate service: you create a "distribution", they give you a host name of your own and you point your DNS at it. Not in Google Cloud.

Cloud CDN is a property of a backend of the global external Application Load Balancer. It is enabled like this:

  • On a backend bucket (bb-catalogo-imagenes) → it caches Cloud Storage objects.
  • On a backend service (bs-catalogo-web) → it caches whatever the MIG instances return, or tomorrow the Cloud Run serverless NEG (DA-001).

Several very convenient things follow from that:

  • You do not change the DNS or the IP. You carry on with alpinashop-lb-ip and the same certificate.
  • The same URL map decides what is cached and what is not, because each path goes to a different backend and each backend has its own cache configuration.
  • Cloud Armor (03-05) is evaluated at the edge, so the security rules also apply to requests that end up being served from the cache. The cache is not a hole through which unfiltered traffic sneaks in.
  • The load balancer logs you enabled in 03-02 already carry the cache fields. There is no separate observability system.

Remember how the alpinashop-url-map map ended up:

Path Backend Cache it?
/imagenes/* bb-catalogo-imagenes (bucket) Yes, aggressively: they are immutable files
/estatico/* bb-catalogo-imagenes (bucket) Yes: versioned CSS and JS
/ and product pages bs-catalogo-web (MIG) Yes, carefully: HTML that changes
/carrito, /cuenta, /pago bs-catalogo-web (MIG) Never: per-user content

That table is the lesson's design decision. Everything else is parameters.

  1. Enabling it on the backend bucket and on the backend service

We start with the easy, highest-impact part: the images.

gcloud config set project alpinashop-prod

# CDN on the backend bucket that serves /imagenes/* and /estatico/*
gcloud compute backend-buckets update bb-catalogo-imagenes \
  --enable-cdn \
  --cache-mode=CACHE_ALL_STATIC \
  --default-ttl=3600 \
  --max-ttl=31536000 \
  --client-ttl=3600 \
  --negative-caching

What each option does:

  • --enable-cdn: turns the cache on. It is the only mandatory one; the rest are adjustments.
  • --cache-mode: the general policy (section 5).
  • --default-ttl: how long a response is kept if the origin says nothing.
  • --max-ttl: the ceiling. Even if the origin asks for a year, this value will not be exceeded. Here we leave it at a year because we trust our own Cache-Control.
  • --client-ttl: the maximum max-age communicated to the browser. It lets you cache a lot at the edge and little at the client, which is very useful (section 7).
  • --negative-caching: caches errors too (section 11).

Now the MIG's backend service. Here we are more conservative, because behind it there is an application generating HTML:

gcloud compute backend-services update bs-catalogo-web \
  --global \
  --enable-cdn \
  --cache-mode=CACHE_ALL_STATIC \
  --default-ttl=300 \
  --max-ttl=3600 \
  --client-ttl=60 \
  --serve-while-stale=86400 \
  --compression-mode=AUTOMATIC
  • --default-ttl=300: five minutes for the static content the application serves that comes with no headers of its own.
  • --serve-while-stale=86400: if the origin does not respond, the edge can carry on serving the expired copy for up to 24 hours while it revalidates. This option turns a MIG outage into a degradation instead of a 502 error for cached content. It is one of the best benefit/effort ratios in the whole module.
  • --compression-mode=AUTOMATIC: the edge compresses with gzip or Brotli the text responses the origin sends uncompressed, saving bandwidth without touching gunicorn.

Check the state with:

gcloud compute backend-buckets describe bb-catalogo-imagenes \
  --format="yaml(cdnPolicy, enableCdn)"

gcloud compute backend-services describe bs-catalogo-web --global \
  --format="yaml(enableCDN, cdnPolicy)"

CDN configuration changes take a few minutes to propagate to every point of presence. And take note: changing the configuration does not flush the cache. The objects already stored stay there with their original TTL.

  1. Cache modes: CACHE_ALL_STATIC, USE_ORIGIN_HEADERS, FORCE_CACHE_ALL

The cache mode answers one question: who decides what gets cached, the origin or the CDN?

Mode What it caches Who is in charge Risk When to use it
USE_ORIGIN_HEADERS Only what the origin explicitly marks as cacheable The origin, always None; if you get it wrong, it simply is not cached Applications that already emit correct Cache-Control and want total control
CACHE_ALL_STATIC Static content by content type and extension, even if the origin says nothing; respects the origin for the rest Shared Low: private and no-store are respected The sensible default. The one AlpinaShop uses
FORCE_CACHE_ALL Every successful response, ignoring private and no-store The CDN, ignoring the origin High: it can serve one user's page to another Only on backends serving exclusively immutable public content

Details that matter:

  • CACHE_ALL_STATIC considers content "static" by its MIME type and its extension: images, video, CSS, JavaScript, fonts, PDF. HTML does not fall into that category, so AlpinaShop's product pages will not be cached unless the application asks for it with its own Cache-Control. That is exactly what we want: explicit control over the HTML.
  • In any mode, a response with Cache-Control: private, no-store or Set-Cookie is not cached (except under FORCE_CACHE_ALL). That is the safety net that prevents the classic disaster.
  • FORCE_CACHE_ALL on a backend that serves session pages is probably the worst security incident you can cause with a checkbox. If you ever need it, apply it to a dedicated backend service that only public paths reach, never to the one serving the whole site.

For AlpinaShop the decision is:

  • bb-catalogo-imagenes → CACHE_ALL_STATIC. Everything there is static by definition and already comes with a one-year Cache-Control from 02-02.
  • bs-catalogo-web → CACHE_ALL_STATIC, and let the Flask application decide path by path.

This is how Dani declares it in the application:

# app.py — explicit cacheability control per path

@app.route("/producto/<sku>")
def product_page(sku):
    product = repository.get(sku)
    response = make_response(render_template("producto.html", producto=product))
    # Public and cacheable for 5 minutes at the edge, 1 minute in the browser.
    # 's-maxage' is read by the CDN; 'max-age' is read by the browser.
    response.headers["Cache-Control"] = "public, max-age=60, s-maxage=300"
    return response


@app.route("/carrito")
def cart():
    response = make_response(render_template("carrito.html"))
    # Never, nowhere, under no circumstances.
    response.headers["Cache-Control"] = "private, no-store"
    return response

A rule worth committing to memory: if a response depends on who is asking, it carries private, no-store. Put it there explicitly and do not rely on the cache mode saving you.

  1. The cache key and the utm_* parameter disaster

The cache key is the string with which the CDN identifies a stored object. If two requests produce the same key, the second is a hit. If they produce different keys, the second is a miss even though the content is identical.

By default the key includes:

  • The protocol (http or https).
  • The host (alpinashop.example).
  • The path (/imagenes/productos/MOCH-2210/web/frontal-800.a3f9c1.webp).
  • The complete query string (?utm_source=newsletter&utm_campaign=otono26).

And there is the problem. AlpinaShop's marketing team launches a campaign and the image URLs start arriving like this:

/imagenes/productos/MOCH-2210/web/frontal-800.a3f9c1.webp?utm_source=newsletter&utm_medium=email&utm_campaign=otono26
/imagenes/productos/MOCH-2210/web/frontal-800.a3f9c1.webp?utm_source=instagram&utm_medium=social&utm_campaign=otono26
/imagenes/productos/MOCH-2210/web/frontal-800.a3f9c1.webp?fbclid=IwAR3x9...

They are the same object with three different keys. Each variant causes a cache miss, a read from the bucket and a cache fill charge. With a big campaign, the hit ratio collapses from 92 % to 40 % and nobody understands why the bill went up exactly when the visitors arrived.

The solution is to exclude those parameters from the key:

# Backend bucket: the images depend on NO parameter at all
gcloud compute backend-buckets update bb-catalogo-imagenes \
  --cache-key-query-string-blacklist=utm_source,utm_medium,utm_campaign,utm_term,utm_content,gclid,fbclid,msclkid,ref

# Backend service: the opposite approach, an allowlist.
# Only these parameters form part of the key; the rest are ignored.
gcloud compute backend-services update bs-catalogo-web --global \
  --cache-key-query-string-whitelist=pagina,orden,talla,color

Two strategies, and it is worth understanding why each one is used where it is:

Strategy Command Advantage Risk
Blocklist (exclude) --cache-key-query-string-blacklist Easy to get started with; breaks nothing existing Every new tracking parameter has to be added by hand
Allowlist (include only) --cache-key-query-string-whitelist Immune to new parameters; maximum ratio If you forget a parameter that does change the content, you serve the wrong page

For the bucket, the blocklist is more than enough; in reality no parameter affects an image. For the backend service, the allowlist is more powerful but requires Dani to review the list every time he adds a new filter to the catalogue. It is a decision to be documented, not improvised.

If the images really do not depend on any parameter, the cleanest option is simply:

gcloud compute backend-buckets update bb-catalogo-imagenes \
  --cache-key-query-string-blacklist=""   # equivalent to ignoring the whole query

Other components that can be added to the key:

  • --cache-key-include-http-header=X-Alpina-Region: useful if you serve variants by header. Each distinct value multiplies the cache entries, so use it only with headers of very low cardinality.
  • --cache-key-include-named-cookie=idioma: variants by a specific cookie. Same warning.
  • --no-cache-key-include-protocol: unifies HTTP and HTTPS into the same entry. Since at AlpinaShop all HTTP is redirected to HTTPS (03-02), it makes no difference.

And a warning that points the other way: if the origin returns a Vary header with values the CDN does not support, the response is simply not cached. Vary: Accept-Encoding is safe; Vary: User-Agent kills the cache completely, because there are tens of thousands of different user agents.

  1. TTL: who decides how long an object lives at the edge

There are five values involved and they are constantly confused. This is the chain of command:

Value Where it is defined Who it affects What it does
Cache-Control: max-age Origin header Browser and CDN General TTL
Cache-Control: s-maxage Origin header Shared caches only (the CDN) Takes priority over max-age at the edge
--default-ttl Backend configuration CDN TTL when the origin says nothing
--max-ttl Backend configuration CDN Ceiling: it trims whatever the origin asks for
--client-ttl Backend configuration Browser Ceiling on the max-age sent to the client

The actual sequence, for a response arriving from the origin with Cache-Control: public, max-age=31536000, immutable and a backend configured with --max-ttl=86400 --client-ttl=3600:

  1. The CDN reads max-age = 31,536,000 s.
  2. It applies --max-ttl: it stores the object for 24 hours, not a year.
  3. It applies --client-ttl: it rewrites the header towards the client as max-age=3600, one hour.

The combination of a high s-maxage + a low max-age is the most useful and the least used:

Cache-Control: public, max-age=60, s-maxage=86400

It means: "CDN, keep it a day; browser, keep it a minute". The advantage: if you need to change the content, you invalidate the CDN cache (which you do control) and within 60 seconds every browser in the world sees the new thing. You can never invalidate somebody else's browser cache; that is why it is best for the client to cache little and the edge to cache a lot — except when the file name is immutable, which is the next case.

  1. Immutable names: the strategy that holds all of the above together

In 02-02 you made a decision that now pays its interest: AlpinaShop's images are stored with a name that is never reused, with a hash or version embedded in it:

productos/MOCH-2210/web/frontal-800.a3f9c1.webp
productos/MOCH-2210/web/frontal-800.b7e204.webp   ← the new photo is a different object

And they are uploaded with Cache-Control: public, max-age=31536000, immutable.

The consequence is enormous and deserves stating explicitly:

If the name is immutable, the invalidation problem disappears. Publishing a new photo does not consist of updating a cached object, but of publishing an object nobody has ever asked for. The HTML page points at the new name and the cache serves the new object from the very first moment, while the old one expires on its own.

With that scheme, the correct policy for the bucket is to cache for a year for real, in the client too:

gcloud compute backend-buckets update bb-catalogo-imagenes \
  --cache-mode=CACHE_ALL_STATIC \
  --default-ttl=3600 \
  --max-ttl=31536000 \
  --client-ttl=31536000

Here we do want a long client-ttl: a returning customer's browser will never ask for the image again. Note the contrast with the previous section: the client TTL is long when the name is immutable and short when it is not. That is the whole rule.

And the corollary for the HTML: a product page cannot have a long TTL, because its URL is stable (/producto/MOCH-2210). That is why it carries s-maxage=300. Content with a stable URL → short TTL. Content with a versioned URL → one-year TTL.

  1. Cache invalidation: the emergency button

Sometimes you have to remove something from the edge before it expires: a wrong price was published, an image with the logo wrong, an incorrect piece of legal text.

# Invalidate a specific object
gcloud compute url-maps invalidate-cdn-cache alpinashop-url-map \
  --path="/imagenes/productos/MOCH-2210/web/frontal-800.a3f9c1.webp"

# Invalidate a whole prefix (the wildcard is only valid at the end)
gcloud compute url-maps invalidate-cdn-cache alpinashop-url-map \
  --path="/imagenes/productos/MOCH-2210/*"

# Invalidate the whole site: the sledgehammer. Avoid it.
gcloud compute url-maps invalidate-cdn-cache alpinashop-url-map \
  --path="/*" --async

# Follow the operation
gcloud compute operations list --global --filter="operationType~invalidate" --limit=5

Notice that invalidation is run against the URL map, not against the backend: it propagates to every backend that map routes to.

Why it should not be your everyday mechanism:

  • It is slow in incident terms. The command has to be propagated to every point of presence in the world; the order of magnitude is minutes, not seconds. During a content crisis, those minutes feel long.
  • It is subject to quota. There is a limit on invalidations per project and unit of time. A deployment that invalidates /* on every release hits the quota sooner than you think.
  • A /* destroys the hit ratio. After a total invalidation, all the traffic in the world hits the origin at once. You have turned a routine release into an involuntary load test.
  • It requires elevated permissions, typically roles/compute.loadBalancerAdmin, which you will not want to hand out lightly (03-04).

The correct alternative is the one in section 8: versioned names for the assets, a short TTL for the HTML. Invalidation is reserved for mistakes, which is what it is for.

  1. Private content in the cache: signed URLs and signed cookies

The alpinashop-catalogo bucket is private (you created it with --public-access-prevention in 02-02) and the backend bucket reads it with its own permissions, so the catalogue images are public through the load balancer but not directly. Perfect for the catalogue.

Now picture the following AlpinaShop case: the videos for the technical climbing guides should only be watched by customers who have bought the in-person course. They are large files, they benefit greatly from caching, but they cannot be public.

That is what Cloud CDN signed URLs and signed cookies are for. The difference from Cloud Storage signed URLs (02-02) is important: those are validated by Cloud Storage and break the cache (each signature is a different URL); these are validated by the CDN edge and allow caching.

# 1) Generate a 16-byte signing key in base64url
head -c 16 /dev/urandom | base64 | tr +/ -_ > clave-cdn.txt

# 2) Register it on the backend
gcloud compute backend-buckets add-signed-url-key bb-catalogo-guias \
  --key-name=clave-guias-2026-01 \
  --key-file=clave-cdn.txt

# 3) State how long the signed response is cached
gcloud compute backend-buckets update bb-catalogo-guias \
  --signed-url-cache-max-age=3600

# 4) Store the key in Secret Manager (03-06) and delete the local file
gcloud secrets create cdn-clave-guias --data-file=clave-cdn.txt
shred -u clave-cdn.txt

Signed cookie versus signed URL:

Signed URL Signed cookie
Scope One specific object A whole path prefix
How it arrives Inside the link A Set-Cookie header with the name Cloud-CDN-Cookie
Good for A one-off download A gallery or a player with many files
Effect on the cache Caches well: the key ignores the signature parameters Caches well, and a single cookie covers the whole session

For the guide videos, the cookie is the right option: the Flask application validates the purchase, issues a signed cookie valid for /guias/ for two hours and the player requests the segments with no further authentication. The edge validates the signature and serves from the cache.

Warning. Signing keys are cryptographic material: they give access to paid content. They must live in Secret Manager (03-06), be rotated periodically (several keys can be registered at once on the same backend precisely so you can rotate without interruption) and never appear in the repository. If the protected content has contractual or rights implications, have a security professional review the design before production.

  1. Negative caching, stale content and other useful options

Negative caching. Without it, if a crawler asks for ten thousand non-existent URLs, your origin generates ten thousand 404s. With it, the edge remembers the error:

gcloud compute backend-services update bs-catalogo-web --global \
  --negative-caching \
  --negative-caching-policy='404=120,410=300,500=0,502=0,503=0'

The interesting nuance is the 5xx codes set to 0: we do not want to cache our own server errors. If the MIG returns a 502 during a deployment and we cache it for two minutes, we have multiplied the deployment's impact. 404s, yes: a product that does not exist will still not exist in two minutes.

Serving stale content (--serve-while-stale). We already enabled it: during that window, if the origin does not respond or returns an error, the edge serves the expired copy and revalidates in the background. It is the difference between "the site is a little out of date" and "the site is down".

Bypassing the cache on demand. Very useful so that Dani can debug in production without touching the configuration:

gcloud compute backend-services update bs-catalogo-web --global \
  --bypass-cache-on-request-headers=x-alpina-debug

From then on, any request with that header goes straight to the origin. Since it can bypass the cache for the whole site, the header should be hard to guess and should not be documented outside the team.

  1. Measurement: hit ratio, the Age and Via headers, and debugging with curl

Without measurement, all of the above is faith. There are three levels of observation.

Level 1: the HTTP response. The fastest way of knowing whether something is being cached:

curl -sI https://alpinashop.example/imagenes/productos/MOCH-2210/web/frontal-800.a3f9c1.webp
HTTP/2 200
content-type: image/webp
cache-control: public, max-age=31536000, immutable
age: 4127
via: 1.1 google

How to read it:

Header What it means
Age: 4127 The copy has been at the edge for 4127 seconds. Age present and greater than zero = cache hit.
No Age Cache miss: the response comes from the origin
Via: 1.1 google It has gone through Google's proxy infrastructure. It appears whether there is a hit or not
Cache-Control What the edge tells the browser, already trimmed by --client-ttl

The classic test: request it twice in a row. The first has no Age; the second does. If the second one does not either, there is a cacheability problem and it is time for level 2.

Level 2: the load balancer logs. The cache fields are already there:

gcloud logging read '
  resource.type="http_load_balancer"
  AND httpRequest.requestUrl:"/imagenes/"
' --project=alpinashop-prod --limit=20 --freshness=1h \
  --format="table(
    httpRequest.requestUrl,
    httpRequest.cacheLookup,
    httpRequest.cacheHit,
    httpRequest.cacheFillBytes,
    jsonPayload.statusDetails
  )"
  • cacheLookup: false → the cache was not even consulted. Almost always it means the CDN is not enabled on that backend, or the request is not cacheable (a method other than GET, an Authorization header…).
  • cacheLookup: true, cacheHit: false → it was looked up and was not there: a legitimate miss or a fragmented cache key (section 6).
  • cacheFillBytes → the bytes fetched from the origin to fill the cache. If this is persistently high, your ratio is bad.
  • statusDetails: response_from_cache → confirmation of a hit.

Level 3: metrics. The hit ratio is obtained from loadbalancing.googleapis.com/https/request_count grouped by the cache_result label, whose values are HIT, MISS, DISABLED and revalidation variants. Marta's dashboard (06-04) should carry:

Metric / view Target for AlpinaShop
Hit ratio on /imagenes/* > 90 %
Overall hit ratio > 60 %
Monthly cacheFillBytes Stable or falling despite traffic growth
p95 of image total_latencies Well below the HTML p95

A guide to debugging a persistent miss. If an object is never cached, check in this order:

  1. --enable-cdn really is on the right backend (describe, not memory).
  2. The response does not carry Cache-Control: private, no-store or no-cache.
  3. The response does not carry Set-Cookie. It is the most frequent cause in Flask applications: merely touching session in the view is enough for a cookie to be emitted and the response to stop being cacheable.
  4. There is no Vary with high-cardinality headers (User-Agent, Cookie).
  5. The request is a GET and carries no Authorization.
  6. The cache key is not fragmented by query parameters.
  7. The response comes with Content-Length or a determinable length.

Point 3 is fixed in two lines and usually explains half the cases:

@app.after_request
def strip_cookie_on_static(response):
    """Stops an accidental session cookie making the catalogue uncacheable."""
    if request.path.startswith(("/estatico/", "/imagenes/")):
        response.headers.pop("Set-Cookie", None)
    return response

  1. The real saving: latency and euros

Let us run an indicative calculation for AlpinaShop. Suppose 2 TB a month of images served to European customers and a hit ratio of 90 %.

Item Without a CDN With a CDN at 90 %
Internet egress from the bucket 2,000 GB × ~$0.11/GB ≈ $220 200 GB (the misses) → cache fill ≈ $2–20
Egress from the CDN cache — 1,800 GB × ~$0.08/GB ≈ $144
Cache lookups — ~2 M requests × $0.0075/10,000 ≈ $1.5
Class B operations on the bucket Millions of reads Only the fill ones: two orders of magnitude fewer
Approximate total ≈ $220 ≈ $150–170

Order-of-magnitude prices for illustration. Egress, cache fill and lookup rates vary by continent and by volume tier, and change over time: always check them in the calculator and in the official pricing documentation before presenting a figure to anyone.

Honest conclusions from that table, which is more interesting than a brochure's "you save 90 %":

  • The saving on the bill is real but moderate: around 25-30 %. Egress from the cache is paid for too.
  • The saving grows with traffic, because the origin stops scaling with the visitors: fewer MIG instances, fewer bucket operations, fewer IOPS.
  • The big win is latency. A customer in Lisbon goes from a time to first byte of around 40-60 ms to under 15 ms; one in Helsinki, from over 60 ms to similar figures. On a product page with twelve images, that translates into hundreds of milliseconds of full page load, which is exactly what Google measures in its page experience signals and what shows up in conversion.
  • And there is a saving that appears on no bill: campaign peaks stop being a capacity problem.

  1. When Cloud CDN is not enough: Media CDN

Google Cloud has a second delivery product, Media CDN, built on the same infrastructure that serves YouTube. It is not "Cloud CDN improved": it is a different product, with a different configuration model (EdgeCacheService, EdgeCacheOrigin, EdgeCacheKeyset) and a different audience.

Cloud CDN Media CDN
Configured on The existing global load balancer Its own resources, independent of the load balancer
Designed for Websites, APIs, images, static assets Video on demand and live, mass downloads
Typical volume Any Tens of TB a month upwards
Extras Those covered in this lesson Advanced routing rules, better economics at large scale

For AlpinaShop, Cloud CDN is the right choice and will continue to be. If one day the technical guides section turns into a video platform with thousands of hours and tens of TB a month, then it would be time to evaluate Media CDN.

Common Mistakes and Tips

  • Believing Cloud CDN is a separate product and looking for where to create the "distribution". It is a checkbox on the backend of the load balancer you already have.
  • Enabling FORCE_CACHE_ALL on the backend that serves the whole site. It is the fastest route to serving one customer another customer's shopping cart. If you need it, do it on a backend dedicated to public paths.
  • Ignoring campaign parameters. utm_* and fbclid fragment the cache key and sink the ratio on exactly the day the traffic arrives. Configure the blocklist before the campaign, not after.
  • Confusing max-age with s-maxage. The second is only read by shared caches and takes priority at the edge. Low max-age + high s-maxage is the pair you want for HTML.
  • Setting a long --client-ttl on content with a stable URL. You cannot invalidate anybody's browser. A long client TTL only with immutable names.
  • Using invalidation as part of the deployment. It is slow, it has a quota and a /* leaves the origin exposed. Version the names.
  • Forgetting about Set-Cookie. A Flask session touched by accident in a catalogue view makes the whole path uncacheable and nobody understands the low ratio.
  • Vary: User-Agent. It nullifies the cache completely. Review which Vary your application and your CMS emit.
  • Caching 5xx. With badly configured negative caching, a deployment with a thirty-second glitch turns into five minutes of errors for everyone. Set the 5xx to 0.
  • Not enabling --serve-while-stale. It is practically free and turns outages into degradations.
  • Measuring "by eye" from the browser. The browser has its own cache and will deceive you. Debug with curl -sI and with the cacheHit/cacheLookup fields in the log.
  • Tip: test from several locations. A hit in Madrid says nothing about Warsaw. Each point of presence has its own cache and its own warm-up.
  • Tip: separate what is cacheable from what is not in the URL map. It is easier to reason about /estatico/* with a one-year cache and /cuenta/* with no cache than about a single backend with subtle rules.

Exercises

Exercise 1 — Diagnosing a 38 % hit ratio

Marta opens the dashboard a week after enabling the CDN and sees a 38 % ratio on /imagenes/*, when she expected more than 90 %. She gathers this data:

  • curl -sI on a specific image, twice in a row: the second response does carry Age: 61.
  • The log shows many requests with cacheLookup: true, cacheHit: false and URLs ending in ?utm_source=....
  • Other requests show cacheLookup: false on /imagenes/promo/* paths.
  • The Cache-Control of the images under /imagenes/promo/ is public, max-age=300.

Explain each symptom separately and write the commands that fix the situation.

Exercise 2 — Cache policy for three new paths

AlpinaShop adds three paths. Define for each one: cache mode, the Cache-Control the application should emit, client and edge TTL, and which components the cache key should carry.

  1. /api/stock/<sku> — returns JSON with the units available. It changes every few minutes and thousands of visitors query it.
  2. /estatico/app.4f2b9c.js — the JavaScript bundle, with a hash in the name.
  3. /cuenta/pedidos — the authenticated customer's order history.

Exercise 3 — Calculating a campaign's impact

The autumn campaign is going to multiply image traffic by five for two weeks: from 2 TB/month to a rate equivalent to 10 TB/month. Estimate the monthly egress cost with and without a CDN assuming a 90 % ratio, and answer: what other effect, more important than cost, does the CDN have in that campaign? What would you configure before it starts?


Solutions

Solution 1

They are three different problems disguised as one:

a) The specific image is being cached. The Age: 61 on the second request proves it. So the backend bucket's base configuration is correct and does not need touching.

b) The utm_* are fragmenting the key. Each campaign variant generates a new entry: cacheLookup: true, cacheHit: false is exactly this problem's signature. It is fixed by excluding the parameters from the key:

gcloud compute backend-buckets update bb-catalogo-imagenes \
  --cache-key-query-string-blacklist=utm_source,utm_medium,utm_campaign,utm_term,utm_content,gclid,fbclid,msclkid

c) /imagenes/promo/* is not even looked up in the cache. cacheLookup: false means the request never got as far as a lookup. If those paths are routed in the URL map towards a different backend (for example a bs-promociones created for the campaign), that backend does not have --enable-cdn. It is checked and fixed like this:

gcloud compute url-maps describe alpinashop-url-map \
  --format="yaml(pathMatchers)"

gcloud compute backend-services update bs-promociones --global \
  --enable-cdn --cache-mode=CACHE_ALL_STATIC \
  --default-ttl=3600 --max-ttl=86400

d) The max-age=300 on the promo images is not an error, but it is inconsistent: if the promotional images also have immutable names, they should carry a year. If they do not, the underlying fix is to adopt the same naming convention as the rest of the catalogue, not to raise the TTL blindly.

Solution 2

Path Mode App's Cache-Control Edge Client Cache key
/api/stock/<sku> CACHE_ALL_STATIC (with an explicit header) public, max-age=0, s-maxage=60 60 s 0 s Protocol + host + path. No query parameters
/estatico/app.4f2b9c.js CACHE_ALL_STATIC public, max-age=31536000, immutable 1 year 1 year Path only; ignore the whole query
/cuenta/pedidos Irrelevant private, no-store Never Never Not applicable

The reasoning:

  1. Stock. It is JSON, CACHE_ALL_STATIC would not cache it on its own, so the explicit header is mandatory. s-maxage=60 with max-age=0 is the key combination: the edge absorbs thousands of requests a minute and the browser always asks, so a stock change is seen at most a minute late. Sixty seconds of lag on a units counter is acceptable; a minute of stale data at the payment step would not be, and that is why the definitive stock check is done at the moment the order is confirmed, not by reading this API.
  2. JS bundle. A hashed name → textbook case: a year everywhere. Publishing a new version means publishing a different name.
  3. Order history. An explicit private, no-store. And very importantly: never on a backend with FORCE_CACHE_ALL. It will also carry a session Set-Cookie, which reinforces the uncacheability, but you must not depend on that side effect: the header is set deliberately.

Solution 3

Indicative cost at 10 TB/month:

  • Without a CDN: 10,000 GB × ~$0.11/GB ≈ $1,100/month.
  • With a CDN at 90 %: 9,000 GB from cache × ~$0.08/GB ≈ $720, plus 1,000 GB of fill (≈$10-40) and the lookups (a few dollars) ≈ $740-770/month. A saving of roughly 30 %.

The most important effect is not the cost, it is the load on the origin. Without a CDN, that fivefold traffic is absorbed by the MIG and the bucket: autoscaling would climb towards the limit of 10 instances, the compute cost would shoot up and latency would get worse on exactly the day of the highest sales. With a 90 % hit rate, the origin only sees 10 % of the image traffic; the MIG stays close to its usual size and the customer experience is better precisely when it matters most.

What to configure before the campaign starts:

  1. The blocklist of parameters utm_*, gclid and fbclid. Without this, the campaign sabotages itself.
  2. --serve-while-stale=86400 on bs-catalogo-web, so that an origin problem under load degrades rather than takes the shop down.
  3. Negative caching with the 5xx at 0 and 404s at a couple of minutes.
  4. Verify with curl -sI that the campaign images return Age on the second request, before the launch.
  5. A dashboard with the hit ratio and cacheFillBytes visible during the campaign, and an alert if the ratio drops below 80 %.
  6. Do not invalidate the cache during the campaign except in a real emergency: it would mean handing the traffic peak to the origin.

Conclusion

AlpinaShop no longer sends the same bytes to Lisbon over and over again. You have understood that Cloud CDN is not a separate product but a property of the load balancer you built in 03-02: it is enabled with --enable-cdn on bb-catalogo-imagenes and on bs-catalogo-web, without touching the DNS, the IP, the certificate or the application, and taking advantage of the same logs and the same metrics.

You know how to choose the cache mode with judgement — CACHE_ALL_STATIC as the sensible default, USE_ORIGIN_HEADERS when the application really is in charge, and FORCE_CACHE_ALL only on backends dedicated to public content — and why a response with private, no-store or Set-Cookie must never be cached. You have mastered the cache key, which is where the ratio is won or lost: you have excluded utm_*, gclid and fbclid so that the autumn campaign does not destroy the work, and you know the trade-off between blocklist and allowlist. You have put the five TTL values in order — max-age, s-maxage, --default-ttl, --max-ttl, --client-ttl — and you have seen why the immutable names adopted in 02-02 are the piece that makes invalidation unnecessary, that slow, quota-limited operation capable of leaving the origin out in the open. You know how to serve paid content from the cache with signed cookies, how to cache 404s but never 500s, how to keep serving during an outage with --serve-while-stale, and how to debug a cache miss by reading Age, Via and the log's cacheHit/cacheLookup fields. And you have an honest calculation of the saving: around 30 % of the egress bill, far more in origin load, and a latency improvement that shows on every product page.

With this, AlpinaShop's edge is complete as far as performance goes: network, load balancing and caching. But there is a problem we have been putting off lesson after lesson. In 03-01 we said the firewall does not distinguish between people and that "only Lucía and Marta can get in over SSH" is solved somewhere else. In 02-02 we limited Lucía to a bucket prefix with a condition we did not explain. In this very lesson we have said that invalidating the cache requires an elevated role that is best not handed around. All of that points to the same place: who can do what, on which resource. In the next lesson, 03-04, Identity and Access Management (IAM), we build AlpinaShop's real permission model: groups instead of people, predefined roles instead of Editor, a bespoke custom role for Lucía, service accounts with no downloaded keys, impersonation, conditions, and the tools for answering any cloud administrator's most frequent question: "why can't this user do this?".

Google Cloud Platform (GCP) Course

Module 1: Introduction to Google Cloud Platform

Module 2: Core GCP Services

Module 3: Networking and Security

Module 4: Data and Analytics

Module 5: Machine Learning and AI

Module 6: DevOps and Monitoring

Module 7: Advanced GCP Topics

Module 8: Final Project

© Copyright 2026. All rights reserved