AlpinaShop's load balancer is working: a global anycast IP, HTTPS, the MIG behind it and the images served straight from the bucket. But the inefficiency we closed the previous lesson with is still there. Every time a customer in Lisbon opens a backpack's product page, the same bytes leave europe-west1 again and cross the peninsula. With a 60 GB catalogue, hundreds of thousands of image requests a month and customers spread across Europe, that is latency the customer notices and egress the bill accumulates.
Cloud CDN solves precisely that problem, and it does so without you having to change a line of the application or move a single object. It is not a separate product with its own console, its own IP and its own domain: it is a checkbox enabled on the backend service or the backend bucket you already built in 03-02. That design detail is what makes this lesson short on commands and long on judgement: enabling it costs one --enable-cdn; doing it well consists of understanding what gets cached, with which key, for how long and how it is measured.
In this lesson you are going to enable caching on bb-catalogo-imagenes and bs-catalogo-web, choose the right cache mode for each, stop the marketing campaigns' utm_* parameters wrecking the hit ratio, decide a TTL policy consistent with the immutable names you already adopted in 02-02, learn why invalidation should not be your everyday tool, and measure the result in milliseconds and in euros.
Contents
- What a CDN is and exactly what problem it solves
- Google's network of points of presence
- Where Cloud CDN lives: on the load balancer, not beside it
- Enabling it on the backend bucket and on the backend service
- Cache modes:
CACHE_ALL_STATIC,USE_ORIGIN_HEADERS,FORCE_CACHE_ALL - The cache key and the
utm_*parameter disaster - TTL: who decides how long an object lives at the edge
- Immutable names: the strategy that holds all of the above together
- Cache invalidation: the emergency button
- Private content in the cache: signed URLs and signed cookies
- Negative caching, stale content and other useful options
- Measurement: hit ratio, the
AgeandViaheaders, and debugging withcurl - The real saving: latency and euros
- When Cloud CDN is not enough: Media CDN
- What a CDN is and exactly what problem it solves
A content delivery network (CDN) is a set of cache servers spread out geographically that keep copies of your content close to the users. When a client asks for an object, the nearest node serves it instead of the origin server.
It sounds simple, but it is worth being precise about the three distinct problems it solves, because they are not the same and they are measured differently:
| Problem | Without a CDN | With a CDN |
|---|---|---|
| Latency | The object travels from europe-west1 to Lisbon on every request |
It is served from the Lisbon point of presence: tens of milliseconds less |
| Egress cost | Every byte served is billed as internet egress from the bucket | Cache hits are billed at the CDN rate, which is cheaper |
| Load on the origin | The MIG and the bucket serve every request | The origin only sees the cache misses: fewer instances, fewer operations |
There is a fourth effect, less obvious and very valuable: spike absorption. When AlpinaShop launches the autumn campaign and the home page receives ten times the usual traffic, if the home page and its images are cached, the MIG barely notices. The autoscaling you configured in 02-01 will still be there as a safety net, but it will trigger far less often.
And one limitation to be clear about from the start: a CDN only speeds up what can be cached. A specific customer's shopping cart page, the result of a payment or Lucía's dashboard are never cached. A CDN does not make your application faster; it makes most of the traffic never reach your application.
- Google's network of points of presence
Google maintains more than two hundred cache locations around the world, on top of dozens of regions and network points of presence. Many of those caches are inside the operators' own networks (what Google calls Google Global Cache), that is, physically at the operator the client is connected to.
What matters for your design is not the number, but the mechanics:
flowchart LR
C1([Lisbon customer])
C2([Helsinki customer])
P1["Lisbon PoP<br/>cache"]
P2["Helsinki PoP<br/>cache"]
LB["Global load balancer<br/>alpinashop-lb-ip"]
O1["Bucket alpinashop-catalogo<br/>europe-west1"]
O2["MIG alpinashop-web-mig<br/>europe-west1"]
C1 -->|1. GET /imagenes/...| P1
C2 -->|1. GET /imagenes/...| P2
P1 -->|2. cache miss| LB
P2 -->|2. hit: goes no further| P2
LB --> O1
LB --> O2
P1 -.->|3. cached response| C1
- The client resolves
alpinashop.exampleand reaches the anycast IP. That IP is announced from every point of presence, so the client enters through the nearest one. - At that point of presence the cache is consulted. If there is a hit, the response leaves from there and never reaches
europe-west1. - If there is a miss, the request travels over Google's private network to the origin — not over the public internet — the response is stored at the edge and delivered to the client.
Two practical consequences:
- The cache is not one, there are many. An object can be warm in Madrid and cold in Warsaw. That is why the hit ratio is never 100 %, and nor should it be.
- Google collapses concurrent requests for the same object (request coalescing, on by default): if a thousand clients ask at once for an image that is not cached, the origin receives one request, not a thousand. This avoids the stampede effect at launches.
- Where Cloud CDN lives: on the load balancer, not beside it
This is the idea that causes the most confusion for people arriving from other platforms. With many providers the CDN is a separate service: you create a "distribution", they give you a host name of your own and you point your DNS at it. Not in Google Cloud.
Cloud CDN is a property of a backend of the global external Application Load Balancer. It is enabled like this:
- On a backend bucket (
bb-catalogo-imagenes) → it caches Cloud Storage objects. - On a backend service (
bs-catalogo-web) → it caches whatever the MIG instances return, or tomorrow the Cloud Run serverless NEG (DA-001).
Several very convenient things follow from that:
- You do not change the DNS or the IP. You carry on with
alpinashop-lb-ipand the same certificate. - The same URL map decides what is cached and what is not, because each path goes to a different backend and each backend has its own cache configuration.
- Cloud Armor (03-05) is evaluated at the edge, so the security rules also apply to requests that end up being served from the cache. The cache is not a hole through which unfiltered traffic sneaks in.
- The load balancer logs you enabled in 03-02 already carry the cache fields. There is no separate observability system.
Remember how the alpinashop-url-map map ended up:
| Path | Backend | Cache it? |
|---|---|---|
/imagenes/* |
bb-catalogo-imagenes (bucket) |
Yes, aggressively: they are immutable files |
/estatico/* |
bb-catalogo-imagenes (bucket) |
Yes: versioned CSS and JS |
/ and product pages |
bs-catalogo-web (MIG) |
Yes, carefully: HTML that changes |
/carrito, /cuenta, /pago |
bs-catalogo-web (MIG) |
Never: per-user content |
That table is the lesson's design decision. Everything else is parameters.
- Enabling it on the backend bucket and on the backend service
We start with the easy, highest-impact part: the images.
gcloud config set project alpinashop-prod
# CDN on the backend bucket that serves /imagenes/* and /estatico/*
gcloud compute backend-buckets update bb-catalogo-imagenes \
--enable-cdn \
--cache-mode=CACHE_ALL_STATIC \
--default-ttl=3600 \
--max-ttl=31536000 \
--client-ttl=3600 \
--negative-cachingWhat each option does:
--enable-cdn: turns the cache on. It is the only mandatory one; the rest are adjustments.--cache-mode: the general policy (section 5).--default-ttl: how long a response is kept if the origin says nothing.--max-ttl: the ceiling. Even if the origin asks for a year, this value will not be exceeded. Here we leave it at a year because we trust our ownCache-Control.--client-ttl: the maximummax-agecommunicated to the browser. It lets you cache a lot at the edge and little at the client, which is very useful (section 7).--negative-caching: caches errors too (section 11).
Now the MIG's backend service. Here we are more conservative, because behind it there is an application generating HTML:
gcloud compute backend-services update bs-catalogo-web \
--global \
--enable-cdn \
--cache-mode=CACHE_ALL_STATIC \
--default-ttl=300 \
--max-ttl=3600 \
--client-ttl=60 \
--serve-while-stale=86400 \
--compression-mode=AUTOMATIC--default-ttl=300: five minutes for the static content the application serves that comes with no headers of its own.--serve-while-stale=86400: if the origin does not respond, the edge can carry on serving the expired copy for up to 24 hours while it revalidates. This option turns a MIG outage into a degradation instead of a 502 error for cached content. It is one of the best benefit/effort ratios in the whole module.--compression-mode=AUTOMATIC: the edge compresses with gzip or Brotli the text responses the origin sends uncompressed, saving bandwidth without touching gunicorn.
Check the state with:
gcloud compute backend-buckets describe bb-catalogo-imagenes \
--format="yaml(cdnPolicy, enableCdn)"
gcloud compute backend-services describe bs-catalogo-web --global \
--format="yaml(enableCDN, cdnPolicy)"CDN configuration changes take a few minutes to propagate to every point of presence. And take note: changing the configuration does not flush the cache. The objects already stored stay there with their original TTL.
- Cache modes:
CACHE_ALL_STATIC, USE_ORIGIN_HEADERS, FORCE_CACHE_ALL
CACHE_ALL_STATIC, USE_ORIGIN_HEADERS, FORCE_CACHE_ALLThe cache mode answers one question: who decides what gets cached, the origin or the CDN?
| Mode | What it caches | Who is in charge | Risk | When to use it |
|---|---|---|---|---|
USE_ORIGIN_HEADERS |
Only what the origin explicitly marks as cacheable | The origin, always | None; if you get it wrong, it simply is not cached | Applications that already emit correct Cache-Control and want total control |
CACHE_ALL_STATIC |
Static content by content type and extension, even if the origin says nothing; respects the origin for the rest | Shared | Low: private and no-store are respected |
The sensible default. The one AlpinaShop uses |
FORCE_CACHE_ALL |
Every successful response, ignoring private and no-store |
The CDN, ignoring the origin | High: it can serve one user's page to another | Only on backends serving exclusively immutable public content |
Details that matter:
CACHE_ALL_STATICconsiders content "static" by its MIME type and its extension: images, video, CSS, JavaScript, fonts, PDF. HTML does not fall into that category, so AlpinaShop's product pages will not be cached unless the application asks for it with its ownCache-Control. That is exactly what we want: explicit control over the HTML.- In any mode, a response with
Cache-Control: private,no-storeorSet-Cookieis not cached (except underFORCE_CACHE_ALL). That is the safety net that prevents the classic disaster. FORCE_CACHE_ALLon a backend that serves session pages is probably the worst security incident you can cause with a checkbox. If you ever need it, apply it to a dedicated backend service that only public paths reach, never to the one serving the whole site.
For AlpinaShop the decision is:
bb-catalogo-imagenes→CACHE_ALL_STATIC. Everything there is static by definition and already comes with a one-yearCache-Controlfrom 02-02.bs-catalogo-web→CACHE_ALL_STATIC, and let the Flask application decide path by path.
This is how Dani declares it in the application:
# app.py — explicit cacheability control per path
@app.route("/producto/<sku>")
def product_page(sku):
product = repository.get(sku)
response = make_response(render_template("producto.html", producto=product))
# Public and cacheable for 5 minutes at the edge, 1 minute in the browser.
# 's-maxage' is read by the CDN; 'max-age' is read by the browser.
response.headers["Cache-Control"] = "public, max-age=60, s-maxage=300"
return response
@app.route("/carrito")
def cart():
response = make_response(render_template("carrito.html"))
# Never, nowhere, under no circumstances.
response.headers["Cache-Control"] = "private, no-store"
return responseA rule worth committing to memory: if a response depends on who is asking, it carries private, no-store. Put it there explicitly and do not rely on the cache mode saving you.
- The cache key and the
utm_* parameter disaster
utm_* parameter disasterThe cache key is the string with which the CDN identifies a stored object. If two requests produce the same key, the second is a hit. If they produce different keys, the second is a miss even though the content is identical.
By default the key includes:
- The protocol (
httporhttps). - The host (
alpinashop.example). - The path (
/imagenes/productos/MOCH-2210/web/frontal-800.a3f9c1.webp). - The complete query string (
?utm_source=newsletter&utm_campaign=otono26).
And there is the problem. AlpinaShop's marketing team launches a campaign and the image URLs start arriving like this:
/imagenes/productos/MOCH-2210/web/frontal-800.a3f9c1.webp?utm_source=newsletter&utm_medium=email&utm_campaign=otono26 /imagenes/productos/MOCH-2210/web/frontal-800.a3f9c1.webp?utm_source=instagram&utm_medium=social&utm_campaign=otono26 /imagenes/productos/MOCH-2210/web/frontal-800.a3f9c1.webp?fbclid=IwAR3x9...
They are the same object with three different keys. Each variant causes a cache miss, a read from the bucket and a cache fill charge. With a big campaign, the hit ratio collapses from 92 % to 40 % and nobody understands why the bill went up exactly when the visitors arrived.
The solution is to exclude those parameters from the key:
# Backend bucket: the images depend on NO parameter at all
gcloud compute backend-buckets update bb-catalogo-imagenes \
--cache-key-query-string-blacklist=utm_source,utm_medium,utm_campaign,utm_term,utm_content,gclid,fbclid,msclkid,ref
# Backend service: the opposite approach, an allowlist.
# Only these parameters form part of the key; the rest are ignored.
gcloud compute backend-services update bs-catalogo-web --global \
--cache-key-query-string-whitelist=pagina,orden,talla,colorTwo strategies, and it is worth understanding why each one is used where it is:
| Strategy | Command | Advantage | Risk |
|---|---|---|---|
| Blocklist (exclude) | --cache-key-query-string-blacklist |
Easy to get started with; breaks nothing existing | Every new tracking parameter has to be added by hand |
| Allowlist (include only) | --cache-key-query-string-whitelist |
Immune to new parameters; maximum ratio | If you forget a parameter that does change the content, you serve the wrong page |
For the bucket, the blocklist is more than enough; in reality no parameter affects an image. For the backend service, the allowlist is more powerful but requires Dani to review the list every time he adds a new filter to the catalogue. It is a decision to be documented, not improvised.
If the images really do not depend on any parameter, the cleanest option is simply:
gcloud compute backend-buckets update bb-catalogo-imagenes \
--cache-key-query-string-blacklist="" # equivalent to ignoring the whole queryOther components that can be added to the key:
--cache-key-include-http-header=X-Alpina-Region: useful if you serve variants by header. Each distinct value multiplies the cache entries, so use it only with headers of very low cardinality.--cache-key-include-named-cookie=idioma: variants by a specific cookie. Same warning.--no-cache-key-include-protocol: unifies HTTP and HTTPS into the same entry. Since at AlpinaShop all HTTP is redirected to HTTPS (03-02), it makes no difference.
And a warning that points the other way: if the origin returns a Vary header with values the CDN does not support, the response is simply not cached. Vary: Accept-Encoding is safe; Vary: User-Agent kills the cache completely, because there are tens of thousands of different user agents.
- TTL: who decides how long an object lives at the edge
There are five values involved and they are constantly confused. This is the chain of command:
| Value | Where it is defined | Who it affects | What it does |
|---|---|---|---|
Cache-Control: max-age |
Origin header | Browser and CDN | General TTL |
Cache-Control: s-maxage |
Origin header | Shared caches only (the CDN) | Takes priority over max-age at the edge |
--default-ttl |
Backend configuration | CDN | TTL when the origin says nothing |
--max-ttl |
Backend configuration | CDN | Ceiling: it trims whatever the origin asks for |
--client-ttl |
Backend configuration | Browser | Ceiling on the max-age sent to the client |
The actual sequence, for a response arriving from the origin with Cache-Control: public, max-age=31536000, immutable and a backend configured with --max-ttl=86400 --client-ttl=3600:
- The CDN reads
max-age= 31,536,000 s. - It applies
--max-ttl: it stores the object for 24 hours, not a year. - It applies
--client-ttl: it rewrites the header towards the client asmax-age=3600, one hour.
The combination of a high s-maxage + a low max-age is the most useful and the least used:
It means: "CDN, keep it a day; browser, keep it a minute". The advantage: if you need to change the content, you invalidate the CDN cache (which you do control) and within 60 seconds every browser in the world sees the new thing. You can never invalidate somebody else's browser cache; that is why it is best for the client to cache little and the edge to cache a lot — except when the file name is immutable, which is the next case.
- Immutable names: the strategy that holds all of the above together
In 02-02 you made a decision that now pays its interest: AlpinaShop's images are stored with a name that is never reused, with a hash or version embedded in it:
productos/MOCH-2210/web/frontal-800.a3f9c1.webp productos/MOCH-2210/web/frontal-800.b7e204.webp ← the new photo is a different object
And they are uploaded with Cache-Control: public, max-age=31536000, immutable.
The consequence is enormous and deserves stating explicitly:
If the name is immutable, the invalidation problem disappears. Publishing a new photo does not consist of updating a cached object, but of publishing an object nobody has ever asked for. The HTML page points at the new name and the cache serves the new object from the very first moment, while the old one expires on its own.
With that scheme, the correct policy for the bucket is to cache for a year for real, in the client too:
gcloud compute backend-buckets update bb-catalogo-imagenes \
--cache-mode=CACHE_ALL_STATIC \
--default-ttl=3600 \
--max-ttl=31536000 \
--client-ttl=31536000Here we do want a long client-ttl: a returning customer's browser will never ask for the image again. Note the contrast with the previous section: the client TTL is long when the name is immutable and short when it is not. That is the whole rule.
And the corollary for the HTML: a product page cannot have a long TTL, because its URL is stable (/producto/MOCH-2210). That is why it carries s-maxage=300. Content with a stable URL → short TTL. Content with a versioned URL → one-year TTL.
- Cache invalidation: the emergency button
Sometimes you have to remove something from the edge before it expires: a wrong price was published, an image with the logo wrong, an incorrect piece of legal text.
# Invalidate a specific object
gcloud compute url-maps invalidate-cdn-cache alpinashop-url-map \
--path="/imagenes/productos/MOCH-2210/web/frontal-800.a3f9c1.webp"
# Invalidate a whole prefix (the wildcard is only valid at the end)
gcloud compute url-maps invalidate-cdn-cache alpinashop-url-map \
--path="/imagenes/productos/MOCH-2210/*"
# Invalidate the whole site: the sledgehammer. Avoid it.
gcloud compute url-maps invalidate-cdn-cache alpinashop-url-map \
--path="/*" --async
# Follow the operation
gcloud compute operations list --global --filter="operationType~invalidate" --limit=5Notice that invalidation is run against the URL map, not against the backend: it propagates to every backend that map routes to.
Why it should not be your everyday mechanism:
- It is slow in incident terms. The command has to be propagated to every point of presence in the world; the order of magnitude is minutes, not seconds. During a content crisis, those minutes feel long.
- It is subject to quota. There is a limit on invalidations per project and unit of time. A deployment that invalidates
/*on every release hits the quota sooner than you think. - A
/*destroys the hit ratio. After a total invalidation, all the traffic in the world hits the origin at once. You have turned a routine release into an involuntary load test. - It requires elevated permissions, typically
roles/compute.loadBalancerAdmin, which you will not want to hand out lightly (03-04).
The correct alternative is the one in section 8: versioned names for the assets, a short TTL for the HTML. Invalidation is reserved for mistakes, which is what it is for.
- Private content in the cache: signed URLs and signed cookies
The alpinashop-catalogo bucket is private (you created it with --public-access-prevention in 02-02) and the backend bucket reads it with its own permissions, so the catalogue images are public through the load balancer but not directly. Perfect for the catalogue.
Now picture the following AlpinaShop case: the videos for the technical climbing guides should only be watched by customers who have bought the in-person course. They are large files, they benefit greatly from caching, but they cannot be public.
That is what Cloud CDN signed URLs and signed cookies are for. The difference from Cloud Storage signed URLs (02-02) is important: those are validated by Cloud Storage and break the cache (each signature is a different URL); these are validated by the CDN edge and allow caching.
# 1) Generate a 16-byte signing key in base64url
head -c 16 /dev/urandom | base64 | tr +/ -_ > clave-cdn.txt
# 2) Register it on the backend
gcloud compute backend-buckets add-signed-url-key bb-catalogo-guias \
--key-name=clave-guias-2026-01 \
--key-file=clave-cdn.txt
# 3) State how long the signed response is cached
gcloud compute backend-buckets update bb-catalogo-guias \
--signed-url-cache-max-age=3600
# 4) Store the key in Secret Manager (03-06) and delete the local file
gcloud secrets create cdn-clave-guias --data-file=clave-cdn.txt
shred -u clave-cdn.txtSigned cookie versus signed URL:
| Signed URL | Signed cookie | |
|---|---|---|
| Scope | One specific object | A whole path prefix |
| How it arrives | Inside the link | A Set-Cookie header with the name Cloud-CDN-Cookie |
| Good for | A one-off download | A gallery or a player with many files |
| Effect on the cache | Caches well: the key ignores the signature parameters | Caches well, and a single cookie covers the whole session |
For the guide videos, the cookie is the right option: the Flask application validates the purchase, issues a signed cookie valid for /guias/ for two hours and the player requests the segments with no further authentication. The edge validates the signature and serves from the cache.
Warning. Signing keys are cryptographic material: they give access to paid content. They must live in Secret Manager (03-06), be rotated periodically (several keys can be registered at once on the same backend precisely so you can rotate without interruption) and never appear in the repository. If the protected content has contractual or rights implications, have a security professional review the design before production.
- Negative caching, stale content and other useful options
Negative caching. Without it, if a crawler asks for ten thousand non-existent URLs, your origin generates ten thousand 404s. With it, the edge remembers the error:
gcloud compute backend-services update bs-catalogo-web --global \
--negative-caching \
--negative-caching-policy='404=120,410=300,500=0,502=0,503=0'The interesting nuance is the 5xx codes set to 0: we do not want to cache our own server errors. If the MIG returns a 502 during a deployment and we cache it for two minutes, we have multiplied the deployment's impact. 404s, yes: a product that does not exist will still not exist in two minutes.
Serving stale content (--serve-while-stale). We already enabled it: during that window, if the origin does not respond or returns an error, the edge serves the expired copy and revalidates in the background. It is the difference between "the site is a little out of date" and "the site is down".
Bypassing the cache on demand. Very useful so that Dani can debug in production without touching the configuration:
gcloud compute backend-services update bs-catalogo-web --global \
--bypass-cache-on-request-headers=x-alpina-debugFrom then on, any request with that header goes straight to the origin. Since it can bypass the cache for the whole site, the header should be hard to guess and should not be documented outside the team.
- Measurement: hit ratio, the
Age and Via headers, and debugging with curl
Age and Via headers, and debugging with curlWithout measurement, all of the above is faith. There are three levels of observation.
Level 1: the HTTP response. The fastest way of knowing whether something is being cached:
HTTP/2 200 content-type: image/webp cache-control: public, max-age=31536000, immutable age: 4127 via: 1.1 google
How to read it:
| Header | What it means |
|---|---|
Age: 4127 |
The copy has been at the edge for 4127 seconds. Age present and greater than zero = cache hit. |
No Age |
Cache miss: the response comes from the origin |
Via: 1.1 google |
It has gone through Google's proxy infrastructure. It appears whether there is a hit or not |
Cache-Control |
What the edge tells the browser, already trimmed by --client-ttl |
The classic test: request it twice in a row. The first has no Age; the second does. If the second one does not either, there is a cacheability problem and it is time for level 2.
Level 2: the load balancer logs. The cache fields are already there:
gcloud logging read '
resource.type="http_load_balancer"
AND httpRequest.requestUrl:"/imagenes/"
' --project=alpinashop-prod --limit=20 --freshness=1h \
--format="table(
httpRequest.requestUrl,
httpRequest.cacheLookup,
httpRequest.cacheHit,
httpRequest.cacheFillBytes,
jsonPayload.statusDetails
)"cacheLookup: false→ the cache was not even consulted. Almost always it means the CDN is not enabled on that backend, or the request is not cacheable (a method other thanGET, anAuthorizationheader…).cacheLookup: true, cacheHit: false→ it was looked up and was not there: a legitimate miss or a fragmented cache key (section 6).cacheFillBytes→ the bytes fetched from the origin to fill the cache. If this is persistently high, your ratio is bad.statusDetails: response_from_cache→ confirmation of a hit.
Level 3: metrics. The hit ratio is obtained from loadbalancing.googleapis.com/https/request_count grouped by the cache_result label, whose values are HIT, MISS, DISABLED and revalidation variants. Marta's dashboard (06-04) should carry:
| Metric / view | Target for AlpinaShop |
|---|---|
Hit ratio on /imagenes/* |
> 90 % |
| Overall hit ratio | > 60 % |
Monthly cacheFillBytes |
Stable or falling despite traffic growth |
p95 of image total_latencies |
Well below the HTML p95 |
A guide to debugging a persistent miss. If an object is never cached, check in this order:
--enable-cdnreally is on the right backend (describe, not memory).- The response does not carry
Cache-Control: private,no-storeorno-cache. - The response does not carry
Set-Cookie. It is the most frequent cause in Flask applications: merely touchingsessionin the view is enough for a cookie to be emitted and the response to stop being cacheable. - There is no
Varywith high-cardinality headers (User-Agent,Cookie). - The request is a
GETand carries noAuthorization. - The cache key is not fragmented by query parameters.
- The response comes with
Content-Lengthor a determinable length.
Point 3 is fixed in two lines and usually explains half the cases:
@app.after_request
def strip_cookie_on_static(response):
"""Stops an accidental session cookie making the catalogue uncacheable."""
if request.path.startswith(("/estatico/", "/imagenes/")):
response.headers.pop("Set-Cookie", None)
return response
- The real saving: latency and euros
Let us run an indicative calculation for AlpinaShop. Suppose 2 TB a month of images served to European customers and a hit ratio of 90 %.
| Item | Without a CDN | With a CDN at 90 % |
|---|---|---|
| Internet egress from the bucket | 2,000 GB × ~$0.11/GB ≈ $220 | 200 GB (the misses) → cache fill ≈ $2–20 |
| Egress from the CDN cache | — | 1,800 GB × ~$0.08/GB ≈ $144 |
| Cache lookups | — | ~2 M requests × $0.0075/10,000 ≈ $1.5 |
| Class B operations on the bucket | Millions of reads | Only the fill ones: two orders of magnitude fewer |
| Approximate total | ≈ $220 | ≈ $150–170 |
Order-of-magnitude prices for illustration. Egress, cache fill and lookup rates vary by continent and by volume tier, and change over time: always check them in the calculator and in the official pricing documentation before presenting a figure to anyone.
Honest conclusions from that table, which is more interesting than a brochure's "you save 90 %":
- The saving on the bill is real but moderate: around 25-30 %. Egress from the cache is paid for too.
- The saving grows with traffic, because the origin stops scaling with the visitors: fewer MIG instances, fewer bucket operations, fewer IOPS.
- The big win is latency. A customer in Lisbon goes from a time to first byte of around 40-60 ms to under 15 ms; one in Helsinki, from over 60 ms to similar figures. On a product page with twelve images, that translates into hundreds of milliseconds of full page load, which is exactly what Google measures in its page experience signals and what shows up in conversion.
- And there is a saving that appears on no bill: campaign peaks stop being a capacity problem.
- When Cloud CDN is not enough: Media CDN
Google Cloud has a second delivery product, Media CDN, built on the same infrastructure that serves YouTube. It is not "Cloud CDN improved": it is a different product, with a different configuration model (EdgeCacheService, EdgeCacheOrigin, EdgeCacheKeyset) and a different audience.
| Cloud CDN | Media CDN | |
|---|---|---|
| Configured on | The existing global load balancer | Its own resources, independent of the load balancer |
| Designed for | Websites, APIs, images, static assets | Video on demand and live, mass downloads |
| Typical volume | Any | Tens of TB a month upwards |
| Extras | Those covered in this lesson | Advanced routing rules, better economics at large scale |
For AlpinaShop, Cloud CDN is the right choice and will continue to be. If one day the technical guides section turns into a video platform with thousands of hours and tens of TB a month, then it would be time to evaluate Media CDN.
Common Mistakes and Tips
- Believing Cloud CDN is a separate product and looking for where to create the "distribution". It is a checkbox on the backend of the load balancer you already have.
- Enabling
FORCE_CACHE_ALLon the backend that serves the whole site. It is the fastest route to serving one customer another customer's shopping cart. If you need it, do it on a backend dedicated to public paths. - Ignoring campaign parameters.
utm_*andfbclidfragment the cache key and sink the ratio on exactly the day the traffic arrives. Configure the blocklist before the campaign, not after. - Confusing
max-agewiths-maxage. The second is only read by shared caches and takes priority at the edge. Lowmax-age+ highs-maxageis the pair you want for HTML. - Setting a long
--client-ttlon content with a stable URL. You cannot invalidate anybody's browser. A long client TTL only with immutable names. - Using invalidation as part of the deployment. It is slow, it has a quota and a
/*leaves the origin exposed. Version the names. - Forgetting about
Set-Cookie. A Flask session touched by accident in a catalogue view makes the whole path uncacheable and nobody understands the low ratio. Vary: User-Agent. It nullifies the cache completely. Review whichVaryyour application and your CMS emit.- Caching 5xx. With badly configured negative caching, a deployment with a thirty-second glitch turns into five minutes of errors for everyone. Set the 5xx to
0. - Not enabling
--serve-while-stale. It is practically free and turns outages into degradations. - Measuring "by eye" from the browser. The browser has its own cache and will deceive you. Debug with
curl -sIand with thecacheHit/cacheLookupfields in the log. - Tip: test from several locations. A hit in Madrid says nothing about Warsaw. Each point of presence has its own cache and its own warm-up.
- Tip: separate what is cacheable from what is not in the URL map. It is easier to reason about
/estatico/*with a one-year cache and/cuenta/*with no cache than about a single backend with subtle rules.
Exercises
Exercise 1 — Diagnosing a 38 % hit ratio
Marta opens the dashboard a week after enabling the CDN and sees a 38 % ratio on /imagenes/*, when she expected more than 90 %. She gathers this data:
curl -sIon a specific image, twice in a row: the second response does carryAge: 61.- The log shows many requests with
cacheLookup: true,cacheHit: falseand URLs ending in?utm_source=.... - Other requests show
cacheLookup: falseon/imagenes/promo/*paths. - The
Cache-Controlof the images under/imagenes/promo/ispublic, max-age=300.
Explain each symptom separately and write the commands that fix the situation.
Exercise 2 — Cache policy for three new paths
AlpinaShop adds three paths. Define for each one: cache mode, the Cache-Control the application should emit, client and edge TTL, and which components the cache key should carry.
/api/stock/<sku>— returns JSON with the units available. It changes every few minutes and thousands of visitors query it./estatico/app.4f2b9c.js— the JavaScript bundle, with a hash in the name./cuenta/pedidos— the authenticated customer's order history.
Exercise 3 — Calculating a campaign's impact
The autumn campaign is going to multiply image traffic by five for two weeks: from 2 TB/month to a rate equivalent to 10 TB/month. Estimate the monthly egress cost with and without a CDN assuming a 90 % ratio, and answer: what other effect, more important than cost, does the CDN have in that campaign? What would you configure before it starts?
Solutions
Solution 1
They are three different problems disguised as one:
a) The specific image is being cached. The Age: 61 on the second request proves it. So the backend bucket's base configuration is correct and does not need touching.
b) The utm_* are fragmenting the key. Each campaign variant generates a new entry: cacheLookup: true, cacheHit: false is exactly this problem's signature. It is fixed by excluding the parameters from the key:
gcloud compute backend-buckets update bb-catalogo-imagenes \
--cache-key-query-string-blacklist=utm_source,utm_medium,utm_campaign,utm_term,utm_content,gclid,fbclid,msclkidc) /imagenes/promo/* is not even looked up in the cache. cacheLookup: false means the request never got as far as a lookup. If those paths are routed in the URL map towards a different backend (for example a bs-promociones created for the campaign), that backend does not have --enable-cdn. It is checked and fixed like this:
gcloud compute url-maps describe alpinashop-url-map \
--format="yaml(pathMatchers)"
gcloud compute backend-services update bs-promociones --global \
--enable-cdn --cache-mode=CACHE_ALL_STATIC \
--default-ttl=3600 --max-ttl=86400d) The max-age=300 on the promo images is not an error, but it is inconsistent: if the promotional images also have immutable names, they should carry a year. If they do not, the underlying fix is to adopt the same naming convention as the rest of the catalogue, not to raise the TTL blindly.
Solution 2
| Path | Mode | App's Cache-Control |
Edge | Client | Cache key |
|---|---|---|---|---|---|
/api/stock/<sku> |
CACHE_ALL_STATIC (with an explicit header) |
public, max-age=0, s-maxage=60 |
60 s | 0 s | Protocol + host + path. No query parameters |
/estatico/app.4f2b9c.js |
CACHE_ALL_STATIC |
public, max-age=31536000, immutable |
1 year | 1 year | Path only; ignore the whole query |
/cuenta/pedidos |
Irrelevant | private, no-store |
Never | Never | Not applicable |
The reasoning:
- Stock. It is JSON,
CACHE_ALL_STATICwould not cache it on its own, so the explicit header is mandatory.s-maxage=60withmax-age=0is the key combination: the edge absorbs thousands of requests a minute and the browser always asks, so a stock change is seen at most a minute late. Sixty seconds of lag on a units counter is acceptable; a minute of stale data at the payment step would not be, and that is why the definitive stock check is done at the moment the order is confirmed, not by reading this API. - JS bundle. A hashed name → textbook case: a year everywhere. Publishing a new version means publishing a different name.
- Order history. An explicit
private, no-store. And very importantly: never on a backend withFORCE_CACHE_ALL. It will also carry a sessionSet-Cookie, which reinforces the uncacheability, but you must not depend on that side effect: the header is set deliberately.
Solution 3
Indicative cost at 10 TB/month:
- Without a CDN: 10,000 GB × ~$0.11/GB ≈ $1,100/month.
- With a CDN at 90 %: 9,000 GB from cache × ~$0.08/GB ≈ $720, plus 1,000 GB of fill (≈$10-40) and the lookups (a few dollars) ≈ $740-770/month. A saving of roughly 30 %.
The most important effect is not the cost, it is the load on the origin. Without a CDN, that fivefold traffic is absorbed by the MIG and the bucket: autoscaling would climb towards the limit of 10 instances, the compute cost would shoot up and latency would get worse on exactly the day of the highest sales. With a 90 % hit rate, the origin only sees 10 % of the image traffic; the MIG stays close to its usual size and the customer experience is better precisely when it matters most.
What to configure before the campaign starts:
- The blocklist of parameters
utm_*,gclidandfbclid. Without this, the campaign sabotages itself. --serve-while-stale=86400onbs-catalogo-web, so that an origin problem under load degrades rather than takes the shop down.- Negative caching with the 5xx at
0and 404s at a couple of minutes. - Verify with
curl -sIthat the campaign images returnAgeon the second request, before the launch. - A dashboard with the hit ratio and
cacheFillBytesvisible during the campaign, and an alert if the ratio drops below 80 %. - Do not invalidate the cache during the campaign except in a real emergency: it would mean handing the traffic peak to the origin.
Conclusion
AlpinaShop no longer sends the same bytes to Lisbon over and over again. You have understood that Cloud CDN is not a separate product but a property of the load balancer you built in 03-02: it is enabled with --enable-cdn on bb-catalogo-imagenes and on bs-catalogo-web, without touching the DNS, the IP, the certificate or the application, and taking advantage of the same logs and the same metrics.
You know how to choose the cache mode with judgement — CACHE_ALL_STATIC as the sensible default, USE_ORIGIN_HEADERS when the application really is in charge, and FORCE_CACHE_ALL only on backends dedicated to public content — and why a response with private, no-store or Set-Cookie must never be cached. You have mastered the cache key, which is where the ratio is won or lost: you have excluded utm_*, gclid and fbclid so that the autumn campaign does not destroy the work, and you know the trade-off between blocklist and allowlist. You have put the five TTL values in order — max-age, s-maxage, --default-ttl, --max-ttl, --client-ttl — and you have seen why the immutable names adopted in 02-02 are the piece that makes invalidation unnecessary, that slow, quota-limited operation capable of leaving the origin out in the open. You know how to serve paid content from the cache with signed cookies, how to cache 404s but never 500s, how to keep serving during an outage with --serve-while-stale, and how to debug a cache miss by reading Age, Via and the log's cacheHit/cacheLookup fields. And you have an honest calculation of the saving: around 30 % of the egress bill, far more in origin load, and a latency improvement that shows on every product page.
With this, AlpinaShop's edge is complete as far as performance goes: network, load balancing and caching. But there is a problem we have been putting off lesson after lesson. In 03-01 we said the firewall does not distinguish between people and that "only Lucía and Marta can get in over SSH" is solved somewhere else. In 02-02 we limited Lucía to a bucket prefix with a condition we did not explain. In this very lesson we have said that invalidating the cache requires an elevated role that is best not handed around. All of that points to the same place: who can do what, on which resource. In the next lesson, 03-04, Identity and Access Management (IAM), we build AlpinaShop's real permission model: groups instead of people, predefined roles instead of Editor, a bespoke custom role for Lucía, service accounts with no downloaded keys, impersonation, conditions, and the tools for answering any cloud administrator's most frequent question: "why can't this user do this?".
Google Cloud Platform (GCP) Course
Module 1: Introduction to Google Cloud Platform
- What is Google Cloud Platform?
- Setting Up Your GCP Account
- A Tour of the GCP Console
- Projects, Resource Hierarchy and Billing
- Regions, Zones and the Shared Responsibility Model
- Cloud Shell and the gcloud CLI
Module 2: Core GCP Services
- Compute Engine: Virtual Machines on Google Cloud
- Cloud Storage: Object Storage
- Cloud SQL: Managed Relational Databases
- App Engine: Platform as a Service
- Google Kubernetes Engine (GKE)
- NoSQL Databases: Firestore, Bigtable and Spanner
- How to Choose the Right Compute Service
Module 3: Networking and Security
- VPC Networks
- Cloud Load Balancing
- Cloud CDN
- Identity and Access Management (IAM)
- Cloud Armor
- Secrets and Encryption: Secret Manager and Cloud KMS
- Cloud DNS, TLS Certificates and Publishing Services Securely
Module 4: Data and Analytics
- BigQuery: The Analytical Data Warehouse
- Cloud Dataflow: Batch and Streaming Data Processing
- Cloud Dataproc: Managed Spark and Hadoop
- Cloud Pub/Sub: Asynchronous Messaging
- Cloud Data Fusion: Code-Free Data Integration
- Orchestrating Pipelines with Cloud Composer and Workflows
- Data Governance and Dashboards with Dataplex and Looker Studio
Module 5: Machine Learning and AI
- Vertex AI: The Machine Learning Platform on GCP
- AutoML: Custom Models Without Writing Code
- TensorFlow on GCP: Training and Serving Models
- Natural Language API
- Vision API
- Generative AI on Vertex AI: Gemini Models and Embeddings
- MLOps: From Model to Product with Vertex AI Pipelines
Module 6: DevOps and Monitoring
- Cloud Build: Continuous Integration on GCP
- Cloud Source Repositories and Source Code Management
- Cloud Functions: Serverless Functions
- Cloud Monitoring (formerly Stackdriver): Metrics, Dashboards and Alerts
- Cloud Deployment Manager and Native Infrastructure as Code
- Cloud Logging and Cloud Trace: Logs, Traces and Diagnostics
- Terraform on GCP: Infrastructure as Code in Practice
Module 7: Advanced GCP Topics
- Hybrid and Multicloud with Anthos
- Serverless Computing with Cloud Run
- Advanced Networking: Shared VPC, Peering and Hybrid Connectivity
- Security Best Practices
- Cost Management and Optimization
- Reliability: SLOs, High Availability and Disaster Recovery
- Governance at Scale: Organization, Policies and Auditing
