AlpinaShop's network is already built: alpinashop-vpc with its subnets, the firewall that only lets through what is needed, Cloud NAT for egress and the database on a private IP. But the catalogue is still served from one particular machine's address. If that machine goes down, the shop goes down. If traffic grows, there is nothing to spread it. And there is no HTTPS, no domain, and no single point at which to apply security. This lesson puts in front of it the piece that solves all of that at once: the load balancer.
Cloud Load Balancing is not an appliance nor a virtual machine you administer. It is a distributed service, implemented in Google's own network infrastructure, with no instances to size or scale. Its global version has a property worth understanding well because it changes the design: a single anycast IP address announced from more than a hundred points of presence around the world, so that a customer in Madrid and one in Helsinki each enter at their nearest point and travel the rest of the way over Google's private network.
On top of that, the load balancer is where everything else plugs in: the TLS certificate (03-07), the Cloud CDN cache (03-03) and the Cloud Armor application firewall (03-05) are all configured on it. Building it properly is the precondition for the next four lessons.
In this lesson you will learn to find your way around the load balancer catalogue without getting lost in the naming, you will assemble AlpinaShop's global external HTTP(S) load balancer piece by piece with gcloud, you will route /imagenes/* straight to the alpinashop-catalogo bucket and everything else to the alpinashop-web-mig MIG, and you will master health checks, which are the cause of 90 % of the problems with a freshly created load balancer.
Contents
- Why a load balancer and what problems it solves
- Google Cloud's load balancer catalogue and how to choose
- Anatomy of a global external HTTP(S) load balancer
- Groundwork: health check and named port
- The backend service over the MIG
- The backend bucket for the images
- The URL map: routing by path
- Target proxy and forwarding rule
- Health checks in depth: why they fail
- Balancing policies, capacity and session affinity
- Redirecting HTTP to HTTPS
- Global scaling: why there is no need to "pre-warm"
- Internal load balancers for internal services
- Serverless NEGs: Cloud Run behind the same load balancer
- Load balancer logging and metrics
- Why a load balancer and what problems it solves
A load balancer does far more than spread requests around. In AlpinaShop's case it simultaneously solves:
| Current problem | How the load balancer solves it |
|---|---|
| A single VM: single point of failure | Spreads across the MIG's instances and takes failing ones out of rotation |
| Traffic only enters through one zone | The regional MIG covers three zones; the load balancer spreads across them |
| Distant customers with high latency | Anycast IP: you enter at the nearest PoP and travel over Google's network |
| No HTTPS and no certificate | Terminates TLS with managed certificates (03-07) |
| The application serves the images | A backend bucket serves alpinashop-catalogo without touching the application |
| Nowhere to put a WAF or a cache | Cloud Armor (03-05) and Cloud CDN (03-03) are applied here |
| Updating the app means an outage | Connection draining and rolling updates of the MIG |
The key idea: the load balancer is the edge of your architecture. Everything you want to do once for all traffic — encrypt, filter, cache, measure — is done at the edge.
- Google Cloud's load balancer catalogue and how to choose
This is where many people get lost, because the naming has changed over the years and the names are long. Let us sort it out with three questions.
Question 1: layer 7 or layer 4? That is, does the load balancer understand HTTP (paths, headers, cookies) or does it just move TCP/UDP packets?
Question 2: external or internal? Is it called from the internet or only from inside your VPC?
Question 3: global or regional? A single IP for the whole world or one IP per region?
With those three answers, the load balancer is determined:
| Load balancer | Layer | External/internal | Scope | Protocols | Use case |
|---|---|---|---|---|---|
| Global external Application Load Balancer | 7 | External | Global | HTTP, HTTPS, HTTP/2, HTTP/3 | AlpinaShop's. Worldwide public web, CDN, WAF |
| Regional external Application Load Balancer | 7 | External | Regional | HTTP, HTTPS | Data residency requirements in one region |
| Internal Application Load Balancer | 7 | Internal | Regional (or cross-region) | HTTP, HTTPS | Internal microservices with path-based routing |
| External proxy Network Load Balancer | 4 (proxy) | External | Global or regional | TCP, SSL | Non-HTTP protocols that need TLS termination |
| External passthrough Network Load Balancer | 4 (passthrough) | External | Regional | TCP, UDP, ESP, ICMP | Preserves the source IP; games, unusual protocols |
| Internal passthrough Network Load Balancer | 4 | Internal | Regional | TCP, UDP | High-performance internal backend; database, gRPC |
Two distinctions worth being clear about:
- Proxy versus passthrough. A proxy load balancer terminates the client's connection and opens another one towards the backend; the backend sees the load balancer's IP and receives the client's in the
X-Forwarded-Forheader. A passthrough load balancer terminates nothing: the packet arrives at the backend with the original source IP. - Global versus regional. The global one uses anycast and a single URL map for all regions; the regional one lives in a single region and so does its IP. The global one is the one that allows CDN and Cloud Armor at worldwide scale.
A practical rule for choosing: if it is public web traffic, it is the global external Application Load Balancer. You only step down from there for a specific requirement: data residency, a protocol that is not HTTP, or the need to preserve the source IP at packet level.
- Anatomy of a global external HTTP(S) load balancer
The load balancer is not an object: it is a chain of objects created separately and linked together. Understanding the chain is understanding the service.
flowchart TB
C([Customer in Europe])
IP["Global anycast IP address<br/>alpinashop-lb-ip"]
FR["Forwarding rule<br/>port 443"]
TP["Target HTTPS proxy<br/>+ TLS certificate (03-07)<br/>+ SSL policy"]
UM["URL map<br/>alpinashop-url-map"]
BS["Backend service<br/>bs-catalogo-web"]
BB["Backend bucket<br/>bb-catalogo-imagenes"]
HC["Health check<br/>hc-catalogo /salud"]
MIG["MIG alpinashop-web-mig<br/>2-10 instances, 3 zones"]
GCS[("Bucket alpinashop-catalogo")]
C --> IP --> FR --> TP --> UM
UM -->|"/imagenes/*"| BB --> GCS
UM -->|"everything else"| BS --> MIG
HC -.checks.-> MIG
BS -.uses.-> HC
Piece by piece, from the outside in:
| Piece | What it does | gcloud resource |
|---|---|---|
| Global IP address | The public anycast IP announced by the PoPs | compute addresses --global |
| Forwarding rule | Ties IP + port to a target proxy | compute forwarding-rules |
| Target proxy | Terminates the connection; in HTTPS, applies the certificate and the SSL policy | compute target-https-proxies |
| URL map | Decides, based on host and path, which backend service serves it | compute url-maps |
| Backend service | Groups backends, defines balancing policy, CDN, Cloud Armor, timeouts | compute backend-services |
| Backend | The actual group: MIG, NEG or bucket | backend-services add-backend |
| Health check | Decides which instances receive traffic | compute health-checks |
They are built from the inside out: first the health check, then the backend services, then the URL map, the proxy and finally the forwarding rule. If you try it the other way round, gcloud complains that the referenced resource does not exist.
- Groundwork: health check and named port
AlpinaShop's Flask application listens on port 8080 under gunicorn and exposes a health path. First of all, make sure that path exists and is cheap:
# app.py — AlpinaShop's health path
@app.route("/salud")
def health():
"""Health check for the load balancer.
It must be CHEAP and FAST: it is called every few seconds from
several checkers. Do not query the database here unless you want
a DB problem to take the whole fleet out of rotation.
"""
return {"status": "ok"}, 200That nuance deserves an explanation, because it is an architectural decision and not a detail:
- If
/saluddoes not query the database, a database problem leaves the backends "healthy" and customers receive 500 errors from the application. Bad, but contained. - If
/saluddoes query the database, a database problem marks all the backends as unhealthy at once, the load balancer runs out of destinations and returns 502 to everybody. Worse: you have turned a partial failure into a total outage.
The usual practice is to have two paths: a shallow /salud (for the load balancer) and a deep /salud/completo (for the monitoring in 06-04, which alerts but does not take anything out of rotation).
Now, the health check and the named port:
gcloud config set project alpinashop-prod
# HTTP health check against /salud on 8080
gcloud compute health-checks create http hc-catalogo \
--port=8080 \
--request-path=/salud \
--check-interval=10s \
--timeout=5s \
--healthy-threshold=2 \
--unhealthy-threshold=3 \
--description="Health of the AlpinaShop catalogue"
# The instance group must declare a "named port":
# the backend service will refer to the name, not the number.
gcloud compute instance-groups managed set-named-ports alpinashop-web-mig \
--region=europe-west1 \
--named-ports=http-catalogo:8080The named port is a common source of confusion. The backend service does not say "send the traffic to 8080"; it says "send the traffic to the port called http-catalogo", and each instance group translates that name into a number. The advantage is that you can change the application's port without touching the load balancer. The drawback is that if you forget to define the named port, the load balancer uses the default http (port 80), finds nothing and everything shows up as unhealthy.
The check's parameters, explained:
| Parameter | Value | Meaning |
|---|---|---|
--check-interval |
10s | How often the check runs |
--timeout |
5s | How long to wait for the reply before counting it as a failure |
--healthy-threshold |
2 | Consecutive successful checks to come back into rotation |
--unhealthy-threshold |
3 | Consecutive failures to leave rotation |
With these values, a broken instance takes about 30 seconds to leave rotation and a recovered one about 20 to come back. Lowering the interval detects sooner but generates more load; raising the failure threshold avoids evictions caused by a transient spike.
- The backend service over the MIG
The backend service is the central piece: it defines how traffic is spread, which check is used, how long to wait for a response, whether there is CDN and whether there is Cloud Armor.
# 1) Create the backend service (EXTERNAL_MANAGED scheme = modern global load balancer)
gcloud compute backend-services create bs-catalogo-web \
--global \
--load-balancing-scheme=EXTERNAL_MANAGED \
--protocol=HTTP \
--port-name=http-catalogo \
--health-checks=hc-catalogo \
--timeout=30s \
--connection-draining-timeout=60s \
--description="Backend for the AlpinaShop web catalogue"
# 2) Add the regional MIG as a backend
gcloud compute backend-services add-backend bs-catalogo-web \
--global \
--instance-group=alpinashop-web-mig \
--instance-group-region=europe-west1 \
--balancing-mode=RATE \
--max-rate-per-instance=80 \
--capacity-scaler=1.0Important options:
--load-balancing-scheme=EXTERNAL_MANAGED: this is the scheme of the current global external Application Load Balancer, based on Google's Envoy proxy. The oldEXTERNALvalue corresponds to the classic load balancer. Always useEXTERNAL_MANAGEDfor new work: it supports advanced routing, retries, header rewriting and traffic mirroring.--timeout=30s: how long the load balancer waits for the complete response from the backend. If the application has any slow operation (a catalogue export, for instance), raise it or move that operation to an asynchronous process. A 502 withbackend_timeoutin the logs is exactly this.--connection-draining-timeout=60s: when an instance leaves the group (through autoscaling or an update), the load balancer stops sending it new requests but gives it 60 seconds to finish the ones in flight. Without draining, every autoscaling scale-down cuts live connections.--balancing-mode=RATEwith--max-rate-per-instance=80: each instance is considered "full" at 80 requests per second. We develop this in section 10.
- The backend bucket for the images
The 60 GB of catalogue images must not go through the Flask application. A backend bucket connects the load balancer directly to alpinashop-catalogo: Google serves the object with no instance involved at all.
gcloud compute backend-buckets create bb-catalogo-imagenes \
--gcs-bucket-name=alpinashop-catalogo \
--description="Product images served directly from Cloud Storage"Advantages over serving them from the application:
| Aspect | From Flask | From a backend bucket |
|---|---|---|
| VM CPU consumption | High (each image occupies a worker) | Zero |
| Scaling | Limited by the MIG | Unlimited, it is Cloud Storage |
| CDN caching | Requires configuring headers in the app | Enabled with a single option (03-03) |
| Cost | Bigger instances | Storage and egress only |
| Latency | One extra hop | Direct |
For the bucket to be servable by the load balancer, its objects must be publicly readable or else accessible through signed cookies (covered in 03-03). AlpinaShop's product images are public by nature — they are in the shop — so:
# Public read ONLY for the web images and thumbnails prefix.
# The high-resolution originals must NOT be public.
gcloud storage buckets add-iam-policy-binding gs://alpinashop-catalogo \
--member=allUsers \
--role=roles/storage.objectViewer \
--condition='expression=resource.name.startsWith("projects/_/buckets/alpinashop-catalogo/objects/productos/") && (resource.name.contains("/web/") || resource.name.contains("/thumb/")),title=solo-web-y-thumb,description=Public only for the content optimised for the shop'Notice the IAM condition: it makes public only the web/ and thumb/ part of the path convention we settled in 02-02, leaving original/ private. IAM conditions are explained in 03-04; here they serve to avoid exposing the high-resolution originals by accident.
- The URL map: routing by path
The URL map is the load balancer's brain: it looks at the host and path of each request and decides which backend serves it.
# 1) Create the map with the default service (the application)
gcloud compute url-maps create alpinashop-url-map \
--default-service=bs-catalogo-web \
--description="Routing for the AlpinaShop site"
# 2) Add the path matcher that diverts /imagenes/* to the bucket
gcloud compute url-maps add-path-matcher alpinashop-url-map \
--path-matcher-name=matcher-principal \
--default-service=bs-catalogo-web \
--backend-bucket-path-rules="/imagenes/*=bb-catalogo-imagenes,/estatico/*=bb-catalogo-imagenes" \
--new-hosts="alpinashop.example,www.alpinashop.example"The result can be viewed and edited as YAML, which is the clearest way to reason about it:
name: alpinashop-url-map
defaultService: .../backendServices/bs-catalogo-web
hostRules:
- hosts:
- alpinashop.example
- www.alpinashop.example
pathMatcher: matcher-principal
pathMatchers:
- name: matcher-principal
defaultService: .../backendServices/bs-catalogo-web
pathRules:
- paths:
- /imagenes/*
service: .../backendBuckets/bb-catalogo-imagenes
- paths:
- /estatico/*
service: .../backendBuckets/bb-catalogo-imagenesHow a request is evaluated, with real AlpinaShop examples:
| Request | Rule that applies | Backend |
|---|---|---|
https://alpinashop.example/ |
The matcher's defaultService |
bs-catalogo-web |
https://alpinashop.example/producto/moc-40 |
The matcher's defaultService |
bs-catalogo-web |
https://alpinashop.example/imagenes/productos/moc-40/web/frontal-800.webp |
/imagenes/* |
bb-catalogo-imagenes |
https://alpinashop.example/estatico/css/tienda.css |
/estatico/* |
bb-catalogo-imagenes |
https://otrodominio.example/ |
No hostRule matches |
The map's defaultService |
Three details that save hours of debugging:
- Path matching is by prefix with a trailing
*, not a regular expression./imagenes/*matches everything starting with/imagenes/. - The path is not rewritten by default. A request to
/imagenes/productos/moc-40/web/frontal-800.webplooks in the bucket for the objectimagenes/productos/moc-40/web/frontal-800.webp, with theimagenes/prefix included. Since our convention from 02-02 stores objects underproductos/<sku>/web/..., without that prefix, we have to rewrite:
# Export, edit and re-import the map to add urlRewrite
gcloud compute url-maps export alpinashop-url-map \
--destination=/tmp/url-map.yaml --global# Fragment to add inside the pathRule for /imagenes/*
pathRules:
- paths:
- /imagenes/*
service: .../backendBuckets/bb-catalogo-imagenes
routeAction:
urlRewrite:
pathPrefixRewrite: /With pathPrefixRewrite: /, the request /imagenes/productos/moc-40/web/frontal-800.webp becomes /productos/moc-40/web/frontal-800.webp before reaching the bucket, which is exactly the object's path.
gcloud compute url-maps validatechecks the map before applying it, including routing tests declared in the YAML itself. It is highly recommended in a pipeline (06-01).
- Target proxy and forwarding rule
The last two pieces. Here we first assemble the HTTP version to validate that everything works; the certificate and the HTTPS proxy belong to 03-07, but we already write down the final form.
# --- HTTP version (to validate the chain before having a certificate) ---
gcloud compute target-http-proxies create alpinashop-http-proxy \
--url-map=alpinashop-url-map
gcloud compute forwarding-rules create alpinashop-fr-http \
--global \
--address=alpinashop-lb-ip \
--target-http-proxy=alpinashop-http-proxy \
--ports=80 \
--load-balancing-scheme=EXTERNAL_MANAGED# --- HTTPS version (completed in 03-07, with the managed certificate) ---
# gcloud compute target-https-proxies create alpinashop-https-proxy \
# --url-map=alpinashop-url-map \
# --ssl-certificates=alpinashop-cert \
# --ssl-policy=alpinashop-ssl-policy
#
# gcloud compute forwarding-rules create alpinashop-fr-https \
# --global --address=alpinashop-lb-ip \
# --target-https-proxy=alpinashop-https-proxy \
# --ports=443 --load-balancing-scheme=EXTERNAL_MANAGEDEnd-to-end check:
IP=$(gcloud compute addresses describe alpinashop-lb-ip --global --format="value(address)")
echo "Load balancer IP: $IP"
# The application
curl -s -o /dev/null -w "%{http_code} %{time_total}s\n" "http://$IP/"
# An image from the bucket
curl -s -o /dev/null -w "%{http_code}\n" "http://$IP/imagenes/productos/moc-40/web/frontal-800.webp"
# Backend health
gcloud compute backend-services get-health bs-catalogo-web --globalBe patient: a freshly created global load balancer takes between 5 and 10 minutes to propagate to every point of presence. During that time it is normal to get 404s or 502s. If it is still failing after fifteen minutes, then there really is a problem.
- Health checks in depth: why they fail
This section is the most profitable one in the lesson. Most broken load balancers are not badly built: they have their backends in UNHEALTHY.
First point, and the one that is most often the cause: health checks do not originate on the internet or in your subnet. Google launches them from two fixed ranges:
130.211.0.0/2235.191.0.0/16
If the VPC firewall does not allow them, the check never arrives, the instance is marked unhealthy and the load balancer returns 502 Server Error without saying why. We already created the rule in 03-01 (fw-allow-lb-health-web), but it is worth knowing how to verify it:
gcloud compute firewall-rules list \
--filter="network:alpinashop-vpc AND sourceRanges:(130.211.0.0/22 OR 35.191.0.0/16)" \
--format="table(name, allowed[].map().firewall_rule().list(), targetTags.list(), disabled)"A complete checklist for when a backend is UNHEALTHY:
| No. | Cause | How to confirm it | Fix |
|---|---|---|---|
| 1 | The firewall does not allow 130.211.0.0/22 and 35.191.0.0/16 |
The previous command returns nothing | Create the rule |
| 2 | The rule's network tag is not on the instances | gcloud compute instances list --format="table(name,tags.items.list())" |
Add the tag to the MIG's template |
| 3 | Named port badly defined | gcloud compute instance-groups managed describe alpinashop-web-mig --region=europe-west1 --format="value(namedPorts)" |
set-named-ports |
| 4 | The /salud path does not exist or returns 301/302 |
curl -i http://INTERNAL_IP:8080/salud from another VM |
Fix the application; the check requires 200 and does not follow redirects |
| 5 | The application listens on 127.0.0.1 instead of 0.0.0.0 |
ss -tlnp inside the VM |
gunicorn --bind 0.0.0.0:8080 |
| 6 | /salud queries the database and the database is slow |
/salud latency greater than the timeout |
Make /salud shallow |
| 7 | The instance is still booting | Just created by autoscaling | Adjust the MIG's --initial-delay |
| 8 | Wrong protocol (HTTPS in the check, HTTP in the app) | describe the health check |
Use http |
Diagnosis from inside a VM, with no public IP, thanks to IAP:
gcloud compute ssh alpinashop-web-1 --zone=europe-west1-b --tunnel-through-iap
# Inside the VM
curl -i http://localhost:8080/salud # does the app respond?
ss -tlnp | grep 8080 # which interface is it listening on?
sudo journalctl -u gunicorn -n 50 --no-pagerAnd the aggregated view, which tells you which specific instance is broken:
gcloud compute backend-services get-health bs-catalogo-web --global \
--format="table(status.healthStatus[].instance, status.healthStatus[].healthState)"There are two kinds of check on the platform and it is important not to confuse them:
| Type | What it is for | Effect of a failure |
|---|---|---|
| Load balancer check | Deciding whether an instance receives traffic | It leaves rotation |
| MIG autohealing check | Deciding whether an instance is recreated | It is destroyed and another one created |
If you use the same aggressive check for both things, a transient application failure can make the MIG destroy the whole fleet one machine after another. The recommended practice is for the autohealing check to be more lenient (a higher failure threshold and a simpler path) than the load balancer's.
- Balancing policies, capacity and session affinity
Balancing mode
Each backend declares when it is considered full. There are three modes:
| Mode | Metric | When to use it |
|---|---|---|
RATE |
Requests per second | Web applications with homogeneous requests. AlpinaShop's |
UTILIZATION |
The group's CPU usage | Workloads with very variable CPU cost per request |
CONNECTION |
Simultaneous connections | Layer 4 load balancers and long-lived connections |
For AlpinaShop, RATE with --max-rate-per-instance=80 means: "each instance is deemed full at 80 req/s". That number is not made up; it is measured. The correct way to obtain it:
- Run a load test against a single instance.
- Raise the rate until the 95th percentile latency degrades or errors appear.
- Take between 60 % and 70 % of that figure as
max-rate-per-instance, leaving headroom for spikes.
If the number is too high, the load balancer saturates the instances before overflowing to other regions. If it is too low, autoscaling adds unnecessary machines and pushes the cost up.
capacity-scaler: the most useful operational lever
--capacity-scaler is a multiplier between 0.0 and 1.0 applied to the declared capacity. It serves to drain a backend without deleting it:
# Drain the backend to half (during a deployment, for example)
gcloud compute backend-services update-backend bs-catalogo-web \
--global \
--instance-group=alpinashop-web-mig \
--instance-group-region=europe-west1 \
--capacity-scaler=0.5
# Drain it completely (it stops receiving new traffic, draining is respected)
gcloud compute backend-services update-backend bs-catalogo-web \
--global --instance-group=alpinashop-web-mig \
--instance-group-region=europe-west1 \
--capacity-scaler=0.0It is the clean way of doing a blue/green deployment or of taking a troubled region out of service.
Session affinity
By default there is no affinity: each request can go to a different instance. That is the right thing for a stateless application, which is how it should be designed.
| Affinity type | Based on | Cost |
|---|---|---|
NONE (default) |
Nothing | Optimal distribution |
CLIENT_IP |
Client IP | Poor distribution behind corporate NAT |
GENERATED_COOKIE |
A cookie generated by the load balancer | Acceptable distribution |
HEADER_FIELD |
A specific header | Requires control of the client |
AlpinaShop does not need affinity, and it is important to understand why: the shopping cart and the session live in Firestore (decided in 02-06), not in the instance's memory. Affinity is a patch for applications with local state, and it always degrades distribution: one instance can end up overloaded while others sit idle, and autoscaling cannot relieve it because the clients are already stuck to it.
Other backend service options
# Retry policy and outlier detection (EXTERNAL_MANAGED only)
gcloud compute backend-services update bs-catalogo-web --global \
--custom-request-header="X-Cliente-Region:{client_region}" \
--custom-response-header="X-Cache-Estado:{cdn_cache_status}"Custom headers with variables ({client_region}, {client_city}, {cdn_cache_status}) are very useful: they let the application know which country the customer is coming from without geolocating anything, and let you debug the cache by looking at a response header (we will use this in 03-03).
- Redirecting HTTP to HTTPS
Serving the shop over unencrypted HTTP is not acceptable. The redirect is done on the load balancer itself, not in the application, so that not a single byte of content travels in the clear.
It is implemented with a dedicated URL map that has no backends, just a redirect action:
# URL map that only redirects
gcloud compute url-maps import alpinashop-redir-map --global --quiet --source=- <<'EOF'
name: alpinashop-redir-map
defaultUrlRedirect:
httpsRedirect: true
redirectResponseCode: MOVED_PERMANENTLY_DEFAULT
stripQuery: false
EOF
# The HTTP proxy points at that map instead of the real one
gcloud compute target-http-proxies update alpinashop-http-proxy \
--url-map=alpinashop-redir-mapThe options explained:
httpsRedirect: true: changes the scheme tohttpswhile keeping host and path.redirectResponseCode: MOVED_PERMANENTLY_DEFAULT: this is a 301. Browsers cache it, so subsequent visits do not even make the cleartext request. If you are testing and do not want it to stick in the browser, useFOUND(302) in the meantime.stripQuery: false: keeps the query string. Setting it totruewould break campaign links withutm_*parameters.
The HSTS header, which makes the browser not even attempt HTTP, is emitted by the application and is explained in 03-07.
- Global scaling: why there is no need to "pre-warm"
Anyone coming from other providers is used to "pre-warming" the load balancer before a campaign: telling the provider that traffic is coming so it can expand the load balancer's capacity. In Google Cloud that does not exist and is not needed.
The reason is architectural:
- The global load balancer is not a set of instances belonging to you. It is a function of Google's network infrastructure, the same one that serves Search, YouTube and Gmail. Its capacity is not sized for you.
- The anycast IP is announced from every PoP simultaneously. A traffic spike does not arrive at one point; it is spread geographically by construction.
- The load balancer also absorbs volumetric layer 3 and 4 attacks without you configuring anything, because the mitigation is in the network itself.
What you do have to size is what sits behind it: the alpinashop-web-mig MIG is configured for 2–10 instances, and if the autumn campaign needs more, the limit is raised in the autoscaler. The bottleneck will never be the load balancer; it will be your application or your database.
The practical consequence for the AlpinaShop team: before the campaign, the checklist does not include "warn the load balancer" but raising the MIG's max-num-replicas, reviewing max-rate-per-instance with real data and checking that the Cloud SQL read replica can cope.
- Internal load balancers for internal services
Not all traffic is public. When AlpinaShop has internal services — the recommendation engine, the warehouse stock API, Lucía's report generator — those services must not have a public IP. The right piece is an internal load balancer:
# Internal Application Load Balancer (layer 7) for an internal service
gcloud compute health-checks create http hc-stock-interno \
--port=8080 --request-path=/salud --region=europe-west1
gcloud compute backend-services create bs-stock-interno \
--region=europe-west1 \
--load-balancing-scheme=INTERNAL_MANAGED \
--protocol=HTTP \
--health-checks=hc-stock-interno \
--health-checks-region=europe-west1
# Proxy-only subnet: mandatory for layer 7 internal load balancers.
# It is a reserved range from which the Google-managed proxies take IPs.
gcloud compute networks subnets create sn-proxy-euw1 \
--network=alpinashop-vpc \
--region=europe-west1 \
--range=10.10.10.0/24 \
--purpose=REGIONAL_MANAGED_PROXY \
--role=ACTIVETwo new things and one warning:
- The scheme is
INTERNAL_MANAGED, notEXTERNAL_MANAGED. - Layer 7 internal load balancers need a proxy-only subnet (
--purpose=REGIONAL_MANAGED_PROXY) per region. It is a range that hosts none of your VMs: it is used by Google's proxies. Without it, creation fails with a message that is not always obvious. Room must be reserved for it in the 03-01 addressing plan: here we have used10.10.10.0/24, inside the free block we set aside. - Firewall rules: the traffic arrives from the proxy-only subnet's range, not from the original client. You have to allow
10.10.10.0/24towards the backends.
When to use each one:
| AlpinaShop service | Load balancer |
|---|---|
| Public catalogue | Global external Application (the one we have built) |
| Stock API consumed only by the shop | Internal Application |
| Session database with a proprietary protocol | Internal passthrough Network |
- Serverless NEGs: Cloud Run behind the same load balancer
In 02-07 we decided (DA-001) that the catalogue will end up on Cloud Run. Cloud Run already gives you an HTTPS URL of its own, so why put a load balancer in front of it? For four reasons you already know:
- Your own domain and a managed certificate, unified (03-07).
- Cloud CDN over the application's content (03-03).
- Cloud Armor: Cloud Run's native URL cannot be protected with a WAF; the load balancer's can (03-05).
- A single URL map distributing between Cloud Run, the bucket and whatever comes next.
The piece that makes it possible is the serverless NEG (serverless network endpoint group): a backend that points not at machines but at a managed service.
# Serverless NEG pointing at the future Cloud Run service alpinashop-web
gcloud compute network-endpoint-groups create neg-alpinashop-web \
--region=europe-west1 \
--network-endpoint-type=serverless \
--cloud-run-service=alpinashop-web
# Backend service that uses it (no health check: Cloud Run manages that)
gcloud compute backend-services create bs-alpinashop-run \
--global \
--load-balancing-scheme=EXTERNAL_MANAGED
gcloud compute backend-services add-backend bs-alpinashop-run \
--global \
--network-endpoint-group=neg-alpinashop-web \
--network-endpoint-group-region=europe-west1And then migrating from the MIG to Cloud Run is a change in the URL map, not a rebuild:
# Send only /nuevo/* to Cloud Run to test in production with real traffic
gcloud compute url-maps add-path-matcher alpinashop-url-map \
--path-matcher-name=matcher-migracion \
--default-service=bs-catalogo-web \
--path-rules="/nuevo/*=bs-alpinashop-run" \
--new-hosts="alpinashop.example"This is what stops today's work being thrown away tomorrow: the load balancer outlives the change of backend. The IP, the domain, the certificate, the Cloud Armor rules and the CDN configuration all stay; only where the map points changes. The details of Cloud Run belong to 07-02.
There are several types of NEG, and it is worth knowing where they fit:
| NEG type | Points at | Use |
|---|---|---|
Zonal (GCE_VM_IP_PORT) |
VM or pod IP:port | GKE with native Ingress/Gateway |
| Serverless | Cloud Run, Cloud Functions, App Engine | The case above |
Internet (INTERNET_FQDN_PORT) |
An external service by name | Putting CDN or Armor in front of a third-party origin |
| Hybrid | An on-premises server | Migrations (07-03) |
- Load balancer logging and metrics
Without observability, a load balancer is a black box. Enable it from the start:
gcloud compute backend-services update bs-catalogo-web \
--global \
--enable-logging \
--logging-sample-rate=1.0--logging-sample-rate=1.0 logs 100 % of requests. On a high-traffic site it is lowered to 0.1 to contain the log cost; at AlpinaShop, with peaks of 20 req/s, 100 % is perfectly affordable and makes diagnosis far easier.
Useful queries:
# Requests that ended in 5xx in the last hour
gcloud logging read '
resource.type="http_load_balancer"
AND httpRequest.status>=500
' --project=alpinashop-prod --limit=20 --freshness=1h \
--format="table(
timestamp,
httpRequest.requestUrl,
httpRequest.status,
jsonPayload.statusDetails
)"The jsonPayload.statusDetails field is the most valuable one in the load balancer log. It translates the generic 502 into a concrete cause:
statusDetails |
Meaning | What to look at |
|---|---|---|
failed_to_pick_backend |
No healthy backend | Health checks (section 9) |
backend_timeout |
The backend did not respond in time | --timeout and the app's performance |
backend_connection_closed_before_data_sent_to_client |
The backend cut the connection | gunicorn's keepalive lower than the load balancer's |
response_sent_by_backend |
Everything normal | The code was returned by the application |
denied_by_security_policy |
Blocked by Cloud Armor | 03-05 |
client_disconnected_before_any_response |
The client went away | Normal on mobile traffic |
The backend_connection_closed_before_data_sent_to_client case deserves a note because it is subtle and very common: the global load balancer keeps persistent connections with the backend for 10 minutes. If gunicorn closes the connection sooner (its default keepalive is 2 seconds), the load balancer finds the connection closed mid-way and returns sporadic 502s, impossible to reproduce by hand. The fix is to configure gunicorn with a keepalive higher than the load balancer's:
Recommended metrics for Marta's dashboard (built in 06-04):
| Metric | What for |
|---|---|
loadbalancing.googleapis.com/https/request_count |
Volume, by response code |
loadbalancing.googleapis.com/https/total_latencies |
End-to-end latency (p50, p95, p99) |
loadbalancing.googleapis.com/https/backend_latencies |
Backend-only latency: separates the problem |
loadbalancing.googleapis.com/https/backend_request_count |
Requests that reached the backend (versus those served by the CDN) |
The difference between total_latencies and backend_latencies is what tells you whether the problem is yours or the network's: if the total rises and the backend's does not, the problem is on the way to the client.
Common Mistakes and Tips
- Forgetting the firewall for
130.211.0.0/22and35.191.0.0/16. It is the number one cause of a freshly created load balancer returning 502. Before debugging anything else, check that rule. - Not defining the named port on the MIG. The backend service uses
--port-name, and if the group does not declare it, the traffic goes to port 80 and nothing answers. - Getting impatient. A global load balancer takes between 5 and 10 minutes to propagate. The first 404s and 502s are normal.
- Letting
/saludquery the database. It turns a partial failure into a total outage: all the backends are marked unhealthy at once. - Using the same check for the load balancer and for MIG autohealing. A transient failure can make the group destroy and recreate the whole fleet.
- Confusing
EXTERNALwithEXTERNAL_MANAGED. For new work useEXTERNAL_MANAGED; mixing them between the backend service and the forwarding rule produces confusing validation errors. - Forgetting the backend bucket's
urlRewrite. If the map sends/imagenes/*to the bucket without rewriting, it looks for an object whose name starts withimagenes/and you get a systematic 404. - Enabling session affinity "just in case". It degrades distribution and hides a design problem. If the application needs affinity, what has to be fixed is the state, not the load balancer.
- Setting
max-rate-per-instanceby eye. Measure it with a load test against one instance and apply a 30-40 % margin. - Not enabling load balancer logging. Without
statusDetailsyou are debugging blind. - gunicorn's
keepalivetoo low. It causes sporadic, irreproducible 502s. Set it above the load balancer's 600 seconds. - Not deleting the resources from a test. The chain has seven objects; they are deleted in the reverse order of creation. Orphaned static IPs keep on being billed.
Exercises
Exercise 1 — Choosing the right load balancer
For each AlpinaShop scenario, say which load balancer it calls for and justify it in one line:
- The public shop
alpinashop.example, with customers across Europe, HTTPS, image caching and a WAF. - An internal stock API consumed only by the catalogue instances inside
alpinashop-vpc. - A warehouse synchronisation service that speaks a proprietary protocol over TCP on port 9000 and needs to see the client's real IP.
- An internal reporting dashboard for Lucía that routes
/informes/*to one service and/exportar/*to another, reachable only from the internal network.
Exercise 2 — Debugging a broken load balancer
Marta has built the load balancer following the lesson, but curl http://$IP/ returns 502 Server Error and gcloud compute backend-services get-health bs-catalogo-web --global shows every instance as UNHEALTHY. It is known that:
- From another VM in
sn-web-euw1,curl http://10.10.0.7:8080/saludreturns200 OK. - The
fw-allow-lb-health-webrule exists and allowstcp:80,tcp:8080from the correct ranges towards theweb-catalogotag.
Write the sequence of diagnostic commands and the most likely cause.
Exercise 3 — Extending the URL map
AlpinaShop wants three new things:
- That
blog.alpinashop.examplego to a different backend service calledbs-blog. - That
/api/*have a 120-second timeout because there is a slow export, without changing it for the rest of the site. - That
/descargas/manual-mochilas.pdfredirect permanently to/imagenes/documentos/manual-mochilas-2026.pdf.
State which resources have to be created or modified for each point and write the URL map YAML fragment for point 3.
Solutions
Solution 1
- Global external Application Load Balancer. It is worldwide public HTTP traffic, and it is the only one that supports Cloud CDN and Cloud Armor with a global anycast IP.
- Internal Application Load Balancer (
INTERNAL_MANAGED), regional ineurope-west1. It must not have a public IP and path-based routing may be useful later on. It requires a proxy-only subnet. - External passthrough Network Load Balancer (layer 4, regional). It is the only one that does not terminate the connection and therefore preserves the real source IP; besides, the protocol is not HTTP, so a layer 7 load balancer is no use.
- Internal Application Load Balancer. It needs path-based routing (layer 7) and must not be reachable from the internet. If access from outside the office were also needed, the correct alternative would be the global external load balancer protected with IAP (03-04), not opening it up.
Solution 2
# 1) Are the instances tagged as web-catalogo?
gcloud compute instances list \
--filter="name~alpinashop-web-mig" \
--format="table(name, zone, status, tags.items.list())"
# 2) Does the MIG have the named port defined?
gcloud compute instance-groups managed describe alpinashop-web-mig \
--region=europe-west1 \
--format="value(namedPorts)"
# 3) Which port and path does the check point at?
gcloud compute health-checks describe hc-catalogo \
--format="yaml(type, httpHealthCheck)"
# 4) Which port-name does the backend service use?
gcloud compute backend-services describe bs-catalogo-web --global \
--format="value(portName, protocol, timeoutSec)"Most likely cause: the named port is missing. The firewall is fine and the application answers on 8080 from inside the subnet, so the check packet can indeed get through. What is happening is that the backend service asks for the port called http-catalogo and the instance group does not have it defined, so the load balancer directs the check and the traffic to the default port (80), where nothing is listening.
gcloud compute instance-groups managed set-named-ports alpinashop-web-mig \
--region=europe-west1 \
--named-ports=http-catalogo:8080The second candidate cause would be that the MIG's instance template does not apply the web-catalogo tag to new machines: the rule would exist but would apply to nobody. Command 1 confirms it.
Solution 3
Point 1 — blog.alpinashop.example. You create the bs-blog backend service and add a new hostRule to the map with its own pathMatcher. You also have to include the name in the managed certificate (03-07) and create the corresponding DNS record.
gcloud compute url-maps add-path-matcher alpinashop-url-map \
--path-matcher-name=matcher-blog \
--default-service=bs-blog \
--new-hosts=blog.alpinashop.examplePoint 2 — timeout for /api/*. The timeout is a property of the backend service, not of the path rule. So you have to create a second backend service (for example bs-catalogo-api) pointing at the same MIG but with --timeout=120s, and route /api/* to it.
gcloud compute backend-services create bs-catalogo-api \
--global --load-balancing-scheme=EXTERNAL_MANAGED \
--protocol=HTTP --port-name=http-catalogo \
--health-checks=hc-catalogo --timeout=120s
gcloud compute backend-services add-backend bs-catalogo-api \
--global --instance-group=alpinashop-web-mig \
--instance-group-region=europe-west1 \
--balancing-mode=RATE --max-rate-per-instance=80Point 3 — permanent redirect. It is done in the URL map itself, without touching the application:
pathMatchers:
- name: matcher-principal
defaultService: .../backendServices/bs-catalogo-web
pathRules:
- paths:
- /descargas/manual-mochilas.pdf
urlRedirect:
pathRedirect: /imagenes/documentos/manual-mochilas-2026.pdf
redirectResponseCode: MOVED_PERMANENTLY_DEFAULT
stripQuery: trueNote: a pathRule with urlRedirect carries no service. And mind the order: more specific paths must be able to beat generic ones; the load balancer applies the longest match, so /descargas/manual-mochilas.pdf beats /descargas/*.
Conclusion
AlpinaShop no longer depends on a single machine. You have understood Google Cloud's load balancer catalogue by reducing it to three questions — layer 7 or layer 4, external or internal, global or regional — and you know that for public web traffic the answer is almost always the global external Application Load Balancer, with a clear distinction between proxy and passthrough.
You have built the complete chain from the inside out: the hc-catalogo check against /salud on 8080, the http-catalogo named port on the MIG, the bs-catalogo-web backend service with connection draining and RATE mode at 80 req/s per instance, the bb-catalogo-imagenes backend bucket that serves the images from alpinashop-catalogo without touching the application, the alpinashop-url-map map that sends /imagenes/* and /estatico/* to the bucket with pathPrefixRewrite and everything else to the MIG, and the forwarding rule on the global IP alpinashop-lb-ip that we reserved in 03-01.
You have learned to diagnose the one thing that really breaks: health checks, with their list of eight causes headed by the 130.211.0.0/22 and 35.191.0.0/16 ranges, and the distinction between the load balancer's check and the MIG's autohealing check. You know how to tune capacity with a measured max-rate-per-instance, how to drain a backend with capacity-scaler, and why AlpinaShop needs no session affinity: the state lives in Firestore, not in the instances. You have redirected HTTP to HTTPS at the edge, you have seen why nothing needs pre-warming ahead of the autumn campaign, where internal load balancers and their proxy-only subnet fit in, and how a serverless NEG will allow the catalogue to move to Cloud Run (DA-001) by changing only the URL map, keeping the IP, domain, certificate, cache and WAF. And you have enabled logging, which turns an opaque 502 into an actionable statusDetails.
One obvious inefficiency remains. Every time a customer in Lisbon opens a product page, the images travel from europe-west1 to Portugal, even though they are the same bytes sent a minute ago to another customer. With a 60 GB catalogue and customers spread across Europe, that is unnecessary latency and an egress bill that grows with every visit. In the next lesson, 03-03, Cloud CDN, we enable caching on the very backend service we have just built: the images will start being served from the point of presence nearest the customer, you will learn to control what gets cached and for how long, to stop campaign utm_* parameters ruining the hit ratio, and to measure the real saving in latency and in euros.
Google Cloud Platform (GCP) Course
Module 1: Introduction to Google Cloud Platform
- What is Google Cloud Platform?
- Setting Up Your GCP Account
- A Tour of the GCP Console
- Projects, Resource Hierarchy and Billing
- Regions, Zones and the Shared Responsibility Model
- Cloud Shell and the gcloud CLI
Module 2: Core GCP Services
- Compute Engine: Virtual Machines on Google Cloud
- Cloud Storage: Object Storage
- Cloud SQL: Managed Relational Databases
- App Engine: Platform as a Service
- Google Kubernetes Engine (GKE)
- NoSQL Databases: Firestore, Bigtable and Spanner
- How to Choose the Right Compute Service
Module 3: Networking and Security
- VPC Networks
- Cloud Load Balancing
- Cloud CDN
- Identity and Access Management (IAM)
- Cloud Armor
- Secrets and Encryption: Secret Manager and Cloud KMS
- Cloud DNS, TLS Certificates and Publishing Services Securely
Module 4: Data and Analytics
- BigQuery: The Analytical Data Warehouse
- Cloud Dataflow: Batch and Streaming Data Processing
- Cloud Dataproc: Managed Spark and Hadoop
- Cloud Pub/Sub: Asynchronous Messaging
- Cloud Data Fusion: Code-Free Data Integration
- Orchestrating Pipelines with Cloud Composer and Workflows
- Data Governance and Dashboards with Dataplex and Looker Studio
Module 5: Machine Learning and AI
- Vertex AI: The Machine Learning Platform on GCP
- AutoML: Custom Models Without Writing Code
- TensorFlow on GCP: Training and Serving Models
- Natural Language API
- Vision API
- Generative AI on Vertex AI: Gemini Models and Embeddings
- MLOps: From Model to Product with Vertex AI Pipelines
Module 6: DevOps and Monitoring
- Cloud Build: Continuous Integration on GCP
- Cloud Source Repositories and Source Code Management
- Cloud Functions: Serverless Functions
- Cloud Monitoring (formerly Stackdriver): Metrics, Dashboards and Alerts
- Cloud Deployment Manager and Native Infrastructure as Code
- Cloud Logging and Cloud Trace: Logs, Traces and Diagnostics
- Terraform on GCP: Infrastructure as Code in Practice
Module 7: Advanced GCP Topics
- Hybrid and Multicloud with Anthos
- Serverless Computing with Cloud Run
- Advanced Networking: Shared VPC, Peering and Hybrid Connectivity
- Security Best Practices
- Cost Management and Optimization
- Reliability: SLOs, High Availability and Disaster Recovery
- Governance at Scale: Organization, Policies and Auditing
