We closed module 1 with a promise: stop planning and start building. This lesson keeps it. Compute Engine is Google Cloud's virtual machine service and, even though more modern options exist today — containers, PaaS, serverless — it remains the natural way in for a company like AlpinaShop, which currently runs its Flask catalogue on a rented physical server and needs to move it to the cloud without rewriting it.
In this lesson you will learn what a managed VM really is and when it is still the right choice; how machine families and machine types are organised, and how to avoid overpaying by choosing badly; which images you can use and how to build your own; the disk types and their performance; how to create the alpinashop-web-1 instance in europe-west1-b from the console and from gcloud, with a startup script that installs the application by itself; how to log in over SSH securely with OS Login; and, above all, how instance templates and managed instance groups solve AlpinaShop's real problem: the traffic peaks of the autumn campaigns.
Contents
- What a managed virtual machine is and when to use one
- Machine families and machine types
- Public and custom images
- Disks: types, performance and snapshots
- Creating
alpinashop-web-1from the console - Creating the same VM with
gcloudand a startup script - SSH access and OS Login
- Service accounts and access scopes
- Instance templates and managed groups (MIG)
- Autoscaling and auto-healing: the autumn peaks
- Spot VMs: cheap, interruptible compute
- The lifecycle of an instance and what is charged in each state
- What a managed virtual machine is and when to use one
A Compute Engine virtual machine is a complete computer running on Google's infrastructure: it has its CPU, its memory, its disks, its operating system and its IP address. From the inside, it is indistinguishable from a physical server. The difference is in everything around it: it is created in 30 seconds with one command, resized by stopping and starting it, replicated a hundred times from a template, and gone when you delete it.
Managed means Google takes care of the hardware, the virtualisation, the physical network, the cooling and the replacement of failed disks. You take care of the operating system upwards: patches, configuration, application, data. It is exactly the IaaS split we saw in the shared responsibility table in lesson 01-05.
When a VM still makes sense in 2026, with so many serverless alternatives available:
- Lift-and-shift migration. You have an application that works and you want to move it without rewriting it. That is AlpinaShop's case.
- You need control of the operating system. A specific kernel, modules, drivers, software licensed against hardware, corporate security agents.
- Long-running or always-on processes. Serverless penalises (or outright forbids) processes that run for hours.
- Software that does not containerise well. Old monolithic applications, systems with state on local disk, licence servers.
- Workloads with special hardware. GPUs for inference, local disks with minimal latency, high-memory machines for in-RAM databases.
- Predictable cost at high, constant usage. A 24×7 VM with a committed use discount can work out cheaper than its serverless equivalent.
When it is NOT the right choice: an HTTP API with intermittent traffic (Cloud Run is better, 07-02), a job that reacts to an event and lasts seconds (Cloud Functions, 06-03) or an already containerised web application that you want to scale horizontally (GKE, 02-05). In lesson 02-07 we will put all of this into a decision table.
- Machine families and machine types
The machine type defines how many virtual CPUs (vCPUs) and how much memory the instance has. Google groups them into families, each optimised for a different workload profile. Choosing well here is the single decision with the biggest impact on your compute bill.
| Family | Profile | Typical vCPU/memory | What it is for | Example type |
|---|---|---|---|---|
| E2 | Economical general purpose | 2–32 vCPU, flexible ratios | Small web servers, development environments, microservices. The default option to start with | e2-medium (2 vCPU, 4 GB) |
| N2 / N2D | Balanced general purpose | 2–128 vCPU | General production workloads. N2D uses AMD EPYC and tends to be somewhat cheaper | n2-standard-4 (4 vCPU, 16 GB) |
| N4 / C4 | Recent generations | Wide range | Modern replacements for N2/C2 with a better price-performance ratio in the regions where they are available | n4-standard-4 |
| C3 / C3D | Compute-optimised | 4–176 vCPU | CPU intensive: rendering, compilation, simulation, game servers, image processing | c3-highcpu-8 |
| M3 / M2 / M1 | Memory-optimised | Up to several TB of RAM | SAP HANA, in-memory databases, analysis of large datasets in RAM | m3-ultramem-32 |
| A3 / G2 | Accelerators | With GPU (NVIDIA) | Model training and inference, video transcoding | a3-highgpu-8g |
| Custom | Made to measure | You choose vCPU and RAM | When no standard type fits and you are paying for memory or CPU you do not use | custom-6-16384 |
Within each family there are suffixes that indicate the memory/CPU ratio:
-highcpu: little memory per vCPU (≈1 GB). For compute-heavy processes.-standard: balanced ratio (≈4 GB per vCPU). The general case.-highmem: a lot of memory per vCPU (≈8 GB). Caches, databases.-ultramem/-megamem: memory extremes, only in the M families.
Custom types. If your application needs 6 vCPU and 16 GB, no standard type gives you exactly that: n2-standard-8 would give you 8 vCPU and 32 GB, and you would pay twice what you need. The syntax of a custom type is <family>-custom-<vCPU>-<MiB of memory>:
# 6 vCPU and 16 GB (16384 MiB) on the N2 family
gcloud compute instances create ejemplo-custom \
--machine-type=n2-custom-6-16384 \
--zone=europe-west1-bBear the restrictions in mind: the number of vCPUs must be even (except for 1), and the memory per vCPU has a minimum and a maximum per family. If you go over the standard maximum, it is billed as "extended memory", which costs more.
What does AlpinaShop choose? Its Flask catalogue serves a few tens of requests per second today, with peaks concentrated in autumn. We start with e2-medium (2 vCPU, 4 GB) in development and e2-standard-2 in production, knowing that growth will be handled horizontally (more instances) and not vertically (one bigger instance). This decision — scale out, not up — is what makes the autoscaling in section 10 possible.
About prices: all the amounts in this course are orders of magnitude to reason with, not price lists. Prices vary by region and change over time; always check the official Google Cloud calculator and the Compute Engine pricing page before committing.
- Public and custom images
An image is the template for the boot disk: the operating system and the preinstalled software. Google maintains a catalogue of public images, grouped into image families that always point at the most recent version.
# List the Debian image families available
gcloud compute images list \
--filter="family~debian" \
--format="table(name, family, status)"The most common ones:
| Image family | Project | Notes |
|---|---|---|
debian-12 |
debian-cloud |
Lightweight, stable, the default choice on GCP |
ubuntu-2404-lts |
ubuntu-os-cloud |
Very common, plenty of third-party documentation |
rocky-linux-9 |
rocky-linux-cloud |
RHEL-compatible, with no licence cost |
rhel-9 |
rhel-cloud |
Official Red Hat, with an added licence cost |
cos-stable |
cos-cloud |
Container-Optimized OS: minimal, designed only to run containers |
windows-2022 |
windows-cloud |
Windows Server, with a licence billed by the hour |
Using the family instead of the exact name (--image-family=debian-12 instead of --image=debian-12-bookworm-v20260101) is good practice: every new VM boots with the latest patched version.
Custom images. When you have a VM configured the way you want it — dependencies installed, agents, company configuration — you can turn it into an image and use it as the basis for all the others. It is the key technique for getting new instances to boot in seconds instead of minutes:
# 1. Stop the instance to guarantee disk consistency
gcloud compute instances stop alpinashop-web-1 --zone=europe-west1-b
# 2. Create the image from its boot disk
gcloud compute images create alpinashop-web-base-v1 \
--source-disk=alpinashop-web-1 \
--source-disk-zone=europe-west1-b \
--family=alpinashop-web \
--description="Debian 12 + Python + Flask + catalogue dependencies"By giving it --family=alpinashop-web, later versions (-v2, -v3) will form a family of their own and you will be able to create VMs with --image-family=alpinashop-web --image-project=alpinashop-prod, always getting the latest one.
In module 6 you will see how to automate building these images in a pipeline; for now, doing it by hand is enough to understand the mechanism.
- Disks: types, performance and snapshots
Every VM needs at least one boot disk, and it can have additional disks. The choice of disk type affects performance as much as the CPU does, or more.
| Type | Name in gcloud |
Performance | Persistence | When to use it |
|---|---|---|---|---|
| Standard persistent | pd-standard |
Low (HDD) | Survives the VM | Cold data, logs, copies. Almost obsolete |
| Balanced persistent | pd-balanced |
Medium-high (SSD) | Survives the VM | The default option. Boot and general data |
| SSD persistent | pd-ssd |
High, elevated IOPS | Survives the VM | Databases, workloads with many random writes |
| Extreme persistent | pd-extreme |
Very high, provisionable IOPS | Survives the VM | Large business-critical databases |
| Hyperdisk | hyperdisk-balanced, hyperdisk-extreme, hyperdisk-throughput |
Configurable: IOPS and throughput independent of size | Survives the VM | The current generation on modern families (C3, N4, M3); lets you tune performance without oversizing capacity |
| Local SSD | local-ssd |
Maximum, minimal latency | Lost when the VM is stopped or deleted | Cache, scratch, rebuildable indexes. Never unique data |
| Regional persistent disk | --region instead of --zone |
Somewhat lower than its zonal equivalent | Synchronously replicated across two zones | When you need the data to survive the loss of a zone |
Points that tend to surprise anyone coming from a physical server:
- On traditional persistent disks, performance scales with size. A 10 GB
pd-balancedhas very few IOPS. If your database is slow, sometimes the fix is to make the disk bigger even if you do not need the space. With Hyperdisk this breaks: performance is provisioned separately. - Performance is also limited by the machine type. A 2 vCPU VM will not reach the disk's maximum throughput even if the disk allows it.
- Persistent disks are already replicated within their zone. You do not need RAID for reliability.
- Local SSD is genuinely ephemeral. If you stop the instance, the contents disappear. It is a frequent and expensive mistake.
Snapshots and disk images. These are two different mechanisms that are often confused:
| Snapshot | Image | |
|---|---|---|
| Purpose | Point-in-time backup | Template for creating new VMs |
| Incremental | Yes (only blocks changed since the previous one) | No |
| Scope | Global, can be restored in another region | Global |
| Typical use | Disaster recovery, daily backup | Standardising how a fleet boots |
Creating a snapshot and scheduling them automatically:
# Point-in-time snapshot of the boot disk
gcloud compute snapshots create alpinashop-web-1-snap-20260805 \
--source-disk=alpinashop-web-1 \
--source-disk-zone=europe-west1-b \
--storage-location=europe-west1
# Daily snapshot policy at 03:00 UTC, with 14 days of retention
gcloud compute resource-policies create snapshot-schedule alpinashop-diario \
--region=europe-west1 \
--max-retention-days=14 \
--daily-schedule \
--start-time=03:00 \
--on-source-disk-delete=apply-retention-policy
# Attach the policy to the disk
gcloud compute disks add-resource-policies alpinashop-web-1 \
--resource-policies=alpinashop-diario \
--zone=europe-west1-bThis block is worth dwelling on. --max-retention-days=14 means snapshots older than 14 days delete themselves, which stops the cost growing indefinitely. --on-source-disk-delete=apply-retention-policy says that if someone deletes the disk, the snapshots are not deleted immediately but instead follow their retention policy: it is your safety net against an accidental deletion. Marta will have daily backups without writing a single cron entry.
- Creating
alpinashop-web-1 from the console
alpinashop-web-1 from the consoleBefore automating, it is worth going through the full form once, because it shows you which options exist. In the console: Compute Engine → VM instances → Create instance.
The fields that matter:
- Name:
alpinashop-web-1. Lower case, hyphens, no accents. It cannot be changed afterwards. - Region and zone:
europe-west1/europe-west1-b, the decision we made in 01-05. - Machine configuration: General purpose family → E2 series → type
e2-medium. - Boot disk: Debian 12,
pd-balanced, 20 GB. - Identity and API access: service account and scopes (section 8).
- Firewall: the "Allow HTTP/HTTPS traffic" checkboxes add network tags and create firewall rules. We will use them for testing, but designing networks and firewalls properly is the subject of lesson 03-01.
- Advanced options → Automation: this is where the startup script is pasted.
- Labels:
entorno=dev,equipo=plataforma,centro-coste=tienda,aplicacion=catalogo, following the scheme we settled on in 01-04.
Before pressing Create, use the "Equivalent command line" link we saw in 01-03: it gives you back the exact gcloud command. It is the best way to learn the syntax and the first step towards infrastructure as code.
- Creating the same VM with
gcloud and a startup script
gcloud and a startup scriptA startup script is a script that the VM runs as root on every boot. It is the simplest form of automatic configuration: without it, you would have to log in over SSH and type the commands by hand on every new instance, which makes autoscaling impossible.
First we write the script in a separate file (more readable and easier to version than inlining it):
cat > ~/startup-catalogo.sh <<'SCRIPT'
#!/bin/bash
set -e
# Install system dependencies
apt-get update
apt-get install -y python3-pip python3-venv
# Create the application environment
mkdir -p /opt/alpinashop
python3 -m venv /opt/alpinashop/venv
/opt/alpinashop/venv/bin/pip install flask gunicorn
# Minimal Flask application for the catalogue
cat > /opt/alpinashop/app.py <<'PYCODE'
from flask import Flask
import socket
app = Flask(__name__)
PRODUCTS = [
{"sku": "MOC-40", "nombre": "Trekking Backpack 40L", "precio": 89.90},
{"sku": "BOT-GTX", "nombre": "Alpina Gore-Tex Boots", "precio": 149.00},
{"sku": "TDA-2P", "nombre": "Ultralight 2-person Tent", "precio": 219.50},
]
@app.route("/")
def catalog():
rows = "".join(
f"<li>{p['sku']} - {p['nombre']} - {p['precio']:.2f} EUR</li>"
for p in PRODUCTS
)
return (
f"<h1>AlpinaShop</h1><ul>{rows}</ul>"
f"<p>Served by: {socket.gethostname()}</p>"
)
@app.route("/salud")
def health():
return "ok", 200
PYCODE
# systemd service so that it starts on its own and restarts if it fails
cat > /etc/systemd/system/alpinashop.service <<'UNIT'
[Unit]
Description=AlpinaShop catalogue
After=network.target
[Service]
WorkingDirectory=/opt/alpinashop
ExecStart=/opt/alpinashop/venv/bin/gunicorn -b 0.0.0.0:80 -w 2 app:app
Restart=always
[Install]
WantedBy=multi-user.target
UNIT
systemctl daemon-reload
systemctl enable --now alpinashop
SCRIPTThree details of the script are worth understanding:
set -eaborts the script as soon as a command fails. Without it, a failedapt-getwould go unnoticed and you would end up with a half-configured VM that looks healthy.<<'SCRIPT'with single quotes stops bash substituting variables while writing the file; we want the contents to reach the VM literally.- The
/saludroute exists so that, later on, the instance group can check whether the application is alive. Returning 200 on a lightweight route is the basis of any health check.
Now we create the instance:
gcloud compute instances create alpinashop-web-1 \
--project=alpinashop-dev \
--zone=europe-west1-b \
--machine-type=e2-medium \
--image-family=debian-12 \
--image-project=debian-cloud \
--boot-disk-size=20GB \
--boot-disk-type=pd-balanced \
--tags=http-server \
--metadata-from-file=startup-script=$HOME/startup-catalogo.sh \
--metadata=enable-oslogin=TRUE \
--labels=entorno=dev,equipo=plataforma,centro-coste=tienda,aplicacion=catalogoFlag by flag:
--image-family+--image-project: always the latest patched Debian 12.--tags=http-server: a network tag that associates the VM with the firewall rule allowing port 80. Firewall rules are studied in 03-01; here it is enough to know that without the right tag the traffic does not arrive.--metadata-from-file=startup-script=...: thestartup-scriptmetadata key is special, the guest agent runs it on every boot.--metadata=enable-oslogin=TRUE: enables OS Login (section 7).--labels: billing labels, not network tags. Do not confuse--tagswith--labels: the former govern the firewall, the latter the accounting.
Let us check that it works:
# Show the external IP assigned
gcloud compute instances describe alpinashop-web-1 \
--zone=europe-west1-b \
--format="value(networkInterfaces[0].accessConfigs[0].natIP)"
# If the firewall rule does not exist yet, create it (detail in 03-01)
gcloud compute firewall-rules create permitir-http \
--allow=tcp:80 \
--target-tags=http-server \
--description="Temporary HTTP access for catalogue testing"
# Test it (the startup script takes 1-2 minutes the first time)
IP=$(gcloud compute instances describe alpinashop-web-1 \
--zone=europe-west1-b \
--format="value(networkInterfaces[0].accessConfigs[0].natIP)")
curl -s "http://$IP/"If it does not respond, the place to look is the script's log:
gcloud compute ssh alpinashop-web-1 --zone=europe-west1-b \
--command="sudo journalctl -u google-startup-scripts.service --no-pager | tail -40"
- SSH access and OS Login
There are three ways to get into a Linux VM:
| Method | How it works | Advantages | Drawbacks |
|---|---|---|---|
| SSH from the browser | The SSH button in the console; Google generates ephemeral keys | Zero configuration, works from anywhere | Requires the VM to have a public IP or IAP configured |
gcloud compute ssh |
Generates a key pair in ~/.ssh/google_compute_engine and publishes it |
Convenient, integrated with your gcloud identity | Depends on the CLI being installed |
| Your own SSH client | You add your public key to the metadata or to OS Login | Compatible with your tooling | Manual key management |
Examples:
# Log in to the VM
gcloud compute ssh alpinashop-web-1 --zone=europe-west1-b
# Run a command without an interactive session
gcloud compute ssh alpinashop-web-1 --zone=europe-west1-b \
--command="systemctl status alpinashop --no-pager"
# Copy a file to the VM
gcloud compute scp ./catalogo.csv alpinashop-web-1:/tmp/ --zone=europe-west1-b
# Log in without a public IP, through IAP (secure tunnel)
gcloud compute ssh alpinashop-web-1 --zone=europe-west1-b --tunnel-through-iapWhy OS Login is the recommendation. Without OS Login, SSH keys are stored in the metadata of the instance or the project. That has three serious problems: anyone with permission to edit metadata can add their key and log in; when someone leaves the company you have to go round the instances deleting their key by hand; and there is no clear record of who logged in.
OS Login ties the Linux SSH account to the Google Cloud identity. With it:
- Access is granted through IAM roles (
roles/compute.osLoginfor a normal user,roles/compute.osAdminLoginfor sudo), not by manipulating files. - When an employee's account is disabled in Cloud Identity, they lose access to every VM immediately.
- Logins are audited in Cloud Logging.
- Two-step verification can be required with
enable-oslogin-2fa=TRUE.
Enable it at project level so that it applies to every instance, present and future:
gcloud compute project-info add-metadata \
--metadata=enable-oslogin=TRUE
# Grant Dani normal SSH access (without sudo)
gcloud projects add-iam-policy-binding alpinashop-dev \
--member="user:[email protected]" \
--role="roles/compute.osLogin"We will look at exactly how IAM roles are composed in 03-04; here it is enough to hold on to the idea: access to servers is managed with identities, not with key files scattered around.
- Service accounts and access scopes
Every VM runs as someone. That someone is a service account: a non-human identity that the application uses to call other Google Cloud APIs. When AlpinaShop's catalogue uploads an image to alpinashop-catalogo (lesson 02-02), it will not use passwords: it will use the VM's service account.
By default the Compute Engine default service account is assigned, which has the Editor role over the whole project. That is far too permissive: if someone compromises the web application, they get almost total control of the project. The correct practice is to create a dedicated service account with only the permissions it needs:
# Create a service account of its own for the catalogue
gcloud iam service-accounts create sa-catalogo-web \
--display-name="AlpinaShop web catalogue"
# Give it only what it needs: read and write objects in Storage
gcloud projects add-iam-policy-binding alpinashop-dev \
--member="serviceAccount:[email protected]" \
--role="roles/storage.objectAdmin"
# Assign it to the VM (requires the VM to be stopped)
gcloud compute instances set-service-account alpinashop-web-1 \
--zone=europe-west1-b \
--service-account="[email protected]" \
--scopes=cloud-platformAbout scopes: they are an old mechanism that limits which APIs the VM can invoke, in addition to what IAM allows. The effective permission is the intersection of both. The current recommendation is to use --scopes=cloud-platform and control the real permissions with IAM roles, which are far more granular. All of this is developed in 03-04.
- Instance templates and managed groups (MIG)
A single VM has two problems: if it goes down, the site goes down; and if a traffic peak arrives, there is nowhere to grow. Compute Engine's answer is instance templates and managed instance groups.
- Instance template: the immutable definition of what a VM looks like (type, image, disks, network, metadata, script). It creates nothing by itself; it is a mould. Because it is immutable, changing something means creating a new template.
- Managed instance group (MIG): a set of identical VMs created from a template, which the system keeps alive and in the number you tell it.
graph TD
A[Instance template<br/>alpinashop-web-tpl-v1] --> B[MIG alpinashop-web-mig]
B --> C[VM alpinashop-web-mig-a1b2]
B --> D[VM alpinashop-web-mig-c3d4]
B --> E[VM alpinashop-web-mig-e5f6]
F[Health check<br/>GET /salud] --> B
G[Autoscaling policy<br/>target CPU 60 per cent] --> B
We create the template out of everything we already know:
gcloud compute instance-templates create alpinashop-web-tpl-v1 \
--machine-type=e2-medium \
--image-family=debian-12 \
--image-project=debian-cloud \
--boot-disk-size=20GB \
--boot-disk-type=pd-balanced \
--tags=http-server \
--metadata-from-file=startup-script=$HOME/startup-catalogo.sh \
--metadata=enable-oslogin=TRUE \
--service-account="[email protected]" \
--scopes=cloud-platform \
--labels=entorno=prod,equipo=plataforma,centro-coste=tienda,aplicacion=catalogoWe define a health check and create the regional group (spread across the zones of europe-west1, so as to survive the loss of a zone, exactly as we reasoned in 01-05):
# Health check: requests /salud every 10 s
gcloud compute health-checks create http hc-catalogo \
--port=80 \
--request-path=/salud \
--check-interval=10s \
--timeout=5s \
--healthy-threshold=2 \
--unhealthy-threshold=3
# Regional managed group with 2 initial instances
gcloud compute instance-groups managed create alpinashop-web-mig \
--template=alpinashop-web-tpl-v1 \
--size=2 \
--region=europe-west1 \
--health-check=hc-catalogo \
--initial-delay=180--initial-delay=180 matters: it tells the group to wait 3 minutes before considering a newly created instance unhealthy. Without that margin, the startup script would still be installing dependencies, the health check would fail and the MIG would enter a loop of destroying and recreating instances that never manage to boot. It is one of the classic mistakes.
- Autoscaling and auto-healing: the autumn peaks
Here is the answer to the problem AlpinaShop has been dragging along since the first lesson: during the autumn campaigns the physical server saturates and the shop goes slow exactly when it sells the most.
gcloud compute instance-groups managed set-autoscaling alpinashop-web-mig \
--region=europe-west1 \
--min-num-replicas=2 \
--max-num-replicas=10 \
--target-cpu-utilization=0.60 \
--cool-down-period=120What it does exactly:
--min-num-replicas=2: it never drops below 2 instances, in different zones. That is the minimum for a zonal outage not to leave the shop without service.--max-num-replicas=10: a safety ceiling. It protects the bill against an abnormal peak or an attack; with no ceiling, a bot could multiply your spend.--target-cpu-utilization=0.60: the autoscaler adds or removes instances to keep the group's average CPU close to 60 %. Margin is left because scaling is not instantaneous: minutes pass between the decision to create a VM and it serving traffic.--cool-down-period=120: it ignores the metrics from each instance's first 120 seconds of life, while it boots.
You can also scale on requests per second or on your own Cloud Monitoring metrics (for example, the length of a queue), which usually works better than CPU for web applications. We will look at that in 06-04.
Auto-healing. With the health check attached, if an instance stops responding on /salud for three checks in a row, the MIG deletes it and creates another one from the template. It is real self-healing, not a restart: the new instance is born clean.
This has a design consequence you have to accept: the group's instances are disposable. Nothing important can live only on their disk, because they will disappear without warning. Data goes to Cloud SQL (02-03), images to Cloud Storage (02-02) and sessions to a shared store (02-06). The server stops being a pet and becomes cattle.
To update the application you create a new template and start a progressive update:
gcloud compute instance-groups managed rolling-action start-update alpinashop-web-mig \
--region=europe-west1 \
--version=template=alpinashop-web-tpl-v2 \
--max-surge=2 \
--max-unavailable=0--max-unavailable=0 combined with --max-surge=2 means: create up to two new instances before retiring the old ones, so that capacity never drops. It is a deployment with no service interruption. If something goes wrong, you rerun the command pointing back at -v1.
What is missing for this to be a complete architecture is a load balancer that distributes traffic across the group's instances behind a single IP and a TLS certificate. That is exactly the content of lesson 03-02, and that is why we stop here.
- Spot VMs: cheap, interruptible compute
Spot VMs (the evolution of the old preemptible ones) use Google's spare capacity and cost on the order of 60–90 % less. In exchange, Google can stop them at any moment with only 30 seconds' notice.
gcloud compute instances create procesador-imagenes-spot \
--zone=europe-west1-b \
--machine-type=e2-standard-4 \
--provisioning-model=SPOT \
--instance-termination-action=DELETE \
--metadata-from-file=startup-script=$HOME/procesar-lote.sh| Aspect | Standard VM | Spot VM |
|---|---|---|
| Price | Normal rate | Much lower (varies with demand) |
| Duration | Indefinite | Can be terminated at any moment |
| Advance warning | — | 30 seconds (ACPI G2 event) |
| SLA | Yes | No |
| Suitable use | Production services | Batches, rendering, testing, CI |
Valid cases for AlpinaShop: regenerating the thumbnails of the 60 GB of product images, recalculating a heavy report, running the nightly test suite. Invalid cases: the shop's web server or anything whose interruption a customer would notice.
Good practice: catch the termination signal with a shutdown script to save progress so that the job can resume where it left off.
- The lifecycle of an instance and what is charged in each state
stateDiagram-v2
[*] --> PROVISIONING: create
PROVISIONING --> STAGING
STAGING --> RUNNING
RUNNING --> STOPPING: stop
STOPPING --> TERMINATED
TERMINATED --> STAGING: start
TERMINATED --> [*]: delete
RUNNING --> SUSPENDING: suspend
SUSPENDING --> SUSPENDED
SUSPENDED --> RUNNING: resume
This is the table that saves the most money in the whole lesson:
| State | CPU/RAM charged? | Disks charged? | External IP charged? | Contents preserved? |
|---|---|---|---|---|
RUNNING |
Yes | Yes | Yes (if static, or ephemeral and in use) | Yes |
TERMINATED (stopped) |
No | Yes | Yes if it is static and unused | Persistent disk yes; local SSD no |
SUSPENDED |
No (the storage of the dumped RAM is charged) | Yes | Same as stopped | Yes, including memory |
| Deleted | No | Only if you ticked "keep disk" | Only if the IP is static | Only what you kept |
Practical lessons:
- Stopping a VM does not remove its cost. The disks keep on being billed. A development VM stopped for months with a 500 GB disk still costs money.
- Reserved static external IPs that are not in use are charged, precisely to discourage hoarding. Release the ones you do not use.
- Switching off development environments outside working hours is the most profitable and simplest optimisation: around 128 hours a week switched off out of 168 is a direct 75 % saving on compute. It can be automated with Cloud Scheduler and a function (06-03).
Common operations:
# Stop and start
gcloud compute instances stop alpinashop-web-1 --zone=europe-west1-b
gcloud compute instances start alpinashop-web-1 --zone=europe-west1-b
# Resize (requires the instance to be stopped)
gcloud compute instances set-machine-type alpinashop-web-1 \
--zone=europe-west1-b \
--machine-type=e2-standard-2
# Grow the disk (live; afterwards you have to grow the file system)
gcloud compute disks resize alpinashop-web-1 --zone=europe-west1-b --size=50GB
gcloud compute ssh alpinashop-web-1 --zone=europe-west1-b \
--command="sudo resize2fs /dev/sda1"
# Protect against accidental deletion
gcloud compute instances update alpinashop-web-1 \
--zone=europe-west1-b --deletion-protection
# Delete while keeping the data disk
gcloud compute instances delete alpinashop-web-1 \
--zone=europe-west1-b --keep-disks=dataNote the resize2fs: growing the disk in Google Cloud does not automatically grow the file system inside the operating system. It is a step that is constantly forgotten and produces the bewilderment of seeing 50 GB in the console and 20 GB in df -h.
Common Mistakes and Tips
- Oversizing the first VM "just in case". It is cheaper to start small and grow: Compute Engine will even tell you, in the console's recommendations, that your instance is underused. Systematic optimisation is tackled in 07-05.
- Confusing
--tagswith--labels. Network tags control firewall rules; labels are billing and organisation metadata. Puttingentorno=prodas a tag does nothing useful. - Storing important data on a local SSD. It is lost when the instance is stopped. It is never "just for a moment".
- Depending on the disk of a MIG instance. The instances of a managed group are disposable by design.
- Forgetting
--initial-delayon the MIG. The group kills the instances while they are booting and enters an infinite loop of creation and destruction. - Leaving the default service account with the Editor role. It is the easiest privilege escalation to exploit if your application is compromised.
- Spreading SSH keys around through metadata. Use OS Login from the start; withdrawing access afterwards is much harder.
- Believing that stopping a VM leaves it at zero cost. Disks and static IPs still count.
- Tip: make the startup script idempotent. It runs on every boot, not only the first one.
- Tip: test the startup script on a disposable VM before putting it into a template. Debugging it inside a MIG that self-destructs is exasperating.
- Tip: name templates with a version (
-v1,-v2). They are immutable and you will need to be able to go back. - Tip: enable deletion protection on any production instance that is not part of a managed group.
Exercises
Exercise 1: choosing machine type and disk
For each AlpinaShop scenario, state the machine family and type, the disk type and a brief justification:
- The production web server for the Flask catalogue, with moderate traffic and horizontal growth planned.
- A nightly process that resizes the 60,000 product images and takes around 3 hours; it can be retried if it fails.
- A future analytics instance that loads a 400 GB dataset into memory for interactive queries.
- A development VM where Dani tests changes for about 6 hours a day.
Exercise 2: creating the instance and verifying the startup script
Starting from the alpinashop-dev project:
- Create
alpinashop-web-2ineurope-west1-cwithe2-medium, Debian 12, a 20 GB balanced disk, OS Login enabled and the four labels from AlpinaShop's scheme. - Use a startup script that installs nginx and writes a page with the instance name.
- Create the firewall rule required and check with
curlthat it responds. - Look at the startup script's log inside the instance.
- Delete the instance while keeping its boot disk and explain what cost it continues to generate.
Exercise 3: managed group with autoscaling
Marta wants to be ready for the autumn campaign:
- Create a template
alpinashop-web-tpl-ejwith the catalogue's startup script. - Create a regional MIG in
europe-west1with 2 instances and a health check on/salud. - Configure autoscaling between 2 and 6 instances with a CPU target of 65 %.
- Simulate a failure: log in to an instance, stop the service and observe what the group does.
- Explain why this architecture is still not usable facing the public and what is missing.
Solutions
Solution 1
| Scenario | Machine | Disk | Justification |
|---|---|---|---|
| 1. Catalogue web | e2-standard-2 (or e2-medium) |
pd-balanced 20 GB |
Light, steady load; growth is handled with more instances in the MIG, not with a bigger machine |
| 2. Image resizing | c3-highcpu-8 or e2-standard-8 in Spot mode |
pd-balanced + local-ssd as scratch |
CPU intensive, tolerant of interruption: Spot cuts the cost a lot. The local SSD is valid because the final result goes to Cloud Storage |
| 3. In-memory analysis | m3-ultramem-32 or similar from the M family |
pd-ssd or Hyperdisk |
Only the memory families offer that much RAM; the fast disk speeds up the initial load |
| 4. Development VM | e2-medium |
pd-balanced 20 GB |
Cheap; what matters is switching it off outside working hours, which removes the CPU and RAM cost |
Solution 2
cat > ~/startup-nginx.sh <<'SCRIPT'
#!/bin/bash
set -e
apt-get update
apt-get install -y nginx
NAME=$(curl -s -H "Metadata-Flavor: Google" \
http://metadata.google.internal/computeMetadata/v1/instance/name)
echo "<h1>AlpinaShop</h1><p>Instance: $NAME</p>" > /var/www/html/index.html
systemctl enable --now nginx
SCRIPT
gcloud compute instances create alpinashop-web-2 \
--project=alpinashop-dev \
--zone=europe-west1-c \
--machine-type=e2-medium \
--image-family=debian-12 --image-project=debian-cloud \
--boot-disk-size=20GB --boot-disk-type=pd-balanced \
--tags=http-server \
--metadata-from-file=startup-script=$HOME/startup-nginx.sh \
--metadata=enable-oslogin=TRUE \
--labels=entorno=dev,equipo=plataforma,centro-coste=tienda,aplicacion=catalogoNotice how the script gets the instance name from the metadata server (metadata.google.internal), available from inside any VM. The Metadata-Flavor: Google header is mandatory and acts as protection against requests forged from outside.
# 3. Firewall and check
gcloud compute firewall-rules create permitir-http \
--allow=tcp:80 --target-tags=http-server 2>/dev/null || true
IP=$(gcloud compute instances describe alpinashop-web-2 --zone=europe-west1-c \
--format="value(networkInterfaces[0].accessConfigs[0].natIP)")
curl -s "http://$IP/"
# 4. Startup script log
gcloud compute ssh alpinashop-web-2 --zone=europe-west1-c \
--command="sudo journalctl -u google-startup-scripts.service --no-pager | tail -30"
# 5. Delete while keeping the disk
gcloud compute instances delete alpinashop-web-2 \
--zone=europe-west1-c --keep-disks=bootAfter point 5 you are still paying for the 20 GB persistent disk, which is left orphaned in the europe-west1-c zone. Orphaned disks like this are a classic source of invisible spend: it is worth listing them periodically with gcloud compute disks list --filter="-users:*".
Solution 3
# 1. Template
gcloud compute instance-templates create alpinashop-web-tpl-ej \
--machine-type=e2-medium \
--image-family=debian-12 --image-project=debian-cloud \
--boot-disk-type=pd-balanced --boot-disk-size=20GB \
--tags=http-server \
--metadata-from-file=startup-script=$HOME/startup-catalogo.sh \
--metadata=enable-oslogin=TRUE
# 2. Health check and regional MIG
gcloud compute health-checks create http hc-catalogo-ej \
--port=80 --request-path=/salud \
--check-interval=10s --unhealthy-threshold=3
gcloud compute instance-groups managed create alpinashop-web-mig-ej \
--template=alpinashop-web-tpl-ej \
--size=2 --region=europe-west1 \
--health-check=hc-catalogo-ej --initial-delay=180
# 3. Autoscaling
gcloud compute instance-groups managed set-autoscaling alpinashop-web-mig-ej \
--region=europe-west1 \
--min-num-replicas=2 --max-num-replicas=6 \
--target-cpu-utilization=0.65 --cool-down-period=120
# 4. Simulate a failure
INST=$(gcloud compute instance-groups managed list-instances alpinashop-web-mig-ej \
--region=europe-west1 --format="value(instance)" | head -1 | xargs basename)
ZONE=$(gcloud compute instances list --filter="name=$INST" --format="value(zone)")
gcloud compute ssh "$INST" --zone="$ZONE" --command="sudo systemctl stop alpinashop"
# Watch the state of the group
watch -n 10 gcloud compute instance-groups managed list-instances \
alpinashop-web-mig-ej --region=europe-west1In point 4 you will see that, after three failed checks (around 30 seconds), the instance's state changes to UNHEALTHY and shortly afterwards the group recreates it: its name changes and it returns to HEALTHY once the startup script has finished. The service has not been restarted: the whole machine has been replaced.
In point 5, what is missing is a load balancer: today each instance has its own ephemeral IP, there is no single entry point, no HTTPS, no domain name and no traffic distribution. On top of that, the instances are directly exposed to the internet on port 80, which is unacceptable in production. All of that is solved in module 3 (03-01 VPC, 03-02 load balancing, 03-07 DNS and TLS).
Remember to clean up when you finish:
gcloud compute instance-groups managed delete alpinashop-web-mig-ej --region=europe-west1 --quiet
gcloud compute instance-templates delete alpinashop-web-tpl-ej --quiet
gcloud compute health-checks delete hc-catalogo-ej --quietConclusion
We have taken the first real step in AlpinaShop's migration. You know what a managed VM is and, more importantly, when it is still the right answer and when it is not. You know the machine families — E2 to start with, N2/N4 for general production, C3 for compute, M3 for memory, and custom types when nothing fits — and you know that the highcpu/standard/highmem suffix governs the memory ratio. You have seen the catalogue of public images, why it is better to use image families instead of fixed versions, and how to create your own base image. You have compared the disk types and taken on board two critical ideas: that the performance of a classic persistent disk depends on its size, and that a local SSD is lost when the machine is stopped. You have scheduled automatic snapshots with retention, the backup Marta had never had.
You have created alpinashop-web-1 in europe-west1-b from the console and from gcloud, with a startup script that installs Python, Flask and gunicorn and leaves the catalogue serving under systemd, including a /salud route that turned out to be the key piece of the rest of the lesson. You have learned the three routes to SSH access and why OS Login — identities instead of key files — is the only defensible option in a company. You have replaced the default service account with sa-catalogo-web, with minimal permissions. And, above all, you have turned a single machine into a template and a regional managed group with autoscaling from 2 to 10 instances and auto-healing: the mechanism that will make AlpinaShop's autumn campaign stop being a problem. Along the way you have taken on the change of mindset the cloud demands: instances are disposable, and therefore data has to live somewhere else.
That "somewhere else" is exactly the subject of the next lesson. In 02-02, Cloud Storage, we will move the 60 GB of product images that today occupy the local disk of the physical server to the alpinashop-catalogo bucket: we will look at the object model and why directories do not exist, choose a storage class and a location, do the real transfer with gcloud storage rsync, protect the bucket with uniform access and signed URLs, automate cost reduction with lifecycle rules and connect the Flask application to Storage with the Python client library. The instances will be able to die without taking anything with them.
Google Cloud Platform (GCP) Course
Module 1: Introduction to Google Cloud Platform
- What is Google Cloud Platform?
- Setting Up Your GCP Account
- A Tour of the GCP Console
- Projects, Resource Hierarchy and Billing
- Regions, Zones and the Shared Responsibility Model
- Cloud Shell and the gcloud CLI
Module 2: Core GCP Services
- Compute Engine: Virtual Machines on Google Cloud
- Cloud Storage: Object Storage
- Cloud SQL: Managed Relational Databases
- App Engine: Platform as a Service
- Google Kubernetes Engine (GKE)
- NoSQL Databases: Firestore, Bigtable and Spanner
- How to Choose the Right Compute Service
Module 3: Networking and Security
- VPC Networks
- Cloud Load Balancing
- Cloud CDN
- Identity and Access Management (IAM)
- Cloud Armor
- Secrets and Encryption: Secret Manager and Cloud KMS
- Cloud DNS, TLS Certificates and Publishing Services Securely
Module 4: Data and Analytics
- BigQuery: The Analytical Data Warehouse
- Cloud Dataflow: Batch and Streaming Data Processing
- Cloud Dataproc: Managed Spark and Hadoop
- Cloud Pub/Sub: Asynchronous Messaging
- Cloud Data Fusion: Code-Free Data Integration
- Orchestrating Pipelines with Cloud Composer and Workflows
- Data Governance and Dashboards with Dataplex and Looker Studio
Module 5: Machine Learning and AI
- Vertex AI: The Machine Learning Platform on GCP
- AutoML: Custom Models Without Writing Code
- TensorFlow on GCP: Training and Serving Models
- Natural Language API
- Vision API
- Generative AI on Vertex AI: Gemini Models and Embeddings
- MLOps: From Model to Product with Vertex AI Pipelines
Module 6: DevOps and Monitoring
- Cloud Build: Continuous Integration on GCP
- Cloud Source Repositories and Source Code Management
- Cloud Functions: Serverless Functions
- Cloud Monitoring (formerly Stackdriver): Metrics, Dashboards and Alerts
- Cloud Deployment Manager and Native Infrastructure as Code
- Cloud Logging and Cloud Trace: Logs, Traces and Diagnostics
- Terraform on GCP: Infrastructure as Code in Practice
Module 7: Advanced GCP Topics
- Hybrid and Multicloud with Anthos
- Serverless Computing with Cloud Run
- Advanced Networking: Shared VPC, Peering and Hybrid Connectivity
- Security Best Practices
- Cost Management and Optimization
- Reliability: SLOs, High Availability and Disaster Recovery
- Governance at Scale: Organization, Policies and Auditing
