So far, AlpinaShop's migration has taken the most conservative route: moving what was there onto virtual machines. It works, but Marta is still maintaining operating systems, writing startup scripts, versioning instance templates and watching for security patches. It is work that does not distinguish AlpinaShop from any other shop: nobody buys more backpacks because a server is well patched.
App Engine was Google Cloud's first service, launched in 2008, and it is still one of the fastest ways to get a web application into production: you upload the code, Google takes care of everything else — servers, scaling, load balancing, TLS certificates, deployments — and you pay for usage. In this lesson you will deploy AlpinaShop's Flask catalogue on App Engine, learn to control its scaling and its cost, split traffic between versions to do a canary deployment, connect to Cloud SQL and Cloud Storage, and finish with an honest assessment of why many teams today choose Cloud Run instead.
Contents
- What a PaaS is and what disappears from your job
- The standard environment and the flexible environment
- Anatomy of
app.yaml - Deploying AlpinaShop's catalogue
- Versions, traffic splitting and canary deployment
- Scaling types and their effect on cost
- Services and routing with
dispatch.yaml - Reaching Cloud SQL and Cloud Storage from App Engine
- Environment variables and configuration
- Logs and errors
- Background work: Cloud Tasks and
cron.yaml - Real limitations and why many teams today choose Cloud Run
- What a PaaS is and what disappears from your job
Platform as a service means the provider manages everything below your code. Compared with what we did in 02-01:
| Responsibility | Compute Engine (IaaS) | App Engine (PaaS) |
|---|---|---|
| Hardware and physical network | ||
| Operating system and patches | You | |
| Runtime (Python, system dependencies) | You | |
| Web server and load balancing | You | |
| TLS certificate and HTTPS | You | Google (automatic) |
| Scaling and health checks | You (MIG + autoscaler) | |
| Deployment and rollback | You (templates + rolling update) | Google (one command) |
| Application code | You | You |
| Data schema and queries | You | You |
The whole of lesson 02-01 — templates, managed groups, health checks, autoscaling, progressive updates — comes down in App Engine to one configuration file and gcloud app deploy. And you get HTTPS with your own domain for free, something that in Compute Engine will require a load balancer and a managed certificate (03-02 and 03-07).
What you pay for it: less control and more coupling. You do not choose the operating system, you do not install arbitrary packages in the standard environment, and app.yaml is a Google-specific format that works nowhere else.
One organisational detail that surprises people: one App Engine application per project, and its region is chosen when the application is created and cannot be changed. If you get it wrong, you have to create a new project. For AlpinaShop, europe-west1, consistent with everything else.
- The standard environment and the flexible environment
App Engine has two environments that behave very differently under the same name:
| Standard | Flexible | |
|---|---|---|
| Execution base | Google-managed sandbox | Compute Engine VM with Docker |
| Languages | Python, Java, Node.js, Go, PHP, Ruby (supported versions) | Any, through your own container |
| Start-up time | Seconds | Minutes |
| Scale to zero | Yes | No (minimum 1 instance) |
| Cost with no traffic | Zero (or nearly) | That of at least one VM 24×7 |
| File system access | Only /tmp, in memory |
The VM's disk, writing allowed |
| Run your own binaries | No | Yes |
| SSH to the instance | No | Yes |
| Maximum request duration | 10 min (automatic scaling) / 24 h (basic and manual) | 60 min |
| Deployment | Seconds to a couple of minutes | Several minutes (it builds an image) |
| Free tier | Yes, a daily quota of instance hours | No |
When to use each. The standard environment is the natural choice for web applications in a supported language: it starts fast, scales to zero and is cheap. The flexible one makes sense if you need a system dependency the sandbox does not allow, your own binary or an unsupported language.
And here is the honest assessment that anticipates section 12: if you are thinking about App Engine flexible, Cloud Run is almost always the better option today — it also runs your container, but it scales to zero, is billed per request and is not tied to the App Engine format.
AlpinaShop uses the standard environment with Python 3.12: the catalogue is pure Flask with dependencies installed via pip, with nothing exotic.
- Anatomy of
app.yaml
app.yamlapp.yaml is the file that describes the whole application. Let us build AlpinaShop's while commenting on each block.
# Managed runtime: Google maintains the base image, the interpreter and the patches.
runtime: python312
# The command that starts the application. Without it, App Engine looks by convention
# for an 'app' object in main.py, but it is better to be explicit.
# 'gunicorn' is the WSGI server; App Engine exposes the app on $PORT.
entrypoint: gunicorn -b :$PORT -w 2 --timeout 60 main:app
# Instance class: defines the CPU and memory of each instance.
# F1 = 256 MB / 600 MHz (cheapest, the free tier one)
# F2 = 512 MB / 1.2 GHz
# F4 = 1 GB / 2.4 GHz
# F4_1G = 2 GB / 2.4 GHz
instance_class: F2
# Scaling policy: see section 6.
automatic_scaling:
min_instances: 0 # scales to zero when there is no traffic
max_instances: 10 # safety ceiling for the bill
target_cpu_utilization: 0.6 # creates instances when CPU goes above 60 %
target_throughput_utilization: 0.6
max_concurrent_requests: 20 # simultaneous requests per instance
min_pending_latency: 100ms # wait before deciding to create an instance
max_pending_latency: 500ms # if a request waits longer, create an instance now
# NON-sensitive environment variables. Secrets go in Secret Manager (03-06).
env_variables:
BUCKET_CATALOGO: "alpinashop-catalogo"
INSTANCIA_SQL: "alpinashop-prod:europe-west1:alpinashop-pedidos"
DB_USER: "app_catalogo"
DB_NAME: "tienda"
ENTORNO: "produccion"
# Connector for reaching Cloud SQL over a private IP or VPC resources (03-01).
vpc_access_connector:
name: projects/alpinashop-prod/locations/europe-west1/connectors/conector-vpc
# Request routing, in order. The first match wins.
handlers:
# Static assets served directly by Google's infrastructure,
# without consuming application instances or CPU time.
- url: /estatico
static_dir: estatico
secure: always # forces HTTPS
expiration: "30d" # Cache-Control header of 30 days
- url: /favicon\.ico
static_files: estatico/favicon.ico
upload: estatico/favicon\.ico
# Internal panel: only users authenticated with an organization account.
- url: /admin/.*
script: auto
secure: always
login: admin
# Everything else goes to the application.
- url: /.*
script: auto
secure: alwaysThree ideas from this file worth holding on to:
- Static-type
handlersare free in CPU time. Google serves those files from its own infrastructure, without starting or occupying an instance. Serving CSS, JS and images that way instead of through Flask reduces cost and latency noticeably. secure: alwaysshould be on every handler. It redirects HTTP to HTTPS automatically.max_concurrent_requestsis the biggest lever on cost. If your application spends most of its time waiting for the database, it can handle many simultaneous requests per instance; raising this value reduces the number of instances needed. If the application is CPU-bound, raising it makes latency worse.
Alongside app.yaml go the usual files of a Python project:
alpinashop-catalogo/
├── app.yaml
├── main.py
├── requirements.txt
├── .gcloudignore
└── estatico/
├── estilo.css
└── favicon.ico# requirements.txt
Flask==3.0.*
gunicorn==22.0.*
SQLAlchemy==2.0.*
cloud-sql-python-connector[pg8000]==1.*
google-cloud-storage==2.*# .gcloudignore: what is NOT uploaded on deployment
.git/
.gitignore
__pycache__/
*.pyc
venv/
.env
tests/
README.mdThe .gcloudignore is no minor detail: without it, you would upload the .git directory and the entire virtual environment, with slow deployments and the very real risk of publishing an .env file with credentials.
- Deploying AlpinaShop's catalogue
The application, with the pieces we already know from the previous lessons:
# main.py
import os
import sqlalchemy
from flask import Flask, jsonify, render_template_string
from google.cloud.sql.connector import Connector, IPTypes
from google.cloud import storage
app = Flask(__name__)
connector = Connector()
storage_client = storage.Client()
bucket = storage_client.bucket(os.environ["BUCKET_CATALOGO"])
def _connect_db():
return connector.connect(
os.environ["INSTANCIA_SQL"],
"pg8000",
user=os.environ["DB_USER"],
password=os.environ["DB_PASS"],
db=os.environ["DB_NAME"],
ip_type=IPTypes.PUBLIC,
)
engine = sqlalchemy.create_engine(
"postgresql+pg8000://",
creator=_connect_db,
pool_size=2, # deliberately low: App Engine creates MANY instances
max_overflow=1,
pool_recycle=1800,
pool_pre_ping=True,
)
@app.route("/")
def catalog():
query = sqlalchemy.text(
"SELECT sku, nombre, precio FROM tienda.productos "
"WHERE activo = true ORDER BY nombre LIMIT 50"
)
with engine.connect() as conn:
products = conn.execute(query).mappings().all()
rows = "".join(
f"<li>{p['sku']} — {p['nombre']} — {p['precio']:.2f} €</li>" for p in products
)
version = os.environ.get("GAE_VERSION", "local")
return render_template_string(
f"<h1>AlpinaShop</h1><ul>{rows}</ul><p>Version: {version}</p>"
)
@app.route("/salud")
def health():
return "ok", 200
if __name__ == "__main__":
# Local development only; on App Engine gunicorn does the starting.
app.run(host="127.0.0.1", port=8080, debug=True)Note pool_size=2. It is a deliberate change from the previous lesson: App Engine can create dozens of instances at a peak, and each one would open its own pool. With pool_size=5 and 20 instances that would be 100 connections from the catalogue alone. In environments that scale aggressively, pools must be small.
Notice GAE_VERSION too: App Engine injects environment variables with information about the runtime environment (GAE_SERVICE, GAE_VERSION, GAE_INSTANCE, GOOGLE_CLOUD_PROJECT). Showing the version on the page is a very practical trick for verifying deployments and traffic splits.
Deployment:
# Test locally before deploying
export BUCKET_CATALOGO=alpinashop-catalogo
export INSTANCIA_SQL=alpinashop-prod:europe-west1:alpinashop-pedidos
export DB_USER=app_catalogo DB_PASS=... DB_NAME=tienda
python main.py
# Deploy with an explicit version identifier
gcloud app deploy app.yaml \
--version=v1-catalogo \
--project=alpinashop-dev
# Open it in the browser
gcloud app browse
# Show the assigned URL
gcloud app describe --format="value(defaultHostname)"When you deploy, App Engine builds the package, installs the dependencies from requirements.txt, creates a new version and — unless you say otherwise — sends it 100 % of the traffic. The default URL is https://<projectId>.<region>.r.appspot.com, with a valid TLS certificate from the very first second and without you having configured anything.
Always use --version with a meaningful name. Without it, App Engine generates a timestamped identifier that is hard to read in the version list and in the traffic-splitting commands.
- Versions, traffic splitting and canary deployment
Every deployment creates an immutable version that stays available. This enables something very powerful: splitting traffic between versions.
# Deploy without sending traffic: the version is available but inactive
gcloud app deploy app.yaml --version=v2-precios --no-promote
# Test it on its own URL, without affecting customers
gcloud app browse --version=v2-precios
# https://v2-precios-dot-alpinashop-dev.ew.r.appspot.com
# Canary: 10 % of the traffic to the new version
gcloud app services set-traffic default \
--splits=v1-catalogo=0.9,v2-precios=0.1
# If all goes well, increase progressively
gcloud app services set-traffic default \
--splits=v1-catalogo=0.5,v2-precios=0.5
# Full promotion
gcloud app services set-traffic default --splits=v2-precios=1
# Immediate rollback if something fails
gcloud app services set-traffic default --splits=v1-catalogo=1The URL with the <version>-dot-<service>-dot-<project> prefix is a very useful detail: any deployed version is directly reachable, which lets you validate in the real environment before directing users to it.
On the splitting criterion, there are two modes and the difference matters:
# Random split per request: each request may go to a different version
gcloud app services set-traffic default --splits=v1=0.9,v2=0.1 --split-by=random
# Split by cookie: the same user always sees the same version
gcloud app services set-traffic default --splits=v1=0.9,v2=0.1 --split-by=cookieFor a canary of a user interface, cookie is almost always the right thing: with random, a customer could see one page on the old version and the next on the new one, with bewildering results.
And a cost warning: old versions with minimum instances configured keep billing even if they receive no traffic. Clean up periodically:
- Scaling types and their effect on cost
App Engine standard offers three scaling modes, and choosing badly is the main source of unexpected bills.
| Automatic | Basic | Manual | |
|---|---|---|---|
| Creates instances based on | Traffic, CPU and latency | Arrival of requests | Never: a fixed number |
| Scales to zero | Yes (if min_instances: 0) |
Yes, after inactivity | No |
| Cold start | Yes | Yes | No |
| Max. request duration | 10 minutes | 24 hours | 24 hours |
| Typical use | Web applications | Long, sporadic tasks | Services with in-memory state, constant load |
# Automatic: the general case
automatic_scaling:
min_instances: 0
max_instances: 10
min_idle_instances: 0 # "warm" instances held in reserve
max_concurrent_requests: 20
target_cpu_utilization: 0.6# Basic: for a batch processing service
basic_scaling:
max_instances: 3
idle_timeout: 10m # shuts the instance down after 10 min with no requestsThe cold start problem. With min_instances: 0 you pay nothing when there is no traffic, but the first request after a period of inactivity has to wait for an instance to start: loading the interpreter, importing the dependencies, opening the connection pool. In Python with Flask it is usually 1–3 seconds; if the code imports heavy libraries, longer.
The solution is to reserve warm instances:
automatic_scaling:
min_instances: 1 # always 1 live instance
min_idle_instances: 1 # plus 1 in reserve, ready for peaksAnd here is the trade-off, with indicative figures:
| Configuration | Approximate cost with no traffic | Latency of the first request |
|---|---|---|
min_instances: 0 |
~€0 | 1–3 s (cold start) |
min_instances: 1, class F2 |
One F2 instance 24×7: on the order of €30–40/month | Immediate |
min_instances: 2 |
Double that | Immediate, with redundancy |
Indicative figures for reasoning about the order of magnitude. Check the official Google Cloud calculator.
AlpinaShop's decision, consistent with its reality: on alpinashop-dev, min_instances: 0 — Dani does not mind waiting two seconds and the environment spends the night with no traffic. In production, min_instances: 1 and max_instances: 10, because a customer who sees the shop take three seconds to load is a customer who leaves. During the autumn campaign the minimum goes up to 3 and the maximum to 20.
Additional tips for reducing cold starts: import heavy libraries lazily, do not do expensive work at module import time, and keep requirements.txt free of dependencies you do not use.
- Services and routing with
dispatch.yaml
dispatch.yamlAn App Engine application can have several services (previously called modules), each with its own app.yaml, its own scaling and its own versions. It is the way to separate components with different profiles.
For AlpinaShop:
alpinashop/ ├── dispatch.yaml ├── web/ -> app.yaml with "service: default" (public catalogue) ├── api/ -> app.yaml with "service: api" (orders API) └── admin/ -> app.yaml with "service: admin" (internal panel)
# api/app.yaml
runtime: python312
service: api
instance_class: F2
entrypoint: gunicorn -b :$PORT -w 2 main:app
automatic_scaling:
min_instances: 1
max_instances: 20# admin/app.yaml
runtime: python312
service: admin
instance_class: F1
entrypoint: gunicorn -b :$PORT -w 1 main:app
basic_scaling: # the panel is barely used: not worth keeping it warm
max_instances: 2
idle_timeout: 15mEach service has its own URL (https://api-dot-alpinashop-prod.ew.r.appspot.com). To serve them all under a single domain with different paths, you use dispatch.yaml:
dispatch:
- url: "*/api/*"
service: api
- url: "*/admin/*"
service: admin
- url: "*/*"
service: defaultAdvantages of separating services: each one scales according to its own load (the administration panel does not need 10 instances just because there is a peak in the shop), they are deployed independently, and a failure in the panel does not bring down the catalogue. It is a separation of failure domains, the same idea we applied to zones in 01-05.
Bear in mind that the rules in dispatch.yaml are evaluated in order and only a limited number are allowed; for complex routing you use a load balancer (03-02).
- Reaching Cloud SQL and Cloud Storage from App Engine
Here you see the value of having used client libraries in the previous lessons: the code does not change.
Every App Engine application has a default service account, <projectId>@appspot.gserviceaccount.com. It is enough to give it the right permissions:
PROJECT=alpinashop-prod
SA="[email protected]"
# Access to Cloud SQL
gcloud projects add-iam-policy-binding $PROJECT \
--member="serviceAccount:$SA" --role="roles/cloudsql.client"
# Reading and writing objects in the catalogue bucket
gcloud storage buckets add-iam-policy-binding gs://alpinashop-catalogo \
--member="serviceAccount:$SA" --role="roles/storage.objectUser"
# Reading the secret holding the database password (03-06)
gcloud secrets add-iam-policy-binding db-password-catalogo \
--member="serviceAccount:$SA" --role="roles/secretmanager.secretAccessor"With that, the same storage.Client() and the same Cloud SQL connector from lessons 02-02 and 02-03 work without touching a line, because the Application Default Credentials now resolve to App Engine's service account.
One peculiarity of the standard environment: the file system is read-only except for /tmp, which additionally lives in memory and counts against the instance's memory. Never store anything there that has to persist, nor large files. Images uploaded by users go straight to Cloud Storage, exactly as we programmed in 02-02 with upload_from_file over the stream.
- Environment variables and configuration
Non-sensitive variables go in app.yaml, under env_variables. Sensitive ones do not: app.yaml ends up in the Git repository and its values are visible to anyone who can see the version's configuration.
The correct pattern is to read the secrets at start-up from Secret Manager:
import os
from google.cloud import secretmanager
def read_secret(name: str, version: str = "latest") -> str:
"""Reads a secret from Secret Manager using App Engine's service account."""
project = os.environ["GOOGLE_CLOUD_PROJECT"]
client = secretmanager.SecretManagerServiceClient()
path = f"projects/{project}/secrets/{name}/versions/{version}"
response = client.access_secret_version(request={"name": path})
return response.payload.data.decode("UTF-8")
# It is read ONCE when the instance starts, not on every request:
# every call to Secret Manager has latency and a cost per operation.
DB_PASS = read_secret("db-password-catalogo")To distinguish environments, the usual approach is to have one app.yaml per environment (app-dev.yaml, app-prod.yaml) and deploy the appropriate one:
gcloud app deploy app-dev.yaml --project=alpinashop-dev
gcloud app deploy app-prod.yaml --project=alpinashop-prodThe full detail of Secret Manager, secret rotation and encryption with Cloud KMS is in lesson 03-06.
- Logs and errors
App Engine automatically sends the requests, and everything the application writes to standard output, to Cloud Logging. Nothing needs configuring.
# Follow the logs in real time
gcloud app logs tail -s default
# Last 50 entries from a specific version
gcloud app logs read --version=v2-precios --limit=50
# Advanced query: 5xx errors from the last hour
gcloud logging read \
'resource.type="gae_app" AND httpRequest.status>=500' \
--limit=20 --freshness=1h --format="value(timestamp, httpRequest.status, httpRequest.requestUrl)"For the logs to be genuinely useful, write them structured as JSON: Cloud Logging indexes them by field and you will be able to filter by sku or by severity instead of searching for strings.
import json
import logging
import sys
def structured_log(message, severity="INFO", **fields):
entry = {"message": message, "severity": severity, **fields}
print(json.dumps(entry), file=sys.stdout)
structured_log("product viewed", sku="MOC-40", duration_ms=42)
structured_log("insufficient stock", severity="WARNING", sku="BOT-GTX", requested=5)Uncaught errors arrive automatically in Error Reporting, which groups them by stack trace, counts occurrences and can warn you when a new error appears. It is one of the most profitable things the platform offers with no configuration at all.
The in-depth treatment of Cloud Logging, metrics and traces is in 06-06 and 06-04.
- Background work: Cloud Tasks and
cron.yaml
cron.yamlAn HTTP request must not do long work: the user is waiting, and under automatic scaling there is a 10-minute limit. When AlpinaShop confirms an order, an email has to be sent, the stock updated and the warehouse notified; none of that should block the response to the customer.
Cloud Tasks is a managed queue: you enqueue a task, which consists of a deferred HTTP request to your own application, and the service delivers it with retries and rate control.
gcloud tasks queues create cola-pedidos \
--location=europe-west1 \
--max-dispatches-per-second=10 \
--max-concurrent-dispatches=20 \
--max-attempts=5import json
from google.cloud import tasks_v2
tasks_client = tasks_v2.CloudTasksClient()
QUEUE = tasks_client.queue_path("alpinashop-prod", "europe-west1", "cola-pedidos")
def enqueue_confirmation(order_id: int):
"""Enqueues the processing that follows the confirmation of an order."""
task = {
"app_engine_http_request": {
"http_method": tasks_v2.HttpMethod.POST,
"relative_uri": "/tareas/procesar-pedido",
"headers": {"Content-Type": "application/json"},
"body": json.dumps({"pedido_id": order_id}).encode(),
}
}
return tasks_client.create_task(request={"parent": QUEUE, "task": task})
@app.route("/tareas/procesar-pedido", methods=["POST"])
def process_order():
# App Engine guarantees that only Cloud Tasks can set this header:
# external requests carrying it are rejected by the infrastructure.
if "X-AppEngine-TaskName" not in request.headers:
return "forbidden", 403
data = request.get_json()
# ... send email, update stock, notify the warehouse ...
# Returning 2xx marks the task as completed.
# Any other code triggers a retry with exponential backoff,
# so the handler MUST be idempotent.
return "", 204Two critical points here: checking the X-AppEngine-TaskName header as an access control, and idempotency. Cloud Tasks guarantees at least once delivery, not exactly once: if your handler is slow and there is a retry, you could send two emails or deduct the stock twice. Record which orders you have already processed and check before acting.
Cron. For scheduled tasks, cron.yaml invokes a URL of your application on a schedule:
cron:
- description: "Daily export of orders for analytics"
url: /tareas/exportar-pedidos
schedule: every day 03:00
timezone: Europe/Madrid
target: api
- description: "Abandoned shopping cart reminder"
url: /tareas/carritos-abandonados
schedule: every 6 hours
- description: "Weekly sales report for Lucia"
url: /tareas/informe-semanal
schedule: every monday 07:00
timezone: Europe/MadridThe endpoints invoked by cron receive the X-Appengine-Cron: true header, which is checked just like the Tasks one. Protect those routes with a login: admin handler in app.yaml as well.
The equivalent service outside App Engine is Cloud Scheduler, which can invoke HTTP, Pub/Sub or Cloud Functions and works with any compute service. It is the option to use if your application does not live on App Engine.
- Real limitations and why many teams today choose Cloud Run
App Engine is still a good product and there are applications that have been running on it for more than a decade without incident. But it is worth being honest about its limits in 2026:
- Coupling to Google's format.
app.yaml,dispatch.yamlandcron.yamldo not exist outside App Engine. Migrating to another provider — or even to another Google service — means redoing the configuration. - Limited runtimes on their own timetable. Only the languages and versions Google supports; when a Python version becomes obsolete, you have to migrate on Google's schedule.
- A restrictive sandbox in the standard environment. No custom binaries, system dependencies or writing to disk.
- The flexible environment has been left behind. It starts in minutes, does not scale to zero and costs more than container-based alternatives.
- Immutable region and one application per project. Rigid for architectures that evolve.
- The investment does not compound. Time spent mastering
app.yamlis of no use anywhere else; time spent mastering containers is of use everywhere.
Cloud Run (lesson 07-02) occupies that space today: it runs your container, scales to zero, is billed by usage, offers traffic splitting between revisions just like App Engine, automatic HTTPS, and what you deploy is a standard container image that would also run on GKE, on your laptop or on another cloud.
| App Engine standard | Cloud Run | |
|---|---|---|
| Unit of deployment | Code + app.yaml |
Container image |
| Languages | The supported ones | Any |
| Scale to zero | Yes | Yes |
| Traffic splitting | Yes | Yes (revisions) |
| Portability | Low | High |
| Configuration | Google-specific file | Standard container + flags |
| System dependencies | Limited | Whatever you put in the Dockerfile |
When App Engine still makes sense: you already use it and it works; you want the shortest route from Python code to a URL with HTTPS; you do not have containers and do not fancy learning them right now; or you specifically need dispatch.yaml and the built-in cron.
When to choose Cloud Run: a new project, you already work with containers, you want portability, or you need control over the execution environment.
For AlpinaShop, the decision we will document in 02-07 is to start with Compute Engine (lift-and-shift, already done) and evolve towards containers. App Engine remains an option that is known and rejected for reasons, which is the best way to reject something.
Common Mistakes and Tips
- Creating the application in the wrong region. It is irreversible: it demands a new project. Check before
gcloud app create. - Deploying without
.gcloudignore. Slow deployments and the risk of uploading.envor.git. - Forgetting
--no-promotewhen testing. You will send 100 % of the traffic to an unvalidated version. - Using
--split-by=randomfor an interface canary. The user sees different versions on consecutive requests. - Leaving old versions with
min_instances > 0. They bill without receiving traffic. - Setting
min_instances: 0in production without measuring. The cold start eats your conversion rate. - Large connection pools. Multiplied by many instances, they exhaust Cloud SQL's
max_connections. - Storing files on the local file system. Only
/tmp, in memory and ephemeral. - Secrets in
env_variablesinapp.yaml. They end up in Git. Use Secret Manager. - Non-idempotent Cloud Tasks handlers. Delivery is at least once.
- Serving static files from Flask. Use
static_dirhandlers: they are faster and consume no instances. - Tip: put
secure: alwayson every handler. - Tip: use
--versionwith readable names. - Tip: show
GAE_VERSIONin the application so you can tell at a glance which version served each request. - Tip: always set
max_instances. It is your emergency brake against an abnormal peak or an attack.
Exercises
Exercise 1: first deployment and controlling the scaling
- Create the App Engine application in
europe-west1onalpinashop-dev. - Write a minimal Flask application that shows "AlpinaShop" and the version (
GAE_VERSION), with a/saludroute. - Write an
app.yamlwith thepython312runtime, classF1, automatic scaling from 0 to 4 instances and a static handler for/estatico. - Deploy it as version
v1and check the URL. - Explain what you would change in
app.yamlfor production and justify the associated cost.
Exercise 2: canary and rollback
- Modify the application to show prices including VAT and deploy it as
v2without sending it traffic. - Check
v2on its specific URL. - Send 20 % of the traffic to
v2with cookie-based splitting. - Check the split by making several requests and observing the version shown.
- Simulate a failure and perform a full rollback to
v1. Then clean up the versions with no traffic.
Exercise 3: services, cron and tasks
- Create a second service
apiwith its ownapp.yamland its own scaling, exposing/api/pedidos. - Write a
dispatch.yamlthat sends/api/*to theapiservice and everything else todefault. - Add a
cron.yamlwith a daily task at 3:00 (Madrid time) that calls/tareas/exportar-pedidoson theapiservice. - Protect the cron endpoint by checking the corresponding header.
- Explain why a task handler must be idempotent and how you would guarantee it for sending an order confirmation email.
Solutions
Solution 1
# main.py
import os
from flask import Flask
app = Flask(__name__)
@app.route("/")
def home():
version = os.environ.get("GAE_VERSION", "local")
instance = os.environ.get("GAE_INSTANCE", "local")[:8]
return f"<h1>AlpinaShop</h1><p>Version: {version}</p><p>Instance: {instance}</p>"
@app.route("/salud")
def health():
return "ok", 200# app.yaml
runtime: python312
entrypoint: gunicorn -b :$PORT -w 2 main:app
instance_class: F1
automatic_scaling:
min_instances: 0
max_instances: 4
max_concurrent_requests: 20
target_cpu_utilization: 0.6
handlers:
- url: /estatico
static_dir: estatico
secure: always
expiration: "7d"
- url: /.*
script: auto
secure: alwaysFor production I would change to instance_class: F2 (more memory for the connection pool and the client libraries), min_instances: 1 to remove the cold start for the first customer, and max_instances: 10 to absorb the autumn peaks. The added cost is that of keeping one F2 instance running 24 hours a day, on the order of €30–40 a month: a small figure compared with the value of the shop's front page always loading in under a second. It is worth checking the amount in the official calculator before committing.
Solution 2
# 1. Deploy without promoting
gcloud app deploy app.yaml --version=v2 --no-promote
# 2. Test the specific version
gcloud app browse --version=v2
# or directly:
curl -s "https://v2-dot-alpinashop-dev.ew.r.appspot.com/"
# 3. 20 % canary with cookie affinity
gcloud app services set-traffic default \
--splits=v1=0.8,v2=0.2 --split-by=cookie
# 4. Check the split (with -c/-b the cookie is kept between requests)
for i in $(seq 1 10); do
curl -s "https://alpinashop-dev.ew.r.appspot.com/" | grep -o "Version: v[0-9]"
done
# 5. Rollback and cleanup
gcloud app services set-traffic default --splits=v1=1
gcloud app versions list --filter="traffic_split=0"
gcloud app versions delete v2 --quietWith --split-by=cookie, if you keep the cookie between requests you will always see the same version: that is the desired behaviour, so as not to bewilder the user. Without a cookie — as in the curl loop above, where every call is a new session — the split behaves as random, and around 2 in every 10 responses will show v2.
Solution 3
# api/app.yaml
runtime: python312
service: api
instance_class: F2
entrypoint: gunicorn -b :$PORT -w 2 main:app
automatic_scaling:
min_instances: 0
max_instances: 5
handlers:
- url: /tareas/.*
script: auto
secure: always
login: admin # blocks external access to the task endpoints
- url: /.*
script: auto
secure: always# dispatch.yaml
dispatch:
- url: "*/api/*"
service: api
- url: "*/tareas/*"
service: api
- url: "*/*"
service: default# cron.yaml
cron:
- description: "Daily export of orders"
url: /tareas/exportar-pedidos
schedule: every day 03:00
timezone: Europe/Madrid
target: api# 4. Protecting the cron endpoint
from flask import request
@app.route("/tareas/exportar-pedidos")
def export_orders():
if request.headers.get("X-Appengine-Cron") != "true":
return "forbidden", 403
# ... export to gs://alpinashop-catalogo/exportaciones/<date>/ ...
return "", 204- The handler must be idempotent because Cloud Tasks guarantees at least once delivery: if the handler takes longer than expected, if it returns a transient error or if the instance is recycled halfway through, the task is retried and the same order arrives twice. With no protection, the customer would receive two confirmation emails and the stock would be deducted twice.
The way to guarantee it for the confirmation email is to record the fact transactionally in the database:
CREATE TABLE tienda.correos_enviados (
pedido_id BIGINT PRIMARY KEY,
tipo TEXT NOT NULL,
enviado_en TIMESTAMPTZ NOT NULL DEFAULT now()
);with engine.begin() as conn:
result = conn.execute(
sqlalchemy.text(
"INSERT INTO tienda.correos_enviados (pedido_id, tipo) "
"VALUES (:pid, 'confirmacion') ON CONFLICT DO NOTHING"
),
{"pid": order_id},
)
if result.rowcount == 0:
return "", 204 # already sent: the task ends without repeating the email
send_confirmation_email(order_id)The key is the ON CONFLICT DO NOTHING on a primary key: the database guarantees that only one execution manages to insert the row, and the others exit without sending anything. Returning 204 on the second attempt also avoids infinite retries.
Conclusion
You have seen what going up a level of abstraction really means. In App Engine the operating system, the patches, the web server, the load balancing, the TLS certificates, the health checks and the instance templates all disappear: everything that filled lesson 02-01 comes down to an app.yaml and a gcloud app deploy. You know the decisive difference between the standard environment — sandbox, start-up in seconds, scaling to zero, cheap — and the flexible one — containers, start-up in minutes, no scaling to zero — and you know that the flexible one has today been displaced by Cloud Run.
You have dissected app.yaml line by line: runtime, entrypoint, instance class, scaling policy, environment variables and handlers, including the idea of serving static files from Google's infrastructure so as not to spend instances. You have deployed AlpinaShop's Flask catalogue and obtained a URL with HTTPS without configuring anything. You have learned to deploy with --no-promote, validate on the version's URL, split traffic with cookie affinity for a canary and roll back with a single command: deployment practices that in Compute Engine required versioned templates and progressive updates. You have understood the three scaling modes and the central trade-off between min_instances: 0 and the cold start, with AlpinaShop's reasoned decision: zero in development, one in production, three during the campaign. You have separated services with dispatch.yaml so that the administration panel and the catalogue scale independently, you have connected Cloud SQL and Cloud Storage without changing a line of code thanks to the default credentials, you have read secrets from Secret Manager at start-up, you have seen structured logs and Error Reporting, and you have moved the heavy work out of the request with Cloud Tasks and cron.yaml, learning along the way what at least once delivery means and how to write an idempotent handler.
And you have finished with an honest assessment: App Engine is convenient but coupled to Google's format, and the investment you make in it does not compound outside it. The alternative that does compound is containers. In 02-05, Google Kubernetes Engine, we will take that step: we will go over Kubernetes from scratch — container, pod, Deployment, Service, node and control plane — compare the Autopilot and Standard modes, create the alpinashop-cluster cluster, build and publish the catalogue image in Artifact Registry, write the Deployment and Service manifests, scale with a HorizontalPodAutoscaler and do updates with no interruption together with their rollback. It is the level of abstraction with the most power and also the most responsibility in the whole module.
Google Cloud Platform (GCP) Course
Module 1: Introduction to Google Cloud Platform
- What is Google Cloud Platform?
- Setting Up Your GCP Account
- A Tour of the GCP Console
- Projects, Resource Hierarchy and Billing
- Regions, Zones and the Shared Responsibility Model
- Cloud Shell and the gcloud CLI
Module 2: Core GCP Services
- Compute Engine: Virtual Machines on Google Cloud
- Cloud Storage: Object Storage
- Cloud SQL: Managed Relational Databases
- App Engine: Platform as a Service
- Google Kubernetes Engine (GKE)
- NoSQL Databases: Firestore, Bigtable and Spanner
- How to Choose the Right Compute Service
Module 3: Networking and Security
- VPC Networks
- Cloud Load Balancing
- Cloud CDN
- Identity and Access Management (IAM)
- Cloud Armor
- Secrets and Encryption: Secret Manager and Cloud KMS
- Cloud DNS, TLS Certificates and Publishing Services Securely
Module 4: Data and Analytics
- BigQuery: The Analytical Data Warehouse
- Cloud Dataflow: Batch and Streaming Data Processing
- Cloud Dataproc: Managed Spark and Hadoop
- Cloud Pub/Sub: Asynchronous Messaging
- Cloud Data Fusion: Code-Free Data Integration
- Orchestrating Pipelines with Cloud Composer and Workflows
- Data Governance and Dashboards with Dataplex and Looker Studio
Module 5: Machine Learning and AI
- Vertex AI: The Machine Learning Platform on GCP
- AutoML: Custom Models Without Writing Code
- TensorFlow on GCP: Training and Serving Models
- Natural Language API
- Vision API
- Generative AI on Vertex AI: Gemini Models and Embeddings
- MLOps: From Model to Product with Vertex AI Pipelines
Module 6: DevOps and Monitoring
- Cloud Build: Continuous Integration on GCP
- Cloud Source Repositories and Source Code Management
- Cloud Functions: Serverless Functions
- Cloud Monitoring (formerly Stackdriver): Metrics, Dashboards and Alerts
- Cloud Deployment Manager and Native Infrastructure as Code
- Cloud Logging and Cloud Trace: Logs, Traces and Diagnostics
- Terraform on GCP: Infrastructure as Code in Practice
Module 7: Advanced GCP Topics
- Hybrid and Multicloud with Anthos
- Serverless Computing with Cloud Run
- Advanced Networking: Shared VPC, Peering and Hybrid Connectivity
- Security Best Practices
- Cost Management and Optimization
- Reliability: SLOs, High Availability and Disaster Recovery
- Governance at Scale: Organization, Policies and Auditing
