So far, AlpinaShop's migration has taken the most conservative route: moving what was there onto virtual machines. It works, but Marta is still maintaining operating systems, writing startup scripts, versioning instance templates and watching for security patches. It is work that does not distinguish AlpinaShop from any other shop: nobody buys more backpacks because a server is well patched.

App Engine was Google Cloud's first service, launched in 2008, and it is still one of the fastest ways to get a web application into production: you upload the code, Google takes care of everything else — servers, scaling, load balancing, TLS certificates, deployments — and you pay for usage. In this lesson you will deploy AlpinaShop's Flask catalogue on App Engine, learn to control its scaling and its cost, split traffic between versions to do a canary deployment, connect to Cloud SQL and Cloud Storage, and finish with an honest assessment of why many teams today choose Cloud Run instead.

Contents

  1. What a PaaS is and what disappears from your job
  2. The standard environment and the flexible environment
  3. Anatomy of app.yaml
  4. Deploying AlpinaShop's catalogue
  5. Versions, traffic splitting and canary deployment
  6. Scaling types and their effect on cost
  7. Services and routing with dispatch.yaml
  8. Reaching Cloud SQL and Cloud Storage from App Engine
  9. Environment variables and configuration
  10. Logs and errors
  11. Background work: Cloud Tasks and cron.yaml
  12. Real limitations and why many teams today choose Cloud Run

  1. What a PaaS is and what disappears from your job

Platform as a service means the provider manages everything below your code. Compared with what we did in 02-01:

Responsibility Compute Engine (IaaS) App Engine (PaaS)
Hardware and physical network Google Google
Operating system and patches You Google
Runtime (Python, system dependencies) You Google
Web server and load balancing You Google
TLS certificate and HTTPS You Google (automatic)
Scaling and health checks You (MIG + autoscaler) Google
Deployment and rollback You (templates + rolling update) Google (one command)
Application code You You
Data schema and queries You You

The whole of lesson 02-01 — templates, managed groups, health checks, autoscaling, progressive updates — comes down in App Engine to one configuration file and gcloud app deploy. And you get HTTPS with your own domain for free, something that in Compute Engine will require a load balancer and a managed certificate (03-02 and 03-07).

What you pay for it: less control and more coupling. You do not choose the operating system, you do not install arbitrary packages in the standard environment, and app.yaml is a Google-specific format that works nowhere else.

One organisational detail that surprises people: one App Engine application per project, and its region is chosen when the application is created and cannot be changed. If you get it wrong, you have to create a new project. For AlpinaShop, europe-west1, consistent with everything else.

gcloud app create --region=europe-west1 --project=alpinashop-dev

  1. The standard environment and the flexible environment

App Engine has two environments that behave very differently under the same name:

Standard Flexible
Execution base Google-managed sandbox Compute Engine VM with Docker
Languages Python, Java, Node.js, Go, PHP, Ruby (supported versions) Any, through your own container
Start-up time Seconds Minutes
Scale to zero Yes No (minimum 1 instance)
Cost with no traffic Zero (or nearly) That of at least one VM 24×7
File system access Only /tmp, in memory The VM's disk, writing allowed
Run your own binaries No Yes
SSH to the instance No Yes
Maximum request duration 10 min (automatic scaling) / 24 h (basic and manual) 60 min
Deployment Seconds to a couple of minutes Several minutes (it builds an image)
Free tier Yes, a daily quota of instance hours No

When to use each. The standard environment is the natural choice for web applications in a supported language: it starts fast, scales to zero and is cheap. The flexible one makes sense if you need a system dependency the sandbox does not allow, your own binary or an unsupported language.

And here is the honest assessment that anticipates section 12: if you are thinking about App Engine flexible, Cloud Run is almost always the better option today — it also runs your container, but it scales to zero, is billed per request and is not tied to the App Engine format.

AlpinaShop uses the standard environment with Python 3.12: the catalogue is pure Flask with dependencies installed via pip, with nothing exotic.

  1. Anatomy of app.yaml

app.yaml is the file that describes the whole application. Let us build AlpinaShop's while commenting on each block.

# Managed runtime: Google maintains the base image, the interpreter and the patches.
runtime: python312

# The command that starts the application. Without it, App Engine looks by convention
# for an 'app' object in main.py, but it is better to be explicit.
# 'gunicorn' is the WSGI server; App Engine exposes the app on $PORT.
entrypoint: gunicorn -b :$PORT -w 2 --timeout 60 main:app

# Instance class: defines the CPU and memory of each instance.
#   F1 = 256 MB / 600 MHz  (cheapest, the free tier one)
#   F2 = 512 MB / 1.2 GHz
#   F4 = 1 GB   / 2.4 GHz
#   F4_1G = 2 GB / 2.4 GHz
instance_class: F2

# Scaling policy: see section 6.
automatic_scaling:
  min_instances: 0            # scales to zero when there is no traffic
  max_instances: 10           # safety ceiling for the bill
  target_cpu_utilization: 0.6 # creates instances when CPU goes above 60 %
  target_throughput_utilization: 0.6
  max_concurrent_requests: 20 # simultaneous requests per instance
  min_pending_latency: 100ms  # wait before deciding to create an instance
  max_pending_latency: 500ms  # if a request waits longer, create an instance now

# NON-sensitive environment variables. Secrets go in Secret Manager (03-06).
env_variables:
  BUCKET_CATALOGO: "alpinashop-catalogo"
  INSTANCIA_SQL: "alpinashop-prod:europe-west1:alpinashop-pedidos"
  DB_USER: "app_catalogo"
  DB_NAME: "tienda"
  ENTORNO: "produccion"

# Connector for reaching Cloud SQL over a private IP or VPC resources (03-01).
vpc_access_connector:
  name: projects/alpinashop-prod/locations/europe-west1/connectors/conector-vpc

# Request routing, in order. The first match wins.
handlers:
  # Static assets served directly by Google's infrastructure,
  # without consuming application instances or CPU time.
  - url: /estatico
    static_dir: estatico
    secure: always            # forces HTTPS
    expiration: "30d"         # Cache-Control header of 30 days

  - url: /favicon\.ico
    static_files: estatico/favicon.ico
    upload: estatico/favicon\.ico

  # Internal panel: only users authenticated with an organization account.
  - url: /admin/.*
    script: auto
    secure: always
    login: admin

  # Everything else goes to the application.
  - url: /.*
    script: auto
    secure: always

Three ideas from this file worth holding on to:

  • Static-type handlers are free in CPU time. Google serves those files from its own infrastructure, without starting or occupying an instance. Serving CSS, JS and images that way instead of through Flask reduces cost and latency noticeably.
  • secure: always should be on every handler. It redirects HTTP to HTTPS automatically.
  • max_concurrent_requests is the biggest lever on cost. If your application spends most of its time waiting for the database, it can handle many simultaneous requests per instance; raising this value reduces the number of instances needed. If the application is CPU-bound, raising it makes latency worse.

Alongside app.yaml go the usual files of a Python project:

alpinashop-catalogo/
├── app.yaml
├── main.py
├── requirements.txt
├── .gcloudignore
└── estatico/
    ├── estilo.css
    └── favicon.ico
# requirements.txt
Flask==3.0.*
gunicorn==22.0.*
SQLAlchemy==2.0.*
cloud-sql-python-connector[pg8000]==1.*
google-cloud-storage==2.*
# .gcloudignore: what is NOT uploaded on deployment
.git/
.gitignore
__pycache__/
*.pyc
venv/
.env
tests/
README.md

The .gcloudignore is no minor detail: without it, you would upload the .git directory and the entire virtual environment, with slow deployments and the very real risk of publishing an .env file with credentials.

  1. Deploying AlpinaShop's catalogue

The application, with the pieces we already know from the previous lessons:

# main.py
import os
import sqlalchemy
from flask import Flask, jsonify, render_template_string
from google.cloud.sql.connector import Connector, IPTypes
from google.cloud import storage

app = Flask(__name__)

connector = Connector()
storage_client = storage.Client()
bucket = storage_client.bucket(os.environ["BUCKET_CATALOGO"])


def _connect_db():
    return connector.connect(
        os.environ["INSTANCIA_SQL"],
        "pg8000",
        user=os.environ["DB_USER"],
        password=os.environ["DB_PASS"],
        db=os.environ["DB_NAME"],
        ip_type=IPTypes.PUBLIC,
    )


engine = sqlalchemy.create_engine(
    "postgresql+pg8000://",
    creator=_connect_db,
    pool_size=2,          # deliberately low: App Engine creates MANY instances
    max_overflow=1,
    pool_recycle=1800,
    pool_pre_ping=True,
)


@app.route("/")
def catalog():
    query = sqlalchemy.text(
        "SELECT sku, nombre, precio FROM tienda.productos "
        "WHERE activo = true ORDER BY nombre LIMIT 50"
    )
    with engine.connect() as conn:
        products = conn.execute(query).mappings().all()

    rows = "".join(
        f"<li>{p['sku']} — {p['nombre']} — {p['precio']:.2f} €</li>" for p in products
    )
    version = os.environ.get("GAE_VERSION", "local")
    return render_template_string(
        f"<h1>AlpinaShop</h1><ul>{rows}</ul><p>Version: {version}</p>"
    )


@app.route("/salud")
def health():
    return "ok", 200


if __name__ == "__main__":
    # Local development only; on App Engine gunicorn does the starting.
    app.run(host="127.0.0.1", port=8080, debug=True)

Note pool_size=2. It is a deliberate change from the previous lesson: App Engine can create dozens of instances at a peak, and each one would open its own pool. With pool_size=5 and 20 instances that would be 100 connections from the catalogue alone. In environments that scale aggressively, pools must be small.

Notice GAE_VERSION too: App Engine injects environment variables with information about the runtime environment (GAE_SERVICE, GAE_VERSION, GAE_INSTANCE, GOOGLE_CLOUD_PROJECT). Showing the version on the page is a very practical trick for verifying deployments and traffic splits.

Deployment:

# Test locally before deploying
export BUCKET_CATALOGO=alpinashop-catalogo
export INSTANCIA_SQL=alpinashop-prod:europe-west1:alpinashop-pedidos
export DB_USER=app_catalogo DB_PASS=... DB_NAME=tienda
python main.py

# Deploy with an explicit version identifier
gcloud app deploy app.yaml \
  --version=v1-catalogo \
  --project=alpinashop-dev

# Open it in the browser
gcloud app browse

# Show the assigned URL
gcloud app describe --format="value(defaultHostname)"

When you deploy, App Engine builds the package, installs the dependencies from requirements.txt, creates a new version and — unless you say otherwise — sends it 100 % of the traffic. The default URL is https://<projectId>.<region>.r.appspot.com, with a valid TLS certificate from the very first second and without you having configured anything.

Always use --version with a meaningful name. Without it, App Engine generates a timestamped identifier that is hard to read in the version list and in the traffic-splitting commands.

  1. Versions, traffic splitting and canary deployment

Every deployment creates an immutable version that stays available. This enables something very powerful: splitting traffic between versions.

# Deploy without sending traffic: the version is available but inactive
gcloud app deploy app.yaml --version=v2-precios --no-promote

# Test it on its own URL, without affecting customers
gcloud app browse --version=v2-precios
# https://v2-precios-dot-alpinashop-dev.ew.r.appspot.com

# Canary: 10 % of the traffic to the new version
gcloud app services set-traffic default \
  --splits=v1-catalogo=0.9,v2-precios=0.1

# If all goes well, increase progressively
gcloud app services set-traffic default \
  --splits=v1-catalogo=0.5,v2-precios=0.5

# Full promotion
gcloud app services set-traffic default --splits=v2-precios=1

# Immediate rollback if something fails
gcloud app services set-traffic default --splits=v1-catalogo=1

The URL with the <version>-dot-<service>-dot-<project> prefix is a very useful detail: any deployed version is directly reachable, which lets you validate in the real environment before directing users to it.

On the splitting criterion, there are two modes and the difference matters:

# Random split per request: each request may go to a different version
gcloud app services set-traffic default --splits=v1=0.9,v2=0.1 --split-by=random

# Split by cookie: the same user always sees the same version
gcloud app services set-traffic default --splits=v1=0.9,v2=0.1 --split-by=cookie

For a canary of a user interface, cookie is almost always the right thing: with random, a customer could see one page on the old version and the next on the new one, with bewildering results.

And a cost warning: old versions with minimum instances configured keep billing even if they receive no traffic. Clean up periodically:

gcloud app versions list --filter="traffic_split=0"
gcloud app versions delete v1-catalogo --quiet

  1. Scaling types and their effect on cost

App Engine standard offers three scaling modes, and choosing badly is the main source of unexpected bills.

Automatic Basic Manual
Creates instances based on Traffic, CPU and latency Arrival of requests Never: a fixed number
Scales to zero Yes (if min_instances: 0) Yes, after inactivity No
Cold start Yes Yes No
Max. request duration 10 minutes 24 hours 24 hours
Typical use Web applications Long, sporadic tasks Services with in-memory state, constant load
# Automatic: the general case
automatic_scaling:
  min_instances: 0
  max_instances: 10
  min_idle_instances: 0        # "warm" instances held in reserve
  max_concurrent_requests: 20
  target_cpu_utilization: 0.6
# Basic: for a batch processing service
basic_scaling:
  max_instances: 3
  idle_timeout: 10m            # shuts the instance down after 10 min with no requests
# Manual: a fixed number of instances, always on
manual_scaling:
  instances: 2

The cold start problem. With min_instances: 0 you pay nothing when there is no traffic, but the first request after a period of inactivity has to wait for an instance to start: loading the interpreter, importing the dependencies, opening the connection pool. In Python with Flask it is usually 1–3 seconds; if the code imports heavy libraries, longer.

The solution is to reserve warm instances:

automatic_scaling:
  min_instances: 1        # always 1 live instance
  min_idle_instances: 1   # plus 1 in reserve, ready for peaks

And here is the trade-off, with indicative figures:

Configuration Approximate cost with no traffic Latency of the first request
min_instances: 0 ~€0 1–3 s (cold start)
min_instances: 1, class F2 One F2 instance 24×7: on the order of €30–40/month Immediate
min_instances: 2 Double that Immediate, with redundancy

Indicative figures for reasoning about the order of magnitude. Check the official Google Cloud calculator.

AlpinaShop's decision, consistent with its reality: on alpinashop-dev, min_instances: 0 — Dani does not mind waiting two seconds and the environment spends the night with no traffic. In production, min_instances: 1 and max_instances: 10, because a customer who sees the shop take three seconds to load is a customer who leaves. During the autumn campaign the minimum goes up to 3 and the maximum to 20.

Additional tips for reducing cold starts: import heavy libraries lazily, do not do expensive work at module import time, and keep requirements.txt free of dependencies you do not use.

  1. Services and routing with dispatch.yaml

An App Engine application can have several services (previously called modules), each with its own app.yaml, its own scaling and its own versions. It is the way to separate components with different profiles.

For AlpinaShop:

alpinashop/
├── dispatch.yaml
├── web/          -> app.yaml with "service: default"  (public catalogue)
├── api/          -> app.yaml with "service: api"      (orders API)
└── admin/        -> app.yaml with "service: admin"    (internal panel)
# api/app.yaml
runtime: python312
service: api
instance_class: F2
entrypoint: gunicorn -b :$PORT -w 2 main:app
automatic_scaling:
  min_instances: 1
  max_instances: 20
# admin/app.yaml
runtime: python312
service: admin
instance_class: F1
entrypoint: gunicorn -b :$PORT -w 1 main:app
basic_scaling:            # the panel is barely used: not worth keeping it warm
  max_instances: 2
  idle_timeout: 15m

Each service has its own URL (https://api-dot-alpinashop-prod.ew.r.appspot.com). To serve them all under a single domain with different paths, you use dispatch.yaml:

dispatch:
  - url: "*/api/*"
    service: api

  - url: "*/admin/*"
    service: admin

  - url: "*/*"
    service: default
gcloud app deploy web/app.yaml api/app.yaml admin/app.yaml dispatch.yaml

Advantages of separating services: each one scales according to its own load (the administration panel does not need 10 instances just because there is a peak in the shop), they are deployed independently, and a failure in the panel does not bring down the catalogue. It is a separation of failure domains, the same idea we applied to zones in 01-05.

Bear in mind that the rules in dispatch.yaml are evaluated in order and only a limited number are allowed; for complex routing you use a load balancer (03-02).

  1. Reaching Cloud SQL and Cloud Storage from App Engine

Here you see the value of having used client libraries in the previous lessons: the code does not change.

Every App Engine application has a default service account, <projectId>@appspot.gserviceaccount.com. It is enough to give it the right permissions:

PROJECT=alpinashop-prod
SA="[email protected]"

# Access to Cloud SQL
gcloud projects add-iam-policy-binding $PROJECT \
  --member="serviceAccount:$SA" --role="roles/cloudsql.client"

# Reading and writing objects in the catalogue bucket
gcloud storage buckets add-iam-policy-binding gs://alpinashop-catalogo \
  --member="serviceAccount:$SA" --role="roles/storage.objectUser"

# Reading the secret holding the database password (03-06)
gcloud secrets add-iam-policy-binding db-password-catalogo \
  --member="serviceAccount:$SA" --role="roles/secretmanager.secretAccessor"

With that, the same storage.Client() and the same Cloud SQL connector from lessons 02-02 and 02-03 work without touching a line, because the Application Default Credentials now resolve to App Engine's service account.

One peculiarity of the standard environment: the file system is read-only except for /tmp, which additionally lives in memory and counts against the instance's memory. Never store anything there that has to persist, nor large files. Images uploaded by users go straight to Cloud Storage, exactly as we programmed in 02-02 with upload_from_file over the stream.

  1. Environment variables and configuration

Non-sensitive variables go in app.yaml, under env_variables. Sensitive ones do not: app.yaml ends up in the Git repository and its values are visible to anyone who can see the version's configuration.

The correct pattern is to read the secrets at start-up from Secret Manager:

import os
from google.cloud import secretmanager


def read_secret(name: str, version: str = "latest") -> str:
    """Reads a secret from Secret Manager using App Engine's service account."""
    project = os.environ["GOOGLE_CLOUD_PROJECT"]
    client = secretmanager.SecretManagerServiceClient()
    path = f"projects/{project}/secrets/{name}/versions/{version}"
    response = client.access_secret_version(request={"name": path})
    return response.payload.data.decode("UTF-8")


# It is read ONCE when the instance starts, not on every request:
# every call to Secret Manager has latency and a cost per operation.
DB_PASS = read_secret("db-password-catalogo")

To distinguish environments, the usual approach is to have one app.yaml per environment (app-dev.yaml, app-prod.yaml) and deploy the appropriate one:

gcloud app deploy app-dev.yaml  --project=alpinashop-dev
gcloud app deploy app-prod.yaml --project=alpinashop-prod

The full detail of Secret Manager, secret rotation and encryption with Cloud KMS is in lesson 03-06.

  1. Logs and errors

App Engine automatically sends the requests, and everything the application writes to standard output, to Cloud Logging. Nothing needs configuring.

# Follow the logs in real time
gcloud app logs tail -s default

# Last 50 entries from a specific version
gcloud app logs read --version=v2-precios --limit=50

# Advanced query: 5xx errors from the last hour
gcloud logging read \
  'resource.type="gae_app" AND httpRequest.status>=500' \
  --limit=20 --freshness=1h --format="value(timestamp, httpRequest.status, httpRequest.requestUrl)"

For the logs to be genuinely useful, write them structured as JSON: Cloud Logging indexes them by field and you will be able to filter by sku or by severity instead of searching for strings.

import json
import logging
import sys

def structured_log(message, severity="INFO", **fields):
    entry = {"message": message, "severity": severity, **fields}
    print(json.dumps(entry), file=sys.stdout)

structured_log("product viewed", sku="MOC-40", duration_ms=42)
structured_log("insufficient stock", severity="WARNING", sku="BOT-GTX", requested=5)

Uncaught errors arrive automatically in Error Reporting, which groups them by stack trace, counts occurrences and can warn you when a new error appears. It is one of the most profitable things the platform offers with no configuration at all.

The in-depth treatment of Cloud Logging, metrics and traces is in 06-06 and 06-04.

  1. Background work: Cloud Tasks and cron.yaml

An HTTP request must not do long work: the user is waiting, and under automatic scaling there is a 10-minute limit. When AlpinaShop confirms an order, an email has to be sent, the stock updated and the warehouse notified; none of that should block the response to the customer.

Cloud Tasks is a managed queue: you enqueue a task, which consists of a deferred HTTP request to your own application, and the service delivers it with retries and rate control.

gcloud tasks queues create cola-pedidos \
  --location=europe-west1 \
  --max-dispatches-per-second=10 \
  --max-concurrent-dispatches=20 \
  --max-attempts=5
import json
from google.cloud import tasks_v2

tasks_client = tasks_v2.CloudTasksClient()
QUEUE = tasks_client.queue_path("alpinashop-prod", "europe-west1", "cola-pedidos")


def enqueue_confirmation(order_id: int):
    """Enqueues the processing that follows the confirmation of an order."""
    task = {
        "app_engine_http_request": {
            "http_method": tasks_v2.HttpMethod.POST,
            "relative_uri": "/tareas/procesar-pedido",
            "headers": {"Content-Type": "application/json"},
            "body": json.dumps({"pedido_id": order_id}).encode(),
        }
    }
    return tasks_client.create_task(request={"parent": QUEUE, "task": task})


@app.route("/tareas/procesar-pedido", methods=["POST"])
def process_order():
    # App Engine guarantees that only Cloud Tasks can set this header:
    # external requests carrying it are rejected by the infrastructure.
    if "X-AppEngine-TaskName" not in request.headers:
        return "forbidden", 403

    data = request.get_json()
    # ... send email, update stock, notify the warehouse ...

    # Returning 2xx marks the task as completed.
    # Any other code triggers a retry with exponential backoff,
    # so the handler MUST be idempotent.
    return "", 204

Two critical points here: checking the X-AppEngine-TaskName header as an access control, and idempotency. Cloud Tasks guarantees at least once delivery, not exactly once: if your handler is slow and there is a retry, you could send two emails or deduct the stock twice. Record which orders you have already processed and check before acting.

Cron. For scheduled tasks, cron.yaml invokes a URL of your application on a schedule:

cron:
  - description: "Daily export of orders for analytics"
    url: /tareas/exportar-pedidos
    schedule: every day 03:00
    timezone: Europe/Madrid
    target: api

  - description: "Abandoned shopping cart reminder"
    url: /tareas/carritos-abandonados
    schedule: every 6 hours

  - description: "Weekly sales report for Lucia"
    url: /tareas/informe-semanal
    schedule: every monday 07:00
    timezone: Europe/Madrid
gcloud app deploy cron.yaml
gcloud app describe --format="value(name)"   # verify the deployment

The endpoints invoked by cron receive the X-Appengine-Cron: true header, which is checked just like the Tasks one. Protect those routes with a login: admin handler in app.yaml as well.

The equivalent service outside App Engine is Cloud Scheduler, which can invoke HTTP, Pub/Sub or Cloud Functions and works with any compute service. It is the option to use if your application does not live on App Engine.

  1. Real limitations and why many teams today choose Cloud Run

App Engine is still a good product and there are applications that have been running on it for more than a decade without incident. But it is worth being honest about its limits in 2026:

  • Coupling to Google's format. app.yaml, dispatch.yaml and cron.yaml do not exist outside App Engine. Migrating to another provider — or even to another Google service — means redoing the configuration.
  • Limited runtimes on their own timetable. Only the languages and versions Google supports; when a Python version becomes obsolete, you have to migrate on Google's schedule.
  • A restrictive sandbox in the standard environment. No custom binaries, system dependencies or writing to disk.
  • The flexible environment has been left behind. It starts in minutes, does not scale to zero and costs more than container-based alternatives.
  • Immutable region and one application per project. Rigid for architectures that evolve.
  • The investment does not compound. Time spent mastering app.yaml is of no use anywhere else; time spent mastering containers is of use everywhere.

Cloud Run (lesson 07-02) occupies that space today: it runs your container, scales to zero, is billed by usage, offers traffic splitting between revisions just like App Engine, automatic HTTPS, and what you deploy is a standard container image that would also run on GKE, on your laptop or on another cloud.

App Engine standard Cloud Run
Unit of deployment Code + app.yaml Container image
Languages The supported ones Any
Scale to zero Yes Yes
Traffic splitting Yes Yes (revisions)
Portability Low High
Configuration Google-specific file Standard container + flags
System dependencies Limited Whatever you put in the Dockerfile

When App Engine still makes sense: you already use it and it works; you want the shortest route from Python code to a URL with HTTPS; you do not have containers and do not fancy learning them right now; or you specifically need dispatch.yaml and the built-in cron.

When to choose Cloud Run: a new project, you already work with containers, you want portability, or you need control over the execution environment.

For AlpinaShop, the decision we will document in 02-07 is to start with Compute Engine (lift-and-shift, already done) and evolve towards containers. App Engine remains an option that is known and rejected for reasons, which is the best way to reject something.

Common Mistakes and Tips

  • Creating the application in the wrong region. It is irreversible: it demands a new project. Check before gcloud app create.
  • Deploying without .gcloudignore. Slow deployments and the risk of uploading .env or .git.
  • Forgetting --no-promote when testing. You will send 100 % of the traffic to an unvalidated version.
  • Using --split-by=random for an interface canary. The user sees different versions on consecutive requests.
  • Leaving old versions with min_instances > 0. They bill without receiving traffic.
  • Setting min_instances: 0 in production without measuring. The cold start eats your conversion rate.
  • Large connection pools. Multiplied by many instances, they exhaust Cloud SQL's max_connections.
  • Storing files on the local file system. Only /tmp, in memory and ephemeral.
  • Secrets in env_variables in app.yaml. They end up in Git. Use Secret Manager.
  • Non-idempotent Cloud Tasks handlers. Delivery is at least once.
  • Serving static files from Flask. Use static_dir handlers: they are faster and consume no instances.
  • Tip: put secure: always on every handler.
  • Tip: use --version with readable names.
  • Tip: show GAE_VERSION in the application so you can tell at a glance which version served each request.
  • Tip: always set max_instances. It is your emergency brake against an abnormal peak or an attack.

Exercises

Exercise 1: first deployment and controlling the scaling

  1. Create the App Engine application in europe-west1 on alpinashop-dev.
  2. Write a minimal Flask application that shows "AlpinaShop" and the version (GAE_VERSION), with a /salud route.
  3. Write an app.yaml with the python312 runtime, class F1, automatic scaling from 0 to 4 instances and a static handler for /estatico.
  4. Deploy it as version v1 and check the URL.
  5. Explain what you would change in app.yaml for production and justify the associated cost.

Exercise 2: canary and rollback

  1. Modify the application to show prices including VAT and deploy it as v2 without sending it traffic.
  2. Check v2 on its specific URL.
  3. Send 20 % of the traffic to v2 with cookie-based splitting.
  4. Check the split by making several requests and observing the version shown.
  5. Simulate a failure and perform a full rollback to v1. Then clean up the versions with no traffic.

Exercise 3: services, cron and tasks

  1. Create a second service api with its own app.yaml and its own scaling, exposing /api/pedidos.
  2. Write a dispatch.yaml that sends /api/* to the api service and everything else to default.
  3. Add a cron.yaml with a daily task at 3:00 (Madrid time) that calls /tareas/exportar-pedidos on the api service.
  4. Protect the cron endpoint by checking the corresponding header.
  5. Explain why a task handler must be idempotent and how you would guarantee it for sending an order confirmation email.

Solutions

Solution 1

gcloud app create --region=europe-west1 --project=alpinashop-dev
# main.py
import os
from flask import Flask

app = Flask(__name__)


@app.route("/")
def home():
    version = os.environ.get("GAE_VERSION", "local")
    instance = os.environ.get("GAE_INSTANCE", "local")[:8]
    return f"<h1>AlpinaShop</h1><p>Version: {version}</p><p>Instance: {instance}</p>"


@app.route("/salud")
def health():
    return "ok", 200
# app.yaml
runtime: python312
entrypoint: gunicorn -b :$PORT -w 2 main:app
instance_class: F1

automatic_scaling:
  min_instances: 0
  max_instances: 4
  max_concurrent_requests: 20
  target_cpu_utilization: 0.6

handlers:
  - url: /estatico
    static_dir: estatico
    secure: always
    expiration: "7d"
  - url: /.*
    script: auto
    secure: always
gcloud app deploy app.yaml --version=v1 --project=alpinashop-dev
gcloud app browse

For production I would change to instance_class: F2 (more memory for the connection pool and the client libraries), min_instances: 1 to remove the cold start for the first customer, and max_instances: 10 to absorb the autumn peaks. The added cost is that of keeping one F2 instance running 24 hours a day, on the order of €30–40 a month: a small figure compared with the value of the shop's front page always loading in under a second. It is worth checking the amount in the official calculator before committing.

Solution 2

# 1. Deploy without promoting
gcloud app deploy app.yaml --version=v2 --no-promote

# 2. Test the specific version
gcloud app browse --version=v2
# or directly:
curl -s "https://v2-dot-alpinashop-dev.ew.r.appspot.com/"

# 3. 20 % canary with cookie affinity
gcloud app services set-traffic default \
  --splits=v1=0.8,v2=0.2 --split-by=cookie

# 4. Check the split (with -c/-b the cookie is kept between requests)
for i in $(seq 1 10); do
  curl -s "https://alpinashop-dev.ew.r.appspot.com/" | grep -o "Version: v[0-9]"
done

# 5. Rollback and cleanup
gcloud app services set-traffic default --splits=v1=1
gcloud app versions list --filter="traffic_split=0"
gcloud app versions delete v2 --quiet

With --split-by=cookie, if you keep the cookie between requests you will always see the same version: that is the desired behaviour, so as not to bewilder the user. Without a cookie — as in the curl loop above, where every call is a new session — the split behaves as random, and around 2 in every 10 responses will show v2.

Solution 3

# api/app.yaml
runtime: python312
service: api
instance_class: F2
entrypoint: gunicorn -b :$PORT -w 2 main:app
automatic_scaling:
  min_instances: 0
  max_instances: 5
handlers:
  - url: /tareas/.*
    script: auto
    secure: always
    login: admin        # blocks external access to the task endpoints
  - url: /.*
    script: auto
    secure: always
# dispatch.yaml
dispatch:
  - url: "*/api/*"
    service: api
  - url: "*/tareas/*"
    service: api
  - url: "*/*"
    service: default
# cron.yaml
cron:
  - description: "Daily export of orders"
    url: /tareas/exportar-pedidos
    schedule: every day 03:00
    timezone: Europe/Madrid
    target: api
# 4. Protecting the cron endpoint
from flask import request

@app.route("/tareas/exportar-pedidos")
def export_orders():
    if request.headers.get("X-Appengine-Cron") != "true":
        return "forbidden", 403
    # ... export to gs://alpinashop-catalogo/exportaciones/<date>/ ...
    return "", 204
gcloud app deploy web/app.yaml api/app.yaml dispatch.yaml cron.yaml
  1. The handler must be idempotent because Cloud Tasks guarantees at least once delivery: if the handler takes longer than expected, if it returns a transient error or if the instance is recycled halfway through, the task is retried and the same order arrives twice. With no protection, the customer would receive two confirmation emails and the stock would be deducted twice.

The way to guarantee it for the confirmation email is to record the fact transactionally in the database:

CREATE TABLE tienda.correos_enviados (
    pedido_id  BIGINT PRIMARY KEY,
    tipo       TEXT NOT NULL,
    enviado_en TIMESTAMPTZ NOT NULL DEFAULT now()
);
with engine.begin() as conn:
    result = conn.execute(
        sqlalchemy.text(
            "INSERT INTO tienda.correos_enviados (pedido_id, tipo) "
            "VALUES (:pid, 'confirmacion') ON CONFLICT DO NOTHING"
        ),
        {"pid": order_id},
    )
    if result.rowcount == 0:
        return "", 204   # already sent: the task ends without repeating the email

send_confirmation_email(order_id)

The key is the ON CONFLICT DO NOTHING on a primary key: the database guarantees that only one execution manages to insert the row, and the others exit without sending anything. Returning 204 on the second attempt also avoids infinite retries.

Conclusion

You have seen what going up a level of abstraction really means. In App Engine the operating system, the patches, the web server, the load balancing, the TLS certificates, the health checks and the instance templates all disappear: everything that filled lesson 02-01 comes down to an app.yaml and a gcloud app deploy. You know the decisive difference between the standard environment — sandbox, start-up in seconds, scaling to zero, cheap — and the flexible one — containers, start-up in minutes, no scaling to zero — and you know that the flexible one has today been displaced by Cloud Run.

You have dissected app.yaml line by line: runtime, entrypoint, instance class, scaling policy, environment variables and handlers, including the idea of serving static files from Google's infrastructure so as not to spend instances. You have deployed AlpinaShop's Flask catalogue and obtained a URL with HTTPS without configuring anything. You have learned to deploy with --no-promote, validate on the version's URL, split traffic with cookie affinity for a canary and roll back with a single command: deployment practices that in Compute Engine required versioned templates and progressive updates. You have understood the three scaling modes and the central trade-off between min_instances: 0 and the cold start, with AlpinaShop's reasoned decision: zero in development, one in production, three during the campaign. You have separated services with dispatch.yaml so that the administration panel and the catalogue scale independently, you have connected Cloud SQL and Cloud Storage without changing a line of code thanks to the default credentials, you have read secrets from Secret Manager at start-up, you have seen structured logs and Error Reporting, and you have moved the heavy work out of the request with Cloud Tasks and cron.yaml, learning along the way what at least once delivery means and how to write an idempotent handler.

And you have finished with an honest assessment: App Engine is convenient but coupled to Google's format, and the investment you make in it does not compound outside it. The alternative that does compound is containers. In 02-05, Google Kubernetes Engine, we will take that step: we will go over Kubernetes from scratch — container, pod, Deployment, Service, node and control plane — compare the Autopilot and Standard modes, create the alpinashop-cluster cluster, build and publish the catalogue image in Artifact Registry, write the Deployment and Service manifests, scale with a HorizontalPodAutoscaler and do updates with no interruption together with their rollback. It is the level of abstraction with the most power and also the most responsibility in the whole module.

Google Cloud Platform (GCP) Course

Module 1: Introduction to Google Cloud Platform

Module 2: Core GCP Services

Module 3: Networking and Security

Module 4: Data and Analytics

Module 5: Machine Learning and AI

Module 6: DevOps and Monitoring

Module 7: Advanced GCP Topics

Module 8: Final Project

© Copyright 2026. All rights reserved