In the previous lesson we reached an uncomfortable conclusion: the instances of a managed group are disposable, they can vanish at any moment and therefore nothing important can live on their disk. Right now AlpinaShop has 60 GB of product images on the local disk of a rented physical server, with no real backup, served by the same process that renders the catalogue. This lesson solves that problem.

Cloud Storage is Google Cloud's object storage: a service where you keep files — images, videos, backups, exports, logs — with extreme durability, without managing disks or servers, and reachable over HTTPS from anywhere. It is probably the most cross-cutting service on the whole platform: it will turn up again in Cloud SQL backups (02-03), as a data source for BigQuery (04-01), as an artefact store in pipelines (06-01) and as the CDN's origin (03-03).

In this lesson you will understand the object model and why "directories" are an illusion; you will choose a location and a storage class on cost grounds; you will actually migrate the 60 GB to the alpinashop-catalogo bucket; you will secure access with uniform IAM and signed URLs; you will automate savings with lifecycle rules; and you will connect the Flask application to Storage in order to upload the image of a new product.

Contents

  1. The object model: buckets, objects and the illusion of folders
  2. Bucket names and location
  3. Storage classes and their real cost
  4. Migrating AlpinaShop's 60 GB
  5. Access control: uniform IAM, ACLs and public buckets
  6. Signed URLs for private content
  7. Object versioning
  8. Lifecycle rules
  9. Retention and object locking
  10. Metadata and Cache-Control
  11. Notifications to Pub/Sub
  12. Access from the Flask application with the Python library
  13. Egress cost and design tips

  1. The object model: buckets, objects and the illusion of folders

Cloud Storage is not a file system. It has exactly two levels:

  • Bucket: the container. Created once, it has a name that is unique across all of Google Cloud, a fixed location and a default class.
  • Object: the file. It has a name (which may contain slashes), binary contents and some metadata.

And nothing else. Directories do not exist. When you see this in the console:

gs://alpinashop-catalogo/
  mochilas/
    moc-40-frontal.jpg
    moc-40-lateral.jpg
  botas/
    bot-gtx-frontal.jpg

what is really there is three objects whose full names are mochilas/moc-40-frontal.jpg, mochilas/moc-40-lateral.jpg and botas/bot-gtx-frontal.jpg. The / slash is just another character in the name; the console and the tools interpret it as a separator and simulate a hierarchy using prefixes.

Very real practical consequences:

  • Renaming a "folder" with many objects is not an atomic operation: it means copying and deleting every object.
  • An empty folder does not exist. If you delete every object with the mochilas/ prefix, the folder disappears from view.
  • Listing is an operation by prefix. Listing gs://bucket/mochilas/ walks the objects starting with that string; in buckets with millions of objects, the choice of prefix matters for performance.
  • Your naming design is your schema. A good prefix convention replaces the folder structure.

For AlpinaShop we settle on this convention, which we will use throughout the course:

gs://alpinashop-catalogo/
  productos/<sku>/original/<fichero>.jpg     # original image uploaded by the team
  productos/<sku>/web/<fichero>-800.webp     # version optimised for the shop
  productos/<sku>/thumb/<fichero>-200.webp   # thumbnail for listings
  estatico/css/, estatico/js/                # website assets
  exportaciones/<yyyy>/<mm>/<dd>/            # files for analytics (module 4)

The date prefix in exportaciones/ is not accidental: it makes lifecycle rules and partitioned loads into BigQuery easier.

Other important properties of the model:

  • Objects are immutable. You do not modify an object: you replace it with another one of the same name. There is no such thing as "writing at byte 500".
  • Strong consistency. After a successful write, any subsequent read returns the new version; listing is consistent too. This was not always the case in other clouds, and it eliminates a whole category of bugs.
  • Advertised durability of eleven nines (99.999999999 %) a year. Do not confuse durability with availability, or with protection against human error: if you delete an object, it is deleted. That is what versioning is for (section 7).

  1. Bucket names and location

The bucket name is globally unique, shared with every Google Cloud customer in the world. If catalogo is taken — it is — you will have to choose another one. The rules:

  • Between 3 and 63 characters, lower case, digits, hyphens, underscores and dots.
  • It cannot begin with goog or resemble google.
  • It cannot look like an IP address.
  • The name is visible to whoever receives a URL: do not put sensitive information in it.

Since the namespace is global, the usual convention is to prefix it with the name of the organization or the project. AlpinaShop will use alpinashop-catalogo; if that were taken, alpinashop-catalogo-prod or similar.

The location is immutable: it is decided when the bucket is created and cannot be changed. To move it you have to create another one and copy. Three types:

Type Example Replicas Typical availability Storage cost When to use it
Regional europe-west1 Several zones of one region High The lowest Data served from that region; compute in the same region (no egress cost between services)
Dual-region eur4 (Netherlands + Finland), or custom pairs Two specific regions Very high Intermediate You need to survive the loss of a whole region with control over where the data sits
Multi-region EU, US, ASIA Several regions of the geographic area Maximum The highest Content served at continental scale, CDN origins

The criterion for AlpinaShop: the images are served to customers in Spain and Portugal, and the compute (the MIG's VMs, App Engine, GKE) is in europe-west1. A regional bucket in europe-west1 is the right choice: minimum cost, minimum latency towards the compute and zero egress charges between services in the same region. Geographic distribution towards the end customer will be handled with Cloud CDN (03-03), not by paying for a multi-region bucket.

gcloud storage buckets create gs://alpinashop-catalogo \
  --project=alpinashop-prod \
  --location=europe-west1 \
  --default-storage-class=STANDARD \
  --uniform-bucket-level-access \
  --public-access-prevention

The last two flags deserve attention right away: --uniform-bucket-level-access disables per-object ACLs and lets all access be governed by IAM (section 5), and --public-access-prevention stops anyone — even by mistake — making the bucket public. Both are security good practices that are much harder to turn on later than at the start.

  1. Storage classes and their real cost

All objects are stored with the same durability and the same access latency (milliseconds, even in Archive). What changes between classes is the balance between what you pay to keep them and what you pay to read them.

Class Minimum duration Storage cost Retrieval cost Use case
Standard None The highest (≈$0.020/GB/month in Europe) None Frequently accessed data: shop images, static website
Nearline 30 days ≈50 % of Standard Yes, per GB read Monthly access: recent copies, historical files that are still consulted
Coldline 90 days ≈25 % of Standard Higher than Nearline Quarterly access: backups, warm archive
Archive 365 days ≈10 % of Standard The highest Annual access or never: legal retention, definitive archive

The amounts are orders of magnitude as of 2026 for reasoning about proportions, not price lists. Always check the official Cloud Storage pricing page and the Google Cloud calculator.

The minimum duration is the most frequent trap. If you store an object in Coldline and delete it after 10 days, you are billed just the same as if it had been there for 90 days. And if you read it several times, the retrieval cost can far exceed the storage saving.

Rule of thumb: the saving from a cold class is only real if the object is almost never read and is going to stay for a long time.

An indicative calculation for AlpinaShop. Its 60 GB of images, with the thumbnails and web versions that will be generated, will come to around 90 GB:

Scenario Storage/month Comment
90 GB in Standard ≈$1.80 Trivial compared with the rest of the bill
90 GB in Nearline ≈$0.90 A saving of $0.90 that evaporates if the CDN fails and things have to be re-read
500 GB of historical copies in Coldline ≈$2.50 Here it does pay off

Conclusion: for the catalogue's live images, Standard. Cold classes will be reserved for old versions and exports, through lifecycle rules (section 8). Optimising 90 cents a month while paying for oversized instances is a poor use of your time; in 07-05 we will systematise where the real money is.

There is also Autoclass, which automatically moves each object between classes according to its actual access pattern, with no retrieval cost for the transitions. It is the sensible option when you do not know how the data will be accessed:

gcloud storage buckets update gs://alpinashop-catalogo --enable-autoclass

  1. Migrating AlpinaShop's 60 GB

Marta has the images in /var/www/imagenes on the physical server. The goal is to get them into gs://alpinashop-catalogo/productos/ without interrupting the current shop.

Step 1: install the CLI and authenticate on the source server. The physical server is not a Google VM, so it needs credentials. The right approach is a service account with write-only permission on that bucket:

# In Cloud Shell: create the service account for the migration
gcloud iam service-accounts create sa-migracion-catalogo \
  --display-name="Initial image migration"

gcloud storage buckets add-iam-policy-binding gs://alpinashop-catalogo \
  --member="serviceAccount:[email protected]" \
  --role="roles/storage.objectCreator"

# Generate the key (JSON file) that will be copied to the physical server
gcloud iam service-accounts keys create ~/sa-migracion.json \
  --iam-account="[email protected]"

Two important warnings. First: the objectCreator role allows objects to be created but not read or deleted; if that key leaked, the damage would be limited. Second: service account JSON keys are long-lived credentials and the most dangerous asset in a project. Delete it as soon as the migration finishes (gcloud iam service-accounts keys delete). In 03-04 and 07-04 you will see keyless alternatives (Workload Identity Federation).

# On the physical server
gcloud auth activate-service-account --key-file=/root/sa-migracion.json
gcloud config set project alpinashop-prod

Step 2: first bulk copy. gcloud storage cp with --recursive uploads the whole tree. The modern version of the CLI parallelises automatically:

gcloud storage cp --recursive /var/www/imagenes/* \
  gs://alpinashop-catalogo/productos/

If you want to control the parallelism (for example, so as not to saturate the office connection during working hours):

gcloud config set storage/process_count 8
gcloud config set storage/thread_count 4

# And for large files, parallel composite upload
gcloud config set storage/parallel_composite_upload_threshold 150M

Parallel composite upload splits a large file into chunks that are uploaded simultaneously and then reassembled on the server. It speeds up files of hundreds of MB a great deal, but it produces objects with a different kind of checksum (composite CRC32C), which can confuse older tools. For images of a few MB, like AlpinaShop's, it does not apply.

Step 3: incremental synchronisation. On cut-over day you do not want to upload 60 GB again, only whatever has changed. That is what rsync is for:

gcloud storage rsync --recursive --delete-unmatched-destination-objects \
  /var/www/imagenes gs://alpinashop-catalogo/productos

Careful with --delete-unmatched-destination-objects: it deletes everything at the destination that is not in the source. It is what you want for an exact mirror, and a catastrophe if you point the source at the wrong place. Before running it for real, use the simulation:

gcloud storage rsync --recursive --dry-run \
  /var/www/imagenes gs://alpinashop-catalogo/productos

Step 4: verify. Check the number of objects and the total size:

# Number of objects
gcloud storage ls --recursive "gs://alpinashop-catalogo/productos/**" | wc -l

# Total size
gcloud storage du --summarize --readable-sizes gs://alpinashop-catalogo/productos

When not to use the CLI. If you had tens of TB or the data were in another cloud, the right tool would be Storage Transfer Service, which handles retries, verification and scheduling without tying up your server; and for petabytes without the bandwidth, Transfer Appliance, a physical device Google sends you. For 60 GB over a decent connection, gcloud storage is more than enough: at 100 Mbps it is around 90 minutes.

  1. Access control: uniform IAM, ACLs and public buckets

There are two historical access control mechanisms, and it is worth being crystal clear about which to use:

ACL (per-object access control) Uniform bucket-level IAM
Granularity Per individual object Per bucket (and per object via IAM conditions)
Where it is defined In each object's metadata In the IAM policy of the bucket or the project
Auditability Hard: you have to inspect object by object Simple: a single policy to read
Current recommendation Avoid Always use

With uniform bucket-level access enabled, ACLs stop having any effect and everything is decided with IAM. That is what we did when creating the bucket, and it has been Google's firm recommendation for years: the reason bucket data historically leaked was not a lack of controls, but that nobody was able to audit millions of individual ACLs.

The most used predefined roles:

Role Allows
roles/storage.objectViewer Read and list objects
roles/storage.objectCreator Create objects (not read or delete)
roles/storage.objectUser Read, create, delete objects (not configure the bucket)
roles/storage.objectAdmin Full control over the objects
roles/storage.admin Full control over the bucket and its objects
# The web application only needs to read images
gcloud storage buckets add-iam-policy-binding gs://alpinashop-catalogo \
  --member="serviceAccount:[email protected]" \
  --role="roles/storage.objectViewer"

# Lucia needs to read the exports, but only those: IAM condition by prefix
gcloud storage buckets add-iam-policy-binding gs://alpinashop-catalogo \
  --member="user:[email protected]" \
  --role="roles/storage.objectViewer" \
  --condition="title=solo-exportaciones,expression=resource.name.startsWith('projects/_/buckets/alpinashop-catalogo/objects/exportaciones/')"

That IAM condition is a good example of least privilege applied precisely: Lucía can read exportaciones/ and nothing else, with no need for a separate bucket. The full mechanism of IAM conditions is studied in 03-04.

Public buckets. Making a bucket readable by the whole internet is as easy as granting objectViewer to allUsers:

# Do NOT do this on the catalogue bucket
gcloud storage buckets add-iam-policy-binding gs://mi-bucket \
  --member=allUsers --role=roles/storage.objectViewer

Real risks:

  • What is public is public forever. A URL indexed by a search engine or cached by a third party cannot be withdrawn.
  • You pay everyone's egress. If someone links your images from a busy forum, the bill is yours.
  • A prefix mistake exposes the whole bucket, not just what you intended.

When it is reasonable: genuinely public static assets (CSS, JS, the logo) in a separate bucket created explicitly for that. AlpinaShop's product images, even though they are not secret, will be served through the CDN with the bucket as a private origin (03-03), not by opening it to the internet. That is why we created the bucket with --public-access-prevention.

  1. Signed URLs for private content

So how does AlpinaShop let a customer download their invoice as a PDF, or let Dani share a file with a supplier, without making anything public and without giving them a Google account?

With a signed URL: a temporary link that embeds a cryptographic signature and works for a limited time. Whoever holds it can get access; when it expires, it stops working.

# Link valid for 15 minutes for a private object
gcloud storage sign-url gs://alpinashop-catalogo/facturas/2026/F-10234.pdf \
  --duration=15m \
  --impersonate-service-account=sa-catalogo-web@alpinashop-prod.iam.gserviceaccount.com

And from the application, which is where it is really used:

from datetime import timedelta
from google.cloud import storage

client = storage.Client()
bucket = client.bucket("alpinashop-catalogo")
blob = bucket.blob("facturas/2026/F-10234.pdf")

url = blob.generate_signed_url(
    version="v4",
    expiration=timedelta(minutes=15),
    method="GET",
    response_disposition='attachment; filename="invoice-F-10234.pdf"',
)
print(url)

Details that matter:

  • version="v4" is the current signing algorithm; always use this one.
  • response_disposition forces the download with a friendly name, instead of opening the PDF in the browser under the object's internal name.
  • The expiry must be short. A signed URL valid for 7 days is practically a public link: it gets forwarded by email and survives.
  • It also works for uploading (method="PUT"): the user's browser uploads straight to Cloud Storage without going through your server, which saves bandwidth and processing time. It is the recommended pattern for large uploads.
  • Signing requires a key. If the application runs on a VM or on Cloud Run with a service account and no key file, it needs the iam.serviceAccounts.signBlob permission (the Service Account Token Creator role) in order to sign through the IAM API.

  1. Object versioning

Objects are immutable, but their names are not: if you upload another object with the same name, the previous one is lost. And if someone runs a badly aimed rsync, a great deal more is lost.

Versioning makes every overwrite or deletion keep the previous version as a noncurrent version, identified by a generation number:

gcloud storage buckets update gs://alpinashop-catalogo --versioning

# List all versions, including noncurrent ones
gcloud storage ls --all-versions gs://alpinashop-catalogo/productos/MOC-40/

# Restore a specific version by its generation number
gcloud storage cp \
  gs://alpinashop-catalogo/productos/MOC-40/web/frontal-800.webp#1754400000000000 \
  gs://alpinashop-catalogo/productos/MOC-40/web/frontal-800.webp

The #<generation> suffix identifies a specific version. Restoring simply means copying it over the current one.

Versioning is the only real protection against human error — accidental deletion, the wrong script, ransomware — but it has a cost: you pay for all the versions. Without a lifecycle rule that cleans up the old ones, the bucket grows indefinitely. They always go together.

  1. Lifecycle rules

A lifecycle rule is a policy that Cloud Storage applies automatically to objects meeting certain conditions. It is defined in JSON:

{
  "lifecycle": {
    "rule": [
      {
        "action": { "type": "SetStorageClass", "storageClass": "NEARLINE" },
        "condition": {
          "age": 30,
          "matchesPrefix": ["exportaciones/"],
          "matchesStorageClass": ["STANDARD"]
        }
      },
      {
        "action": { "type": "SetStorageClass", "storageClass": "COLDLINE" },
        "condition": {
          "age": 180,
          "matchesPrefix": ["exportaciones/"],
          "matchesStorageClass": ["NEARLINE"]
        }
      },
      {
        "action": { "type": "Delete" },
        "condition": {
          "daysSinceNoncurrentTime": 30,
          "numNewerVersions": 3
        }
      },
      {
        "action": { "type": "AbortIncompleteMultipartUpload" },
        "condition": { "age": 7 }
      }
    ]
  }
}

Rule by rule:

  1. At 30 days, exports move to Nearline. Lucía's reports are consulted a lot in the first week and hardly ever afterwards. matchesStorageClass stops the rule reprocessing objects that have already been moved.
  2. At 180 days they move to Coldline. They are kept in case a history needs rebuilding, at a quarter of the cost.
  3. Noncurrent versions are deleted 30 days after they stop being current, provided at least 3 newer versions exist. The combination of the two conditions is deliberate: it guarantees a recovery time window and a minimum number of versions kept.
  4. Incomplete multipart uploads are aborted after 7 days. An interrupted large upload leaves fragments that take up space and are billed without appearing when you list the bucket. It is a classic phantom cost; this rule should be in all your buckets.

Applying and checking:

gcloud storage buckets update gs://alpinashop-catalogo \
  --lifecycle-file=ciclo-vida.json

gcloud storage buckets describe gs://alpinashop-catalogo \
  --format="json(lifecycle)"

Warnings about the behaviour: the rules are evaluated once a day asynchronously, so do not expect an immediate effect; and transitions to cold classes restart the minimum duration, so moving something to Coldline when you are going to delete it in a month is counterproductive.

Notice that we apply no transition at all to the catalogue's live images: they are read constantly and must stay in Standard.

  1. Retention and object locking

When the law or a contract requires data to be kept for a period, versioning is not enough: anyone with sufficient permissions can delete the lot. That is what two stronger mechanisms are for.

Bucket retention policy: no object can be deleted or overwritten until it reaches the stated age.

# Retain invoices for 6 years (a common commercial obligation in Spain)
gcloud storage buckets update gs://alpinashop-facturas \
  --retention-period=6y

Locking the policy (lock): it makes retention irreversible. Neither the project owner, nor the organization administrator, nor Google support can reduce or remove it. The period can only be increased.

gcloud storage buckets update gs://alpinashop-facturas --lock-retention-period

It is one of those operations that admit no second thoughts: if you lock 6 years, that bucket cannot be emptied for 6 years, and nor can it be deleted while it contains objects. That is exactly its purpose (regulatory compliance, protection against an attacker with administrator credentials), but it demands complete certainty.

In addition, retention on individual objects lets you mark specific objects with a retention date, and event-based holds and temporary holds freeze an object indefinitely until the hold is released (useful, for example, in the face of litigation).

  1. Metadata and Cache-Control

Every object has metadata, some of which turn directly into HTTP headers when it is served:

Metadata HTTP header What for
Content-Type Content-Type So the browser knows what it is. If it is wrong, the browser downloads the image instead of displaying it
Cache-Control Cache-Control How long it can be cached, in the browser and in the CDN
Content-Encoding Content-Encoding Compressed content (gzip)
Content-Disposition Content-Disposition Force a download and a file name
Custom metadata x-goog-meta-* Your own labels: SKU, photographer, shoot date
# Product images: cache for a year, they are immutable by design
gcloud storage objects update "gs://alpinashop-catalogo/productos/**/web/*.webp" \
  --cache-control="public, max-age=31536000, immutable"

# The JSON catalogue changes daily: cache briefly and revalidate
gcloud storage objects update gs://alpinashop-catalogo/estatico/catalogo.json \
  --cache-control="public, max-age=300, must-revalidate"

# Custom metadata with the SKU
gcloud storage objects update gs://alpinashop-catalogo/productos/MOC-40/web/frontal-800.webp \
  --custom-metadata=sku=MOC-40,fotografo=estudio-norte

Why Cache-Control matters so much. When we put Cloud CDN in front of the bucket (03-03), this header is what decides how long the CDN keeps each object at the points of presence. A high max-age means most requests are served from the edge of the network: faster for the customer and far cheaper, because they never leave the bucket.

The problem is what happens when an image changes. That is why the correct technique is to name objects immutably: if the file includes a hash or a version (frontal-800.a3f9c1.webp), it is never overwritten, and publishing a new image means publishing a new name. That way you can cache for a year without fear. It is the same idea you already used without realising it when versioning the instance templates in the previous lesson.

By default, Cloud Storage objects are served with Cache-Control: public, max-age=3600, which is rarely what you want. Set it explicitly when uploading.

  1. Notifications to Pub/Sub

Cloud Storage can publish a message to a Pub/Sub topic every time an object is created, deleted, archived or has its metadata changed. It is the basis of event-driven architectures:

gcloud storage buckets notifications create gs://alpinashop-catalogo \
  --topic=imagenes-subidas \
  --event-types=OBJECT_FINALIZE \
  --object-prefix=productos/ \
  --payload-format=json

The natural use case for AlpinaShop: when the team uploads an original image to productos/<sku>/original/, an event is published, and a consumer automatically generates the web version and the thumbnail. The OBJECT_FINALIZE event fires when the write has completed, not when it starts.

We will not develop it here: Pub/Sub has a lesson of its own (04-04) and the natural consumer would be a Cloud Function (06-03). Hold on to the idea that the bucket is not a passive store but a source of events.

  1. Access from the Flask application with the Python library

The moment has come to connect AlpinaShop's catalogue to the bucket. Installation:

pip install google-cloud-storage

About authentication: the code carries no passwords and no paths to key files. The library uses the Application Default Credentials we saw in 01-06, and resolves them in this order: the GOOGLE_APPLICATION_CREDENTIALS variable if it exists, then the credentials from gcloud auth application-default login on your laptop, and finally — when the code runs on a VM, on GKE or on Cloud Run — the attached service account, obtained from the metadata server. The same code works locally and in production with no changes. This is one of the great strengths of the Google Cloud model.

import os
import uuid
from flask import Flask, request, redirect, render_template_string
from google.cloud import storage
from werkzeug.utils import secure_filename

app = Flask(__name__)

BUCKET = os.environ.get("BUCKET_CATALOGO", "alpinashop-catalogo")
ALLOWED_EXTENSIONS = {".jpg", ".jpeg", ".png", ".webp"}

# The client is created ONCE at start-up, not on every request:
# it opens reusable connections and resolves credentials, which is expensive.
client = storage.Client()
bucket = client.bucket(BUCKET)


@app.route("/producto/<sku>/imagen", methods=["POST"])
def upload_image(sku):
    uploaded = request.files.get("imagen")
    if uploaded is None or uploaded.filename == "":
        return "Missing file", 400

    name = secure_filename(uploaded.filename)
    extension = os.path.splitext(name)[1].lower()
    if extension not in ALLOWED_EXTENSIONS:
        return f"Extension not allowed: {extension}", 400

    # Immutable name: never overwritten, can be cached for a year
    destination = f"productos/{sku}/original/{uuid.uuid4().hex}{extension}"
    blob = bucket.blob(destination)

    blob.cache_control = "public, max-age=31536000, immutable"
    blob.metadata = {"sku": sku, "subido-por": "panel-interno"}

    # upload_from_file takes the stream directly: nothing is written to local disk,
    # which is essential on ephemeral instances like the MIG's in lesson 02-01.
    blob.upload_from_file(uploaded.stream, content_type=uploaded.mimetype)

    return {"objeto": destination, "uri": f"gs://{BUCKET}/{destination}"}, 201


@app.route("/producto/<sku>/imagenes")
def list_images(sku):
    # list_blobs with a prefix: the correct way to "list a folder"
    blobs = client.list_blobs(BUCKET, prefix=f"productos/{sku}/web/")
    urls = [
        b.generate_signed_url(version="v4", expiration=900, method="GET")
        for b in blobs
    ]
    html = "".join(f'<img src="{u}" width="200">' for u in urls)
    return render_template_string(f"<h2>{sku}</h2>{html}")

Teaching points from the code:

  • storage.Client() outside the view. Creating it on every request is a common performance mistake: it means resolving credentials and opening new connections every time.
  • Extension validation and secure_filename. Never build the object name directly from what the user sends: it could contain ../ or unexpected characters.
  • upload_from_file with the stream. It uploads in streaming without touching the instance's disk, which can disappear at any moment.
  • A name with a UUID. It avoids collisions and makes the object immutable, enabling the aggressive caching from section 10.
  • Signed URLs when listing. The bucket is private; the images are shown with temporary links. In production this will be replaced by the CDN with a private origin (03-03).

Downloading an object is just as direct:

blob = bucket.blob("exportaciones/2026/08/pedidos.csv")

content = blob.download_as_text()             # into memory, as text
blob.download_to_filename("/tmp/pedidos.csv")  # to disk

if blob.exists():
    blob.reload()  # refreshes the metadata from the server
    print(blob.size, blob.updated, blob.content_type, blob.storage_class)

  1. Egress cost and design tips

Storage in Cloud Storage is cheap. What surprises people on the bill is almost always the egress: the traffic leaving Google Cloud towards the internet.

Type of traffic Indicative cost
Ingress (uploading data to GCP) Free
Bucket → GCP service in the same region Free
Bucket → GCP service in another region of the same continent Low, per GB
Bucket → internet (Europe/North America) ≈$0.08–0.12/GB, with volume discounts
Bucket → internet through Cloud CDN Lower than direct egress, and many requests never even reach the bucket
Operations (class A: writes and listings; class B: reads) Cents per 10,000, but relevant with millions of small objects

An indicative example for AlpinaShop: if the shop serves 200 GB of images a month straight from the bucket, the egress comes to around $16–24/month, against less than $2 of storage. In other words, moving the data costs ten times more than keeping it. That is why the CDN is not a performance luxury but a cost measure too.

Design tips that follow from all of the above:

  • Put the compute in the same region as the bucket. europe-west1 in both cases: zero transfer cost between services.
  • Serve the images through the CDN with a long Cache-Control and immutable names.
  • Optimise the weight before uploading. Converting to WebP and resizing reduces egress proportionally; it is the most profitable cost optimisation in this lesson.
  • Watch out for very small objects. Millions of tiny files pay more in operations than in storage; sometimes it is worth grouping them.
  • Always enable the rule that aborts incomplete multipart uploads.
  • Label the buckets with entorno, equipo, centro-coste and aplicacion so that spend can be apportioned (01-04).

Common Mistakes and Tips

  • Thinking in folders. They do not exist. Renaming a prefix with 50,000 objects means copying and deleting 50,000 objects.
  • Choosing the location badly. It is immutable. An EU multi-region bucket to serve Spanish customers costs more without adding anything over europe-west1 + CDN.
  • Putting everything in Coldline "to save money". If the objects are read, the retrieval cost and the minimum durations work out dearer than Standard.
  • Enabling versioning without a lifecycle rule. The bucket grows forever and nobody notices until the bill.
  • Making a bucket public "for a quick test". What is public gets indexed, and turning off access does not delete what has already been copied.
  • Using per-object ACLs. Impossible to audit. Enable uniform access and work with IAM.
  • Signed URLs with an expiry of days. They amount to public links: they get forwarded.
  • Forgetting Cache-Control. The default of 1 hour ruins the CDN's effectiveness.
  • Creating the Storage client on every Flask request. An unnecessary latency penalty on every request.
  • Tip: use --dry-run before any rsync with deletion.
  • Tip: name objects immutably (a hash or a UUID). It simplifies caching, versioning and publishing.
  • Tip: delete service account JSON keys as soon as the one-off task that justified them is over.
  • Tip: review spend per bucket with the billing export to BigQuery you configured in 01-04.

Exercises

Exercise 1: creating and organising the catalogue bucket

  1. Create a bucket with a unique name (alpinashop-catalogo-<yourinitials>) in europe-west1, Standard class, uniform access and public access prevention.
  2. Create locally a structure simulating three products with one image each and upload it respecting the productos/<sku>/original/ convention.
  3. Check that no real "folders" exist: list the objects with their full name.
  4. Work out the total size and the number of objects.
  5. Apply a one-year Cache-Control to every image under productos/.

Exercise 2: versioning, lifecycle and recovery

  1. Enable versioning on your bucket.
  2. Overwrite one of the images with different content and list all the versions.
  3. Restore the original version by its generation number.
  4. Write and apply a lifecycle file that: moves to Nearline anything under exportaciones/ older than 30 days, deletes noncurrent versions older than 15 days while keeping at least 2 newer ones, and aborts incomplete multipart uploads after 7 days.
  5. Verify the configuration applied.

Exercise 3: private access and the Flask application

  1. Create a service account sa-tienda-lectura with read-only object permission on your bucket.
  2. Generate a signed URL valid for 10 minutes for one of the images and check it with curl.
  3. Check that the direct URL (unsigned) returns 403.
  4. Write a Flask endpoint that receives a file by POST, validates the extension, uploads it with an immutable name under productos/<sku>/original/ and returns a 5-minute signed URL.
  5. Explain why the code contains no credentials.

Solutions

Solution 1

BUCKET="alpinashop-catalogo-jcm"

gcloud storage buckets create "gs://$BUCKET" \
  --location=europe-west1 \
  --default-storage-class=STANDARD \
  --uniform-bucket-level-access \
  --public-access-prevention

# 2. Local structure and upload
mkdir -p ~/catalogo/{MOC-40,BOT-GTX,TDA-2P}
for sku in MOC-40 BOT-GTX TDA-2P; do
  echo "simulated image of $sku" > ~/catalogo/$sku/frontal.jpg
  gcloud storage cp ~/catalogo/$sku/frontal.jpg \
    "gs://$BUCKET/productos/$sku/original/frontal.jpg"
done

# 3. Full names: there are no folders
gcloud storage ls --recursive "gs://$BUCKET/**"

# 4. Size and number of objects
gcloud storage du --summarize --readable-sizes "gs://$BUCKET"
gcloud storage ls --recursive "gs://$BUCKET/**" | wc -l

# 5. Cache-Control
gcloud storage objects update "gs://$BUCKET/productos/**" \
  --cache-control="public, max-age=31536000, immutable"

In point 3 you will see output like gs://<bucket>/productos/MOC-40/original/frontal.jpg: a single object whose name contains slashes. There is no "folder" entity you can describe or assign permissions to.

Solution 2

# 1. Versioning
gcloud storage buckets update "gs://$BUCKET" --versioning

# 2. Overwrite and list versions
echo "modified version" > /tmp/frontal.jpg
gcloud storage cp /tmp/frontal.jpg "gs://$BUCKET/productos/MOC-40/original/frontal.jpg"
gcloud storage ls --all-versions "gs://$BUCKET/productos/MOC-40/original/"

# 3. Restore (replace GEN with the generation number of the old version)
gcloud storage cp \
  "gs://$BUCKET/productos/MOC-40/original/frontal.jpg#GEN" \
  "gs://$BUCKET/productos/MOC-40/original/frontal.jpg"
{
  "lifecycle": {
    "rule": [
      {
        "action": { "type": "SetStorageClass", "storageClass": "NEARLINE" },
        "condition": {
          "age": 30,
          "matchesPrefix": ["exportaciones/"],
          "matchesStorageClass": ["STANDARD"]
        }
      },
      {
        "action": { "type": "Delete" },
        "condition": { "daysSinceNoncurrentTime": 15, "numNewerVersions": 2 }
      },
      {
        "action": { "type": "AbortIncompleteMultipartUpload" },
        "condition": { "age": 7 }
      }
    ]
  }
}
gcloud storage buckets update "gs://$BUCKET" --lifecycle-file=ciclo-vida.json
gcloud storage buckets describe "gs://$BUCKET" --format="json(lifecycle,versioning)"

An important note: the restore works because versioning was on before the overwrite. Enabling it afterwards recovers nothing. That is the reason to enable it the day the bucket is created, not the day the accident happens.

Solution 3

# 1. Service account with read access
gcloud iam service-accounts create sa-tienda-lectura \
  --display-name="Catalogue read access"

gcloud storage buckets add-iam-policy-binding "gs://$BUCKET" \
  --member="serviceAccount:sa-tienda-lectura@$(gcloud config get-value project).iam.gserviceaccount.com" \
  --role="roles/storage.objectViewer"

# 2. 10-minute signed URL
URL=$(gcloud storage sign-url \
  "gs://$BUCKET/productos/MOC-40/original/frontal.jpg" \
  --duration=10m --format="value(signed_url)")
curl -s -o /dev/null -w "%{http_code}\n" "$URL"   # 200

# 3. Direct URL without a signature
curl -s -o /dev/null -w "%{http_code}\n" \
  "https://storage.googleapis.com/$BUCKET/productos/MOC-40/original/frontal.jpg"  # 403
# 4. Upload endpoint returning a signed URL
import os, uuid
from datetime import timedelta
from flask import Flask, request
from google.cloud import storage
from werkzeug.utils import secure_filename

app = Flask(__name__)
client = storage.Client()
bucket = client.bucket(os.environ["BUCKET_CATALOGO"])
ALLOWED = {".jpg", ".jpeg", ".png", ".webp"}


@app.route("/producto/<sku>/imagen", methods=["POST"])
def upload(sku):
    f = request.files.get("imagen")
    if not f or not f.filename:
        return {"error": "missing file"}, 400

    ext = os.path.splitext(secure_filename(f.filename))[1].lower()
    if ext not in ALLOWED:
        return {"error": f"extension not allowed: {ext}"}, 400

    destination = f"productos/{sku}/original/{uuid.uuid4().hex}{ext}"
    blob = bucket.blob(destination)
    blob.cache_control = "public, max-age=31536000, immutable"
    blob.upload_from_file(f.stream, content_type=f.mimetype)

    url = blob.generate_signed_url(
        version="v4", expiration=timedelta(minutes=5), method="GET"
    )
    return {"objeto": destination, "url": url}, 201
  1. The code contains no credentials because the client library uses the Application Default Credentials. On Dani's laptop they are the ones from gcloud auth application-default login; on the MIG's VM, on GKE or on Cloud Run they are those of the attached service account, obtained from the metadata server and rotated automatically. The same code, with no changes and no key files, works in both places. That is the identity model we will go deeper into in 03-04.

Cleaning up when you finish:

gcloud storage rm --recursive "gs://$BUCKET"

Conclusion

AlpinaShop's 60 GB of images no longer depend on the disk of a physical server. You have understood that Cloud Storage is not a file system but a flat object store, where folders are an illusion built on top of prefixes, and you have settled on a naming convention (productos/<sku>/original|web|thumb/, exportaciones/<date>/) that will hold up the rest of the course. You know that the bucket name is global and that the location is irreversible, and you have reasoned why a regional bucket in europe-west1 is a better choice than a multi-region one for a shop that will serve its images through a CDN. You know the four storage classes, their minimum durations and the trap of the retrieval cost, and you have arrived at an honest conclusion: live images stay in Standard, and cold classes are reserved for old exports.

You have done the real migration with gcloud storage cp and rsync, with controlled parallelism, a prior simulation and subsequent verification, using a service account with creation permission only. You have secured the bucket with uniform access, public access prevention, minimal roles and an IAM condition that limits Lucía to exportaciones/; and you have learned to share private content with short-lived v4 signed URLs, both for reading and for direct upload. You have enabled versioning — the only net against human error — always accompanied by lifecycle rules that stop the bucket growing out of control, including the rule that aborts incomplete uploads. You have seen locked retention for legal obligations, you have tuned Cache-Control and you have understood why immutable names are the key to caching. And you have connected the Flask catalogue to the bucket with the Python library, without a single credential in the code.

One piece of AlpinaShop is still anchored to the physical server, and it is the most delicate of all: the tienda database, that single PostgreSQL Marta administers by hand, with no replica, no failover and with backups nobody has ever tried to restore. In 02-03, Cloud SQL, we will migrate it to the managed instance alpinashop-pedidos: we will compare which tasks Marta stops doing, create the instance with regional high availability and a read replica for Lucía's reports, configure automatic backups and point-in-time recovery, connect the Flask application securely with the Python connector and carry out the real migration with pg_dump and an import from the bucket we have just created.

Google Cloud Platform (GCP) Course

Module 1: Introduction to Google Cloud Platform

Module 2: Core GCP Services

Module 3: Networking and Security

Module 4: Data and Analytics

Module 5: Machine Learning and AI

Module 6: DevOps and Monitoring

Module 7: Advanced GCP Topics

Module 8: Final Project

© Copyright 2026. All rights reserved