Ask Marta today how she knows that the AlpinaShop shop is working. The honest answer is: because nobody has complained. And if something goes wrong, she finds out in one of three ways: a customer writes to customer support, Dani happens to open the website and notices it is slow, or Lucía sees the next day that yesterday's sales were odd.

All three have the same defect: detection is done by a person, late and by chance. And the cost of finding out that way is high and silent. A twenty-minute incident in the middle of the autumn campaign can go completely unnoticed and take hundreds of orders with it.

This lesson changes that. By the end, AlpinaShop will have a dashboard that shows the state of the shop at a glance, checks that watch it from four continents, alerts that reach Marta's phone before any customer writes, and an automatic exception grouper that warns when a deployment introduces a new error.

A note on the name first. Stackdriver was a company Google acquired in 2014 whose product gave rise to the whole operations suite. The name was retired in 2020 and today the services are called Cloud Monitoring, Cloud Logging, Cloud Trace, Cloud Profiler and Error Reporting. You will see "Stackdriver" in old documentation, in legacy library names and in conversations with people who have been around a while; it is worth recognising, but it is a historical name.

Contents

  1. Observability and its three pillars
  2. The metrics data model
  3. The metrics scope: seeing all three projects at once
  4. Platform metrics you already have for free
  5. Exploring: aligners, aggregations and why the average lies
  6. The autumn campaign dashboard
  7. The dashboard as versionable JSON
  8. MQL and PromQL for advanced queries
  9. Alerts: the anatomy of a policy
  10. Notification channels and AlpinaShop's three alerts
  11. Alert fatigue: the uncomfortable conversation
  12. Custom metrics from the application
  13. Uptime checks
  14. Error Reporting: grouped exceptions
  15. The Ops Agent on the MIG's VMs
  16. The cost of observability

  1. Observability and its three pillars

Monitoring is checking whether the things you knew could fail are failing. Observability is the property of a system that lets you understand what is going on inside it from what it emits outwards, including the failures you did not anticipate.

The difference is not academic. A dashboard with the VMs' CPU is monitoring: it answers "is the CPU high?". Being able to answer "why has checkout been slow for customers in the Canary Islands since Tuesday?" without deploying anything new is observability.

It rests on three pillars:

Pillar What it answers Nature Service Lesson
Metrics How much? How many? What trend? Numbers aggregated over time Cloud Monitoring This one
Logs What exactly happened in this case? Discrete events with detail Cloud Logging 06-06
Traces Where did the time go in this request? The journey between components Cloud Trace 06-06

Each has a different profile of cost and usefulness, and that is why they coexist:

  • Metrics are cheap to store and query because they are aggregates, they let you see months of history and they are the natural basis for alerts. But they do not tell you what happened to a specific customer: they tell you the p95 went up.
  • Logs have all the detail of every event, and that is why they are expensive and you have to be selective. They answer specific questions about specific cases.
  • Traces show how a request's time is distributed between services, and they are the only thing that answers "who is being slow?" in a system with several pieces.

The typical journey through an incident uses all three in order: an alert on a metric warns that something is happening, the metrics dashboard narrows down what and where, the logs say exactly what is failing, and the trace explains why. That complete journey is walked from start to finish in 06-06.

A note on scope: here we build metrics, dashboards and alerts. Service level objectives (SLOs) and error budgets — formally defining what "working well" means and how much failure is tolerable — arrive in 07-06. Here the instruments are installed; there it is decided which numbers are acceptable.

  1. The metrics data model

Almost every problem people have with Cloud Monitoring comes from not being clear about this model. It deserves five minutes.

A time series is the fundamental unit, and it is identified by the combination of three things:

time series = metric type + monitored resource + label values
Component What it is Example
Metric type What is measured loadbalancing.googleapis.com/https/request_count
Monitored resource Which object emits it https_lb_rule with its project and its name
Labels Dimensions for filtering and grouping response_code_class="500", matched_url_path_rule="/catalogo"
Points (instant, value) pairs (2026-08-05T10:00:00Z, 142)

That design explains something that is surprising at first: the number of time series multiplies with the labels. If a metric has 3 response code classes, 5 paths and 2 regions, that is 30 series. If somebody adds a label with the customer identifier, and there are 50,000 customers, that is 1,500,000 series. That is called a cardinality explosion, you pay for it and it can end up making the metric unusable. We will come back to it in section 12, which is where the mistake gets made.

The three metric kinds

Telling them apart is essential because each one is queried differently:

Kind What it represents Example How it is queried
Gauge An instantaneous value that goes up and down CPU at 42 %, 87 open connections Read directly
Cumulative counter A total that only grows from an origin Total requests served It has to be derived: rate per second
Distribution A histogram of values in an interval Latencies of every request Percentiles are extracted

The counter is the one that causes the most confusion. The load balancer's request_count is cumulative: its value grows indefinitely. If you plot it as it is, you see an upward ramp that says nothing. What you want is the rate: how many requests per second. That is why counters always get a rate or delta aligner applied.

The distribution is the most valuable and the most under-used. A latency metric as a distribution does not store one number per interval: it stores a complete histogram. That lets you ask for the 50th, 95th and 99th percentile after the fact, without having decided in advance which one interested you. If instead you had stored only the average latency, you would have lost the information forever, and the average is precisely the least useful statistic here, as we will see in section 5.

  1. The metrics scope: seeing all three projects at once

AlpinaShop has its resources spread across three projects: the shop in alpinashop-prod, testing in alpinashop-dev and analytics in alpinashop-datos. By default, each project has its own metrics view, which forces you to jump between three consoles to understand an incident. Unsustainable.

A metrics scope solves this: one project acts as the scoping project and aggregates the metrics of several monitored projects.

flowchart TD
    A[alpinashop-prod<br/>SCOPING PROJECT] --> D[Single metrics scope:<br/>dashboards and alerts for all three]
    B[alpinashop-dev] --> D
    C[alpinashop-datos] --> D
# Add the other two projects to the alpinashop-prod metrics scope
gcloud beta monitoring metrics-scopes create \
  --monitored-project=alpinashop-dev \
  --project=alpinashop-prod

gcloud beta monitoring metrics-scopes create \
  --monitored-project=alpinashop-datos \
  --project=alpinashop-prod

A design detail worth thinking about before running that: which project should be the scoping one? Using alpinashop-prod is the most common and the simplest. The alternative — creating a dedicated project, for example alpinashop-observabilidad — is conceptually cleaner because it separates "operating the platform" from "being the platform", and it allows read access to the metrics to be granted without granting any access to production. For AlpinaShop, with three projects and three people, alpinashop-prod is enough; in an organization with twenty projects, the dedicated project clearly wins.

And an important warning: the metrics scope unifies the view, not the permissions. Anyone consulting a dashboard that shows metrics from alpinashop-datos needs monitoring read permissions on that project. It is consistent with 03-04, and it stops aggregating the view becoming an indirect way of exposing information.

  1. Platform metrics you already have for free

Here is good news a lot of people do not know about: GCP is already collecting hundreds of metrics from your resources, without you having done anything and at no extra cost. All the infrastructure from modules 2, 3 and 4 has been emitting data for months that nobody has looked at.

The ones that matter for AlpinaShop:

Resource Metric What it indicates
Load balancer https/request_count Requests per second, with the code class
Load balancer https/total_latencies End-to-end latency (distribution)
Load balancer https/backend_latencies Backend-only latency
Load balancer https/backend_request_count Requests reaching the backend (as opposed to the CDN)
Compute Engine / MIG instance/cpu/utilization CPU saturation
Compute Engine instance/disk/read_ops_count Disk pressure
Cloud SQL database/cpu/utilization CPU of alpinashop-pedidos
Cloud SQL database/network/connections Open connections
Cloud SQL database/disk/utilization Space used
Cloud SQL database/replication/replica_lag Replica lag
Pub/Sub subscription/num_undelivered_messages Unacknowledged messages
Pub/Sub subscription/oldest_unacked_message_age Age of the oldest one
BigQuery query/scanned_bytes Bytes processed: a direct cost
GKE container/cpu/limit_utilization Usage against the limit
Cloud Functions function/execution_count Invocations (remember 06-03)
Cloud Functions function/execution_times Duration (distribution)
Cloud Run request_count, request_latencies The same as the load balancer
Cloud CDN https/request_count filtered by cache_result Hit ratio

The table that really gets used in an incident is the inverse one: from the symptom to the metric.

Observed symptom Metric to look at What it means if it is bad
"The website is slow" https/total_latencies p95 and p99 Confirms or refutes the perception
"It is slow" and the backend is fine Compare total_latencies with backend_latencies If they differ a lot, the problem is the network or the client
500 errors https/request_count filtered by response_code_class=500 How many and since when
Slowness + high CPU in the MIG instance/cpu/utilization Not enough capacity: check the autoscaling
Slowness + low CPU everywhere database/network/connections Probable exhaustion of the connection pool
Errors after a deployment request_count 5xx + Error Reporting Regression introduced by the deployment
Stale analytics data num_undelivered_messages The consumer cannot keep up or is down
High BigQuery bill query/scanned_bytes per user Somebody is running SELECT * over the history
High network bill https/request_count by cache_result The CDN hit ratio has dropped

That last row connects directly with 03-03: if Cloud CDN's hit ratio drops, traffic goes to the backend, the egress bill goes up and latency gets worse. It is a perfect silent failure — nothing breaks, everything degrades — and it is only detected by looking at the metric.

  1. Exploring: aligners, aggregations and why the average lies

Metrics Explorer is the interactive query tool. Its model has four steps and it is worth understanding them because they are the same ones that later show up in dashboards and alerts.

flowchart LR
    A[1. Select<br/>metric and resource] --> B[2. Filter<br/>by labels]
    B --> C[3. ALIGN<br/>within each series]
    C --> D[4. AGGREGATE<br/>across series]
    D --> E[Chart]

Aligning turns the raw points of each series into regular points at a fixed interval. Aggregating combines several series into fewer series or into a single one. The order matters and confusing them produces meaningless numbers.

Aligner What it does For which kind Example
rate Change per second Counters Requests per second
delta Difference over the interval Counters Requests in 60 s
mean Average over the interval Gauges Average CPU per minute
max / min Extremes of the interval Gauges Peak connections
percentile_95 Percentile within the interval Distributions p95 latency
Aggregation What it does Example
sum Sum across series Total requests from all the instances
mean Average across series Average CPU of the MIG
max Maximum across series The busiest instance
count How many series there are Number of live instances
Group by label One series per value Latency per path

Why the average lies

This is the most important concept in the section, and probably the one most often explained badly in the industry.

Imagine 1,000 requests to the catalogue in a minute. 950 take 100 ms and 50 take 8 seconds. The average is:

(950 × 100 ms + 50 × 8000 ms) / 1000 = 495 ms

A dashboard showing the average shows 495 ms. That looks acceptable. And yet 50 customers a minute are waiting eight seconds, which is more than enough time to go to another website. The average has hidden exactly the problem you were looking for.

The percentiles tell a different story:

Statistic Value What it really means
Average 495 ms A number that happens to nobody
p50 (median) 100 ms Half the customers see this
p95 ~100 ms 95 % are fine
p99 8,000 ms The worst 1 % suffer 8 seconds

Three practical rules worth adopting without exception:

  1. For latency, always percentiles, never the average. The p50 tells you what the typical experience is like; the p95 and the p99 tell you how bad the bad is.
  2. Never average percentiles across series. The average of ten instances' p95s is not the p95 of the whole: it is a number with no statistical meaning. Cloud Monitoring computes the percentile over the combined distribution if you ask for percentile_95 as the aligner and sum as the aggregation of distributions; doing it the other way round produces rubbish.
  3. The p99 matters more than its name suggests. With a million requests a day, the p99 is 10,000 requests. And they do not fall at random on 10,000 different customers: they tend to concentrate on those with large baskets or long histories, that is to say, the best customers.

  1. The autumn campaign dashboard

The autumn campaign is AlpinaShop's commercial peak. Marta wants a screen that answers "is the shop all right?" at a glance.

The design principle, imported from SRE practice: a dashboard must tell a story from top to bottom, starting with what the customer perceives and ending with the technical causes. If you have to think while looking at it, the dashboard is wrong.

Row Chart Metric Why it is there
1. Is traffic arriving? Requests/s https/request_count with rate and sum A collapse is as serious as a spike
1 5xx error rate request_count filtered / total The most direct indicator of pain
2. Is it fast? Latency p50/p95/p99 https/total_latencies Three lines on the same chart
2 Latency per path The same, grouped by matched_url_path_rule Isolates which part is slow
3. Is it holding up? MIG CPU instance/cpu/utilization, mean and max The average hides the hot instance
3 Active instances count of MIG series Is it scaling?
4. And the DB? Cloud SQL connections database/network/connections Common cause of slowness with no CPU
4 Cloud SQL CPU database/cpu/utilization Heavy queries
5. And the CDN? Hit ratio request_count by cache_result Cost and latency
5 Bytes served https/response_bytes_count Detects pattern changes

Two design decisions deserve an explanation.

The MIG's CPU is shown twice, mean and max. If ten instances are at 20 % and one is at 98 %, the average says 27 % — apparently healthy — while that instance is saturated and serving extremely slow requests. It is the same conceptual error as averaging latencies, applied to capacity. The average hides imbalances; the maximum gives them away.

Latency and errors are at the top, the infrastructure at the bottom. The order reflects the question that is really asked in an incident: first "are the customers suffering?" and only then "why?". A dashboard that starts with the CPU invites you to look at infrastructure metrics that may be perfectly fine while the shop is down.

  1. The dashboard as versionable JSON

A dashboard built by clicking in the console is a fragile asset: nobody knows who changed it, it cannot be replicated in another project and it disappears if somebody deletes it. Every Cloud Monitoring dashboard has a JSON representation that can be exported, versioned in Git — following 06-02 — and recreated.

# Export an existing dashboard
gcloud monitoring dashboards list --project=alpinashop-prod --format=json
gcloud monitoring dashboards describe DASHBOARD_ID \
  --project=alpinashop-prod --format=json > paneles/campana-otono.json

# Create or update from the file
gcloud monitoring dashboards create --config-from-file=paneles/campana-otono.json \
  --project=alpinashop-prod

A fragment with the two charts from row 1, which shows the structure without overwhelming you:

{
  "displayName": "AlpinaShop - Autumn campaign",
  "mosaicLayout": {
    "columns": 12,
    "tiles": [
      {
        "width": 6, "height": 4, "xPos": 0, "yPos": 0,
        "widget": {
          "title": "Requests per second",
          "xyChart": {
            "dataSets": [{
              "timeSeriesQuery": {
                "timeSeriesFilter": {
                  "filter": "metric.type=\"loadbalancing.googleapis.com/https/request_count\" resource.type=\"https_lb_rule\" resource.label.\"url_map_name\"=\"alpinashop-url-map\"",
                  "aggregation": {
                    "alignmentPeriod": "60s",
                    "perSeriesAligner": "ALIGN_RATE",
                    "crossSeriesReducer": "REDUCE_SUM"
                  }
                }
              },
              "plotType": "LINE"
            }]
          }
        }
      },
      {
        "width": 6, "height": 4, "xPos": 6, "yPos": 0,
        "widget": {
          "title": "Latency p50 / p95 / p99 (ms)",
          "xyChart": {
            "dataSets": [
              {
                "legendTemplate": "p50",
                "timeSeriesQuery": {
                  "timeSeriesFilter": {
                    "filter": "metric.type=\"loadbalancing.googleapis.com/https/total_latencies\" resource.type=\"https_lb_rule\"",
                    "aggregation": {
                      "alignmentPeriod": "60s",
                      "perSeriesAligner": "ALIGN_PERCENTILE_50",
                      "crossSeriesReducer": "REDUCE_MEAN"
                    }
                  }
                },
                "plotType": "LINE"
              },
              {
                "legendTemplate": "p99",
                "timeSeriesQuery": {
                  "timeSeriesFilter": {
                    "filter": "metric.type=\"loadbalancing.googleapis.com/https/total_latencies\" resource.type=\"https_lb_rule\"",
                    "aggregation": {
                      "alignmentPeriod": "60s",
                      "perSeriesAligner": "ALIGN_PERCENTILE_99",
                      "crossSeriesReducer": "REDUCE_MEAN"
                    }
                  }
                },
                "plotType": "LINE"
              }
            ]
          }
        }
      }
    ]
  }
}

Notice how it maps onto section 5: perSeriesAligner is the aligner — ALIGN_RATE for the counter, ALIGN_PERCENTILE_99 for the distribution — and crossSeriesReducer is the aggregation. The JSON is nothing more than the written form of what you did in the explorer.

The practical way to work is to build the dashboard in the console until it looks right, export it, save it in alpinashop-infra/paneles/ and manage it as code from then on. And when Terraform arrives in 06-07, the google_monitoring_dashboard resource consumes exactly this JSON, so the dashboard comes to be created alongside the infrastructure it monitors.

  1. MQL and PromQL for advanced queries

The filter and aggregation interface covers most cases, but there are questions it cannot express: ratios between two metrics, joins, derived calculations. For that there are two languages.

MQL (Monitoring Query Language) is Cloud Monitoring's own language, with a pipeline syntax:

fetch https_lb_rule
| metric 'loadbalancing.googleapis.com/https/request_count'
| filter resource.url_map_name == 'alpinashop-url-map'
| align rate(1m)
| every 1m
| group_by [metric.response_code_class], [requests: sum(value.request_count)]

And the case where MQL becomes essential, the error rate as a proportion:

{
  fetch https_lb_rule :: loadbalancing.googleapis.com/https/request_count
  | filter metric.response_code_class == '500'
  | align rate(1m) | every 1m | group_by [], [errors: sum(value.request_count)]
  ;
  fetch https_lb_rule :: loadbalancing.googleapis.com/https/request_count
  | align rate(1m) | every 1m | group_by [], [total: sum(value.request_count)]
}
| join
| value [error_rate: errors / total * 100]

That — dividing two different series — cannot be expressed with filters and aggregations. And it is exactly what you want in order to alert: 2 % of errors is what is relevant, not "200 errors", because 200 errors out of 100,000 requests is noise and out of 500 it is a catastrophe.

PromQL is Prometheus's language, supported in Cloud Monitoring through Managed Service for Prometheus. The same query:

sum(rate(loadbalancer_googleapis_com:https_request_count{response_code_class="500"}[1m]))
/
sum(rate(loadbalancer_googleapis_com:https_request_count[1m])) * 100
Criterion MQL PromQL
Reach GCP only Industry standard
Portability None High: Prometheus, Grafana, other clouds
GKE metrics It works Natural: it is what the applications emit
Community and examples Limited Enormous
2026 recommendation For occasional cases on GCP Preferable if you come from Kubernetes

For AlpinaShop, which has the shop on GKE and aspires not to tie itself too closely to one provider, PromQL is the more sensible bet in the medium term. MQL remains useful for one-off queries over platform metrics that have no Prometheus equivalent.

  1. Alerts: the anatomy of a policy

A dashboard is only useful if somebody looks at it. An alert works for you.

An alerting policy has four pieces, and each one answers a different question:

flowchart LR
    A[CONDITION<br/>what to measure and what threshold] --> B[DURATION<br/>how long it has to be bad]
    B --> C[AGGREGATION<br/>over which set]
    C --> D[NOTIFICATION<br/>to whom and how]
    D --> E[DOCUMENTATION<br/>what to do on receiving it]

The duration is the most underrated piece. Without it, any instantaneous spike fires an alert. A 15-second latency spike because the autoscaler started an instance is not an incident: it is the system working. An alert that fires for that teaches the team to ignore it, which is the worst thing that can happen to an alert.

The criterion for choosing the duration is honest and simple: how long does this have to be bad before it is worth waking somebody up? If the answer is "five minutes", set five minutes.

A complete policy in JSON, AlpinaShop's error rate one:

{
  "displayName": "AlpinaShop - High 5xx error rate",
  "combiner": "OR",
  "conditions": [
    {
      "displayName": "5xx errors above 2% for 5 minutes",
      "conditionMonitoringQueryLanguage": {
        "query": "{ fetch https_lb_rule :: loadbalancing.googleapis.com/https/request_count | filter metric.response_code_class == '500' | align rate(1m) | every 1m | group_by [], [e: sum(value.request_count)] ; fetch https_lb_rule :: loadbalancing.googleapis.com/https/request_count | align rate(1m) | every 1m | group_by [], [t: sum(value.request_count)] } | join | value [err_rate: e / t * 100] | condition err_rate > 2 '%'",
        "duration": "300s",
        "trigger": { "count": 1 }
      }
    }
  ],
  "alertStrategy": {
    "autoClose": "1800s"
  },
  "notificationChannels": [
    "projects/alpinashop-prod/notificationChannels/CANAL_SLACK_ALERTAS",
    "projects/alpinashop-prod/notificationChannels/CANAL_CORREO_INFRA"
  ],
  "documentation": {
    "mimeType": "text/markdown",
    "content": "## High 5xx error rate\n\n**What it means:** more than 2 % of requests to the shop are returning a server error.\n\n**Impact:** customers are seeing error pages. Orders are being lost.\n\n**First steps:**\n1. *Autumn campaign* dashboard: has the latency gone up too?\n2. Was there a deployment in the last 30 minutes? Check Cloud Build.\n3. Logs Explorer: `severity=ERROR` in the last 15 minutes (06-06).\n4. Check the Cloud SQL connections: a frequent cause.\n\n**Rollback:** promote the previous image with the pipeline from 06-01.\n\n**Escalation:** if there is no diagnosis within 15 minutes, notify Marta."
  }
}
gcloud alpha monitoring policies create \
  --policy-from-file=alertas/tasa-error-5xx.json \
  --project=alpinashop-prod

The documentation field is what separates a useful alert from an alert that generates anxiety. It arrives on somebody's phone at eleven at night; with no instructions, that person starts from scratch, sleepy and in a hurry. With the first four checks written down, they know what to look at. And the 30-minute autoClose stops an already-resolved incident leaving an incident open forever.

  1. Notification channels and AlpinaShop's three alerts

A channel defines how the alert arrives:

Channel Latency Interrupts Use at AlpinaShop
Email Minutes No Informational notices, summaries
Slack Seconds A little The #alpinashop-alertas channel, the main one
SMS Seconds Yes Only for the critical stuff
PagerDuty / Opsgenie Seconds Yes, with escalation When there are formal on-call rotas (07-06)
Webhook Seconds It depends Your own automations
Pub/Sub Seconds No Reacting with a function (06-03)
gcloud beta monitoring channels create \
  --display-name="AlpinaShop alerts Slack" \
  --type=slack \
  --channel-labels=channel_name="#alpinashop-alertas" \
  --project=alpinashop-prod

The Pub/Sub channel deserves a mention for what it enables: an alert can publish to a topic and a Cloud Function from 06-03 can react automatically — create an issue, run a diagnostic, or even take a narrowly scoped corrective action. It is the door to automated response, with the obvious caveat that an automation acting on production must be very tightly scoped.

The three alerts AlpinaShop starts with

Deliberately three. Not thirty.

Alert 1 — 5xx error rate > 2 % for 5 minutes. Channel: Slack + SMS to Marta.

Reasoning behind the threshold: the shop's normal error rate is 0.1-0.3 %, dominated by requests to URLs that no longer exist. 2 % is an order of magnitude above the noise: it does not fire from natural variation, and at that level there are customers seeing errors. It is a proportion, not an absolute value, for the reason given in section 8. Five minutes filters out deployment spikes.

Alert 2 — p95 latency > 2 seconds for 10 minutes. Channel: Slack.

Reasoning: the shop's usual p95 is around 400 ms. Two seconds is the point where perception changes from "it works fine" to "it is slow" and basket abandonment starts to rise. p95 is chosen rather than p99 because the p99 is more volatile and would generate more noise; and a long duration, 10 minutes, is chosen because transient slowness is common and does not require immediate action. No SMS: it is a problem to be dealt with during working hours.

Alert 3 — Unacknowledged messages on pedidos-nuevos > 1,000 for 15 minutes. Channel: Slack + email.

Reasoning: in normal operation the queue is practically empty because the consumers keep up. A thousand messages piled up for fifteen minutes means the consumer is down or cannot keep up, and that means there are orders that are not being processed. It is the example of an alert that detects a failure that is invisible from the outside: the website works, customers buy, everything looks fine, and the orders are piling up unprocessed.

Alert Threshold Duration Channels What it protects
5xx errors > 2 % 5 min Slack + SMS Customers seeing errors
p95 latency > 2 s 10 min Slack Degraded experience
Order queue > 1,000 msg 15 min Slack + email Failure invisible from the outside

Notice the pattern: each alert detects a different kind of failure — visible error, visible degradation and invisible failure. Ten more alerts about CPU would have added nothing, because high CPU already shows up in the latency.

  1. Alert fatigue: the uncomfortable conversation

This deserves a section of its own because it is the real problem of observability, and it is almost never said clearly.

Alert fatigue is the state in which a team receives so many alerts that it stops reacting to them. It is not a failure of discipline: it is a rational response. If nineteen out of twenty daily alerts are noise, ignoring them all is a strategy with a 95 % success rate and an enormous cost when it fails.

How you get there, almost always by the same route:

  1. Somebody sets up an alert for every available metric, "just in case".
  2. The thresholds are set by eye, without looking at historical values.
  3. There is no duration, so any spike fires.
  4. They all go to the same channel with the same urgency.
  5. Nobody reviews or retires the ones that add nothing.

The five antidotes, in order of importance:

Antidote Concrete rule
Alert on symptoms, not on causes Alert if customers are suffering, not if the CPU is at 80 %
Thresholds from historical data Look at the real value for the last 4 weeks before deciding
Always a duration No alert without a time window
Levels of urgency Does this justify waking somebody up? If not, email or Slack
Periodic review Every quarter: did it fire? was it useful? If not, retire it

The first is the most important and the most counter-intuitive. A CPU at 90 % is not a problem if the customers do not notice it: it may be the system making good use of its resources. Alerting on CPU produces false positives — high CPU with no impact — and false negatives — customers suffering with low CPU, for example because of a lock in the database. Causes are investigated with the dashboard after a symptom has raised the alarm.

And a control question that is worth more than any list, for every alert you are about to create:

If this alert fires at 3 in the morning, is there something somebody must do immediately?

If the answer is no, it is not an alert: it is a chart on a dashboard. Marta and Dani have only three alerts configured, and that is deliberately the sign of a healthy alerting system.

  1. Custom metrics from the application

Platform metrics tell you what the infrastructure is doing. They do not tell you what the business is doing. No GCP metric knows how many orders per minute come into AlpinaShop, and that is probably the most important metric of all: if orders drop to zero, something is wrong even if every technical indicator is green.

# catalogo/metricas.py
from google.cloud import monitoring_v3
import time, os

client = monitoring_v3.MetricServiceClient()
PROJECT = f"projects/{os.environ['GOOGLE_CLOUD_PROJECT']}"

def record_order(amount_eur: float, channel: str) -> None:
    """Writes a point to the custom orders metric."""
    series = monitoring_v3.TimeSeries()
    series.metric.type = "custom.googleapis.com/tienda/pedidos"

    # LABELS: few and of LOW cardinality. See the warning below.
    series.metric.labels["canal"] = channel       # web | mobile | phone
    series.metric.labels["entorno"] = os.environ.get("ENTORNO", "prod")

    series.resource.type = "generic_task"
    series.resource.labels.update({
        "project_id": os.environ["GOOGLE_CLOUD_PROJECT"],
        "location": "europe-west1",
        "namespace": "tienda",
        "job": "catalogo-web",
        "task_id": os.environ.get("HOSTNAME", "unknown"),
    })

    now = time.time()
    point = monitoring_v3.Point({
        "interval": {"end_time": {"seconds": int(now)}},
        "value": {"double_value": 1.0},
    })
    series.points = [point]
    client.create_time_series(name=PROJECT, time_series=[series])

The warning about labels is the critical part of this section. It is tempting to add series.metric.labels["id_cliente"] = customer_id. Do not. With 50,000 active customers you would create 50,000 time series for every combination of the other labels, and that is the cardinality explosion from section 2: high cost, slow queries and a useless metric.

Label Possible values Acceptable?
canal 3 Yes
entorno 2-3 Yes
categoria_producto ~20 Yes
codigo_postal ~11,000 No
id_cliente 50,000+ Never
id_pedido Unlimited Never

The rule: a metric label must have tens of possible values, not thousands. The high-cardinality stuff — the customer identifier, the order identifier — goes into the logs, which are designed for it, and is queried from there. It is one of the reasons the pillars are three and not one.

There is an alternative route that in 2026 is usually better if you already live in Kubernetes: expose a /metrics endpoint in Prometheus format and let Managed Service for Prometheus scrape it. It avoids an API call per event, it is the industry standard and it fits with PromQL. For the catalogue on GKE, it is the option I would recommend.

There are also log-based metrics: counting log entries that match a filter and turning them into a metric, without touching the application's code. It is a very powerful technique for quickly instrumenting what is already being logged, and it is developed in 06-06.

  1. Uptime checks

Everything above measures the system from the inside. Uptime checks measure it from the outside: Google sends requests from several regions of the world and checks that the shop responds.

That difference in point of view detects entire classes of failure that no internal metric sees: an expired TLS certificate, a badly propagated DNS record, an over-aggressive Cloud Armor rule blocking legitimate customers, a misconfigured load balancer or a complete regional outage.

gcloud monitoring uptime create tienda-alpinashop \
  --resource-type=uptime-url \
  --resource-labels=host=www.alpinashop.example,project_id=alpinashop-prod \
  --path=/salud \
  --port=443 \
  --protocol=https \
  --period=1 \
  --timeout=10 \
  --regions=EUROPE,USA_OREGON,ASIA_PACIFIC,SOUTH_AMERICA \
  --content-matchers-content='"estado":"ok"' \
  --project=alpinashop-prod

Four decisions worth understanding:

  • The /salud path is the same endpoint as the load balancer's hc-catalogo health check from 03-02. Reusing it makes sense, but with one important nuance: the load balancer's check decides whether an instance receives traffic and must be fast and shallow; the uptime one measures whether the service is genuinely healthy. A /salud that just returns 200 OK without checking anything will be green with the database down. The endpoint should verify at least connectivity to Cloud SQL, with a short timeout so it does not itself become a problem.
  • --content-matchers-content checks that the response contains what is expected, not just that the code is 200. A server that returns an error page with code 200 — more common than it sounds — is detected this way.
  • Four spread-out regions distinguish a global problem from a regional one. If only the check from Asia fails, the problem is network or DNS, not the application.
  • A 1-minute period is the reasonable balance between fast detection and cost.

And the associated alert, with one deliberate detail:

gcloud alpha monitoring policies create \
  --notification-channels=CANAL_SMS_MARTA \
  --display-name="AlpinaShop - Shop not responding" \
  --condition-display-name="Uptime check failing in 2+ regions" \
  --condition-filter='metric.type="monitoring.googleapis.com/uptime_check/check_passed" resource.type="uptime_url"' \
  --duration=180s \
  --project=alpinashop-prod

The condition requires failure from two regions or more. A failure from a single region is usually a problem with the intermediate network, not with the shop, and alerting on it generates exactly the noise from section 11. Requiring two regions makes this the most reliable alert in the whole system: if the shop does not respond from two continents, the shop is down. It is AlpinaShop's only alert that goes straight to SMS without passing through Slack.

  1. Error Reporting: grouped exceptions

When the Flask catalogue raises an uncaught exception, the stack trace ends up in the logs. With thousands of requests a day, finding a new error in the logs is looking for a needle in a haystack.

Error Reporting solves exactly that: it collects the exceptions, groups them by their signature — the exception type and the stack, not the literal message — and presents a list of distinct problems with their frequency, their first occurrence and their last.

The practical difference: instead of 4,000 log lines, you see "KeyError: 'talla' in catalogo/carrito.py:87, 340 occurrences, first seen 2 hours ago, affecting 89 users". And that "first seen 2 hours ago" is pure gold, because it usually coincides with a deployment.

Instrumenting Flask is straightforward:

# catalogo/app.py
import google.cloud.logging
from google.cloud.error_reporting import Client as ErrorClient

logging_client = google.cloud.logging.Client()
logging_client.setup_logging()           # Python logs go to Cloud Logging
error_client = ErrorClient(service="catalogo-web", version=os.environ["VERSION_IMAGEN"])

@app.errorhandler(Exception)
def handle_error(e):
    # report_exception captures the full traceback from the current context
    error_client.report_exception()
    app.logger.exception("Unhandled error while serving %s", request.path)
    return render_template("error.html"), 500

The version parameter is more important than it looks: by passing the SHA of the deployed image, Error Reporting knows in which version each error appeared. That connects directly with 06-01 and turns a difficult question — "is this error new?" — into a fact. With that information, Error Reporting notifies regressions: it warns when an error type that did not exist appears, and when one that had been marked as resolved comes back.

Capability What it provides
Grouping by signature 4,000 log lines → 12 distinct problems
Count and trend Is it getting worse?
First and last occurrence Correlation with deployments
Affected users Prioritises by real impact, not by volume
Notification of new errors Detects regressions automatically
Link to the source code Jumps to the exact line (06-02 integration)

An important piece of advice about exception design: Error Reporting groups by the signature of the stack, so a generic error repeated in twenty different places groups badly. Specific exceptions and messages with context — but with no personal data — make the grouping useful.

  1. The Ops Agent on the MIG's VMs

There is a gap in the platform metrics that surprises everybody the first time: Google cannot see inside your virtual machines. It knows the CPU, the network and the disk operations, because the hypervisor measures those. It does not know the memory used, the free space on the disk or the processes running, because those can only be seen from inside the operating system.

The Ops Agent is a single agent that collects system metrics and sends logs from the VMs. It replaces the old separate Stackdriver monitoring and logging agents.

Installing it on the templates of the alpinashop-web-mig MIG, via the startup script from 02-01:

#!/bin/bash
# Fragment of the instance template's startup script
curl -sSO https://dl.google.com/cloudagents/add-google-cloud-ops-agent-repo.sh
sudo bash add-google-cloud-ops-agent-repo.sh --also-install

And its configuration, in /etc/google-cloud-ops-agent/config.yaml:

logging:
  receivers:
    catalogo_app:
      type: files
      include_paths: [/var/log/alpinashop/catalogo.log]
  service:
    pipelines:
      catalogo:
        receivers: [catalogo_app]

metrics:
  receivers:
    hostmetrics:
      type: hostmetrics
      collection_interval: 60s
  service:
    pipelines:
      default:
        receivers: [hostmetrics]

What shows up in Cloud Monitoring once it is installed:

Metric Why it matters
agent.googleapis.com/memory/percent_used The most common cause of unexplained restarts
agent.googleapis.com/disk/percent_used A full disk brings the application down silently
agent.googleapis.com/processes/count_by_state Zombie processes, process leaks
agent.googleapis.com/swap/percent_used If swap is active, memory is short

Memory deserves emphasis: a memory leak in the application ends with the system killing the process. From the outside it looks like random restarts and lost requests, with no metric to explain it. Without the agent, that diagnosis is practically impossible.

And a note on the future: when AlpinaShop moves the catalogue to Cloud Run (DA-001, 07-02), the agent is no longer needed because the platform exposes those metrics itself. It is a real advantage of managed services that rarely gets mentioned.

  1. The cost of observability

It is easy for the observability bill to grow without anybody noticing. It is worth knowing the model — always with the current prices from the official documentation:

Component You pay for Free tier Risk
Platform metrics Nothing All of them None
Custom metrics Time series ingested An initial volume High: cardinality
Logs GB ingested A generous monthly volume The biggest of all (06-06)
Traces Spans ingested A monthly volume Medium: controlled with sampling
Uptime checks Executions Ample Low
Dashboards and alerts Nothing — None

The fact that dashboards, alerts and platform metrics are free is the best news in this lesson: everything built in sections 4 to 11 adds no cost. What you pay for is what you ingest: custom metrics and, above all, logs.

The five tips for keeping it from running away:

  1. Watch the cardinality of your custom metrics. It is trap number one: a customer identifier as a label multiplies the cost by thousands.
  2. Do not invent metrics that already exist. Before instrumenting, check whether GCP already emits it. Many custom metrics duplicate free metrics.
  3. Exclude the noise from the logs. Health probes and static assets generate an enormous volume with no diagnostic value. Exclusion filters are the subject of 06-06.
  4. Tune the retention. Not all logs need keeping for 30 days; for some, 7 is enough, and what has to be kept for a long time is cheaper exported to Cloud Storage.
  5. Set a budget alert on the cost of observability itself, using what you learned in 01-04. It is ironic and it is exactly right.

And the perspective to keep hold of: observability is expensive compared with nothing and cheap compared with an incident. Twenty minutes of the shop being down during a campaign costs more than a month of logs. The decision is not "how much to spend on observability", but "what do I need to see so as not to be blind, and what am I paying to see that nobody ever looks at".

Common Mistakes and Tips

Using the average for latency. The most widespread conceptual error. The average hides exactly the problem you are looking for. Percentiles always: p50 for the typical, p95 and p99 for the bad.

Averaging percentiles across instances. The average of the p95s is not the p95 of the whole. It is a meaningless number. Configure the aligner and the aggregation correctly.

Alerting on causes instead of symptoms. An alert on CPU at 80 % produces false positives and false negatives. Alert when customers are suffering; investigate the cause afterwards with the dashboard.

Alerts with no duration. Any spike fires, the team learns to ignore them and the alerting system dies. No alert without a time window.

Absolute thresholds instead of proportions. "200 errors" means nothing without knowing how many requests there were. Alert on percentages.

High-cardinality labels on custom metrics. A customer or order identifier as a label: runaway cost and a useless metric. That goes into the logs.

Creating thirty alerts on day one. Fatigue appears quickly and is hard to reverse. Start with three, live with them for a month, and only add more when a real incident proves one was missing.

Alerts with no documentation. It reaches the phone at eleven at night and whoever receives it does not know what to look at. The documentation field with the first four steps is mandatory.

Forgetting the Ops Agent on the VMs. Without it there are no memory or disk metrics, and those are the causes of the hardest failures to diagnose.

A /salud endpoint that checks nothing. Returning 200 OK without verifying dependencies makes the check green with the database down. Have it check the essentials, with a short timeout.

Dashboards built by clicking and never exported. They get lost, they cannot be replicated and nobody knows who changed them. Export the JSON to Git.

A final tip: the best time to set up observability is before the incident. During an incident, with the shop down and the phone ringing, nobody has time to build a dashboard. And observability is not judged by how pretty it is on a normal day, but by how fast it takes you to the cause on the worst day of the year.

Exercises

Exercise 1: design an alert from historical data

AlpinaShop wants to alert on connections to alpinashop-pedidos. The historical data for the last 4 weeks: median 45 connections, p95 120, maximum observed 180 during a weekend campaign that worked correctly, and the instance's configured limit is 250. Design the complete alerting policy — metric, threshold, duration, channel and documentation — justifying each number, and explain what would happen with a threshold of 100 and with one of 240.

Exercise 2: choose what to measure for a problem nobody sees

Lucía notices that last week's sales report shows 40 fewer orders than expected on Tuesday mornings, consistently. No alert has fired, the dashboard is green and the uptime check has never failed. Propose which metrics you would look at, in what order, and what new instrumentation you would add so that this kind of problem is detected automatically in future.

Exercise 3: redesign an alerting system suffering from fatigue

A company similar to AlpinaShop has 34 alerts configured. Over the last month they fired 280 times; of those, 11 corresponded to real incidents. The team has muted the Slack channel and reviews the alerts "when it can". Diagnose the problem, propose a concrete method for reducing the number of alerts and define the minimum set you would start again with, explaining what kind of failure each one detects.

Solutions

Solution 1

The policy:

Element Value Justification
Metric cloudsql.googleapis.com/database/network/connections The direct one
Aligner ALIGN_MAX over 60 s The peak is what matters, not the minute's average
Threshold 200 connections 80 % of the limit of 250
Duration 5 minutes Filters out legitimate peaks
Channel Slack + email to Marta Serious, but not at 3 a.m.
Severity Warning It is an early warning, not an outage

Why 200. It is 80 % of the hard limit of 250. It leaves room to act before the instance starts refusing connections, and it is well above the maximum historically observed in normal operation (180), so it does not fire on the most extreme known legitimate use. The general rule: alert at a percentage of the limit, not at a multiple of the median, because what does the damage is exhausting the resource.

With a threshold of 100: it would fire constantly. The historical p95 is 120, that is to say, 5 % of the time 120 connections are exceeded in completely normal operation. An alert at 100 would fire several times a day with nothing wrong. It is the direct route to the fatigue of section 11 and to the channel being muted.

With a threshold of 240: it would arrive too late. At 240 out of 250 there are ten connections of headroom; at the speed this counter grows during a peak, the time between the alert and exhaustion is measured in seconds. Marta would receive the notification at the same time as the customers' first connection errors, and the alert would have stopped being preventive and become a notice that it is already too late.

Associated documentation, which is half the value:

## High Cloud SQL connection count

**What it means:** alpinashop-pedidos is above 200 of its 250 maximum connections.

**Impact if it reaches 250:** the application starts receiving connection errors
and the shop returns 500s. It has not happened yet.

**First steps:**
1. Is there a legitimate traffic peak? Autumn campaign dashboard, requests/s.
2. If there is NO traffic peak: probable connection leak after a deployment.
   Check Cloud Build: was there a deployment in the last 2 hours?
3. Check the number of MIG instances and pods: has it scaled a lot?
4. Check Cloud Functions with a high --max-instances connected to the DB (06-03).

**Immediate mitigation:** reduce --max-instances on the connected functions.
**Underlying mitigation:** review the application's connection pool.

And one further recommendation that improves the design considerably: add a second condition on the trend, not just the level. A jump from 45 to 190 connections in five minutes is a far more reliable sign of a leak than the absolute value, and it would warn sooner. Most connection leaks show up as a ramp that goes up and never comes down, and that is detectable well before reaching 80 % of the limit.

Solution 2

The first thing is to recognise what kind of problem this is: a partial, silent failure. There is no outage, no mass errors, no generalised slowness. Something is failing for a subset of users or of operations, and all the aggregate indicators dilute it. These are the most expensive failures because they last for weeks.

Metrics to look at, in this order:

Step What to look at What you are looking for
1 https/request_count on Tuesday mornings versus other days Is traffic dropping or only orders?
2 5xx error rate grouped by matched_url_path_rule An error localised to a specific path
3 p95 latency per path, especially /carrito and /pago Localised slowness
4 request_count grouped by user_agent or country Is a segment affected?
5 Error Reporting filtered by time of day New or increasing exceptions
6 Metrics of Tuesday's scheduled jobs The most likely hypothesis

The main hypothesis, and why. The regularity — Tuesday mornings, systematically — points strongly at something scheduled that runs on Tuesdays: a backup, a reindexing process, a heavy report over alpinashop_analitica, a maintenance task. That process competes for resources with the shop: it saturates Cloud SQL's CPU, occupies connections or locks tables, and during that window a fraction of the purchase attempts fail or are abandoned because of slowness.

Step 1 is especially informative and deserves spelling out: if traffic is normal but orders drop, the problem is in the purchase funnel; if traffic drops too, the problem is one of access — DNS, CDN, a marketing campaign that did not go out. They are completely different diagnoses and that single comparison separates them.

New instrumentation to add, which is the real goal of the exercise:

Instrumentation What it detects Priority
Business metric tienda/pedidos (section 12) The drop in orders, directly High
Funnel metric per stage: views → basket → payment → confirmed At which step people are lost High
Alert on orders/hour compared with the same hour the week before Anomalies, not fixed thresholds High
p95 latency per path on the dashboard Localised slowness Medium
Duration metric for the scheduled jobs Correlation with the degradation window Medium

The funnel metric is the key to the exercise. With counters per stage, the question "where are the orders being lost?" has an immediate answer: if product views are normal, baskets are normal and confirmed payments drop, the problem is in the gateway or in the final step. Without it, you have to reconstruct it by hand from the logs every time.

And the right alert is not a fixed threshold, but a comparative one. "Fewer than X orders per hour" is useless because the volume varies enormously between 4 in the morning and 8 in the evening, and between a Tuesday in July and a campaign Saturday. What works is comparing with the same time slot the week before and alerting on a deviation greater than 30 %. That comparison absorbs the daily and weekly seasonality.

And the underlying lesson: every technical indicator was green while 40 orders a week were being lost. Technical observability does not replace business observability. The most important metric of a shop is how many orders come in, and no platform metric knows it. This case is the best possible argument in favour of the custom metrics from section 12.

Solution 3

The numerical diagnosis, which is devastating. 280 alerts for 11 real incidents means a precision of 3.9 %: 96 out of every 100 alerts are noise. With that ratio, muting the channel is the rational decision, and the team has done nothing blameworthy. The alerting system is broken, not the people.

The real cost is not the 280 interruptions: it is that the 11 real incidents were buried. A system with that precision is not merely useless, it is worse than having no alerts, because it creates the illusion of being watched over.

The almost certain causes, deducible from the pattern:

Cause Symptom Fix
Alerts on causes and not symptoms Lots about CPU, memory, disk Retire them: they are charts, not alerts
Thresholds set by eye They fire during normal operation Recalculate with 4 weeks of history
No duration They fire on any spike Minimum 5 minutes on all of them
Absolute thresholds "N errors" with no context Convert to proportions
Everything to the same channel The urgent and the informational mixed together Separate by level
Never reviewed 34 accumulated alerts Mandatory quarterly review

A concrete reduction method, in five steps:

Step 1, measure each alert. For each of the 34: how many times it fired, how many corresponded to a real incident and what action it prompted. It is an afternoon's work with the data from Cloud Monitoring.

Step 2, apply the 3-in-the-morning question. For each alert: if this fires in the middle of the night, is there something somebody must do immediately? All the ones that answer "no" get retired. My bet: of 34, about 20 go.

Step 3, classify the survivors.

Class Criterion Destination
High precision and a clear action It fired rarely and was always real Kept
Right concept, bad threshold It detects the right thing but fires too much Recalibrated
Duplicate Another alert detects the same thing sooner Retired
Never fired Zero firings in a month Review whether it still makes sense

Step 4, recalibrate with data. For each survivor, look at the real value for the last four weeks and set the threshold above the historical p99 or at a percentage of the hard limit, with a minimum duration of 5 minutes.

Step 5, start from scratch with the minimum set. And here comes the part that is hard to accept: it is preferable to disable all 34 and start with 4 than to try to fix the 34. An alerting system that has already been discredited does not win back trust by being tweaked; it has to be refounded.

The minimum set to start with, each one detecting a different class of failure:

Alert Threshold Duration Failure it detects Channel
Uptime from 2+ regions Failure 3 min Total outage SMS
5xx error rate > 2 % 5 min Visible functional failure Slack + SMS
p95 latency > 2 s 10 min Visible degradation Slack
Orders/hour versus the previous week −30 % 30 min Silent business failure Slack

Four alerts. And the criterion for adding a fifth is strict and very useful: an alert is only added when a real incident has happened and none of the existing ones detected it. Every new alert has to be justified by a specific incident, not by a "just in case". That way the system grows guided by reality, and every alert that exists has a story backing it up.

Closing the exercise, with the lesson to take away: an alerting system is judged by its precision, not by its coverage. Four alerts with 80 % precision are worth infinitely more than thirty-four with 4 %, because the first get attended to and the second get muted. And a muted alert protects nothing.

Conclusion

AlpinaShop no longer finds out about its problems from a customer's email.

You know what observability is — understanding what is going on inside a system from what it emits outwards, including the failures you did not anticipate — and its three pillars, with their different profiles: cheap aggregated metrics for trends and alerts, detailed logs for specific cases, traces for knowing where the time went. And you know that Stackdriver is just a historical name.

You have mastered the data model: a time series is metric type + monitored resource + labels, with the consequence that labels multiply series and produce a cardinality explosion. You can tell apart gauge, counter and distribution, with the practical implications: a counter has to be derived with rate or it means nothing, and a distribution keeps the complete histogram so that percentiles can be asked for after the fact.

You have the three projects in a single metrics scope, with the warning that it unifies the view but not the permissions. You know the free platform metrics that have been accumulating for months without anybody looking at them, and the table that really gets used in an incident: from the symptom to the metric.

You know how to explore with aligners and aggregations in the right order, and above all why the average lies: 495 ms on average while 50 customers a minute wait eight seconds. Percentiles always, never average percentiles across series, and the p99 matters more than its name suggests because it usually falls on the best customers.

You have the autumn campaign dashboard built to tell a story from top to bottom — first whether the customers are suffering, then why — with the MIG's CPU shown as average and maximum because the average hides the hot instance. And you have it as versionable JSON in Git, ready to move to Terraform in 06-07. You know MQL for what filters cannot express — the error rate as a proportion — and PromQL as the more portable bet if you live in Kubernetes.

You know how to set up an alerting policy with its four pieces, with the duration as the underrated piece and the documentation field as what separates a useful alert from one that generates anxiety. You have AlpinaShop's three alerts, each with its reasoned threshold and each detecting a different kind of failure: visible error, visible degradation and failure invisible from the outside. And you have had the uncomfortable conversation about alert fatigue, with the control question that is worth the whole section: if this fires at 3 in the morning, is there anything to be done immediately? If not, it is a chart, not an alert.

You know how to emit business custom metrics — the orders per minute no GCP metric knows about — with the golden rule about labels: tens of values, never thousands, and the high-cardinality stuff goes into the logs. You have uptime checks from four continents that detect what no internal metric sees, with the two-region condition that makes them the most reliable alert in the system. You have Error Reporting grouping Flask exceptions by signature and notifying regressions, with the version that links every error to the deployment that introduced it. And you have the Ops Agent on the MIG's VMs, without which there are no memory or disk metrics and unexplained restarts stay unexplained.

And you know the cost: dashboards, alerts and platform metrics are free; what you pay for is what you ingest, with cardinality and log volume as the two traps. With the right perspective: observability is expensive compared with nothing and cheap compared with twenty minutes of the shop being down during a campaign.

But look at what metrics cannot do. When the 5xx error alert fires at eleven at night, it tells you that 4 % of requests are failing. It does not tell you what is failing. It does not tell you which customer it happened to, or with which product, or on which line of code. And when the p99 latency shoots up, the metric tells you that something is taking three seconds, but not who is taking it: the application, the database, the Vision API call, the network?

For that you need the other two pillars. In 06-06 come Cloud Logging and Cloud Trace: the logs that answer what exactly happened in this case, and the traces that answer where the time went. And with them, the complete journey through an AlpinaShop incident from the alert to the line of code.

Before that, however, there is an outstanding debt that can no longer be postponed. All this infrastructure — the VPC, the load balancer, the cluster, the alerts you have just created — still exists only because somebody ran the right commands in the right order, and that somebody was Marta, and the record of what she did is in her terminal history. In 06-05 that problem is tackled head-on.

Google Cloud Platform (GCP) Course

Module 1: Introduction to Google Cloud Platform

Module 2: Core GCP Services

Module 3: Networking and Security

Module 4: Data and Analytics

Module 5: Machine Learning and AI

Module 6: DevOps and Monitoring

Module 7: Advanced GCP Topics

Module 8: Final Project

© Copyright 2026. All rights reserved