Ask Marta today how she knows that the AlpinaShop shop is working. The honest answer is: because nobody has complained. And if something goes wrong, she finds out in one of three ways: a customer writes to customer support, Dani happens to open the website and notices it is slow, or Lucía sees the next day that yesterday's sales were odd.
All three have the same defect: detection is done by a person, late and by chance. And the cost of finding out that way is high and silent. A twenty-minute incident in the middle of the autumn campaign can go completely unnoticed and take hundreds of orders with it.
This lesson changes that. By the end, AlpinaShop will have a dashboard that shows the state of the shop at a glance, checks that watch it from four continents, alerts that reach Marta's phone before any customer writes, and an automatic exception grouper that warns when a deployment introduces a new error.
A note on the name first. Stackdriver was a company Google acquired in 2014 whose product gave rise to the whole operations suite. The name was retired in 2020 and today the services are called Cloud Monitoring, Cloud Logging, Cloud Trace, Cloud Profiler and Error Reporting. You will see "Stackdriver" in old documentation, in legacy library names and in conversations with people who have been around a while; it is worth recognising, but it is a historical name.
Contents
- Observability and its three pillars
- The metrics data model
- The metrics scope: seeing all three projects at once
- Platform metrics you already have for free
- Exploring: aligners, aggregations and why the average lies
- The autumn campaign dashboard
- The dashboard as versionable JSON
- MQL and PromQL for advanced queries
- Alerts: the anatomy of a policy
- Notification channels and AlpinaShop's three alerts
- Alert fatigue: the uncomfortable conversation
- Custom metrics from the application
- Uptime checks
- Error Reporting: grouped exceptions
- The Ops Agent on the MIG's VMs
- The cost of observability
- Observability and its three pillars
Monitoring is checking whether the things you knew could fail are failing. Observability is the property of a system that lets you understand what is going on inside it from what it emits outwards, including the failures you did not anticipate.
The difference is not academic. A dashboard with the VMs' CPU is monitoring: it answers "is the CPU high?". Being able to answer "why has checkout been slow for customers in the Canary Islands since Tuesday?" without deploying anything new is observability.
It rests on three pillars:
| Pillar | What it answers | Nature | Service | Lesson |
|---|---|---|---|---|
| Metrics | How much? How many? What trend? | Numbers aggregated over time | Cloud Monitoring | This one |
| Logs | What exactly happened in this case? | Discrete events with detail | Cloud Logging | 06-06 |
| Traces | Where did the time go in this request? | The journey between components | Cloud Trace | 06-06 |
Each has a different profile of cost and usefulness, and that is why they coexist:
- Metrics are cheap to store and query because they are aggregates, they let you see months of history and they are the natural basis for alerts. But they do not tell you what happened to a specific customer: they tell you the p95 went up.
- Logs have all the detail of every event, and that is why they are expensive and you have to be selective. They answer specific questions about specific cases.
- Traces show how a request's time is distributed between services, and they are the only thing that answers "who is being slow?" in a system with several pieces.
The typical journey through an incident uses all three in order: an alert on a metric warns that something is happening, the metrics dashboard narrows down what and where, the logs say exactly what is failing, and the trace explains why. That complete journey is walked from start to finish in 06-06.
A note on scope: here we build metrics, dashboards and alerts. Service level objectives (SLOs) and error budgets — formally defining what "working well" means and how much failure is tolerable — arrive in 07-06. Here the instruments are installed; there it is decided which numbers are acceptable.
- The metrics data model
Almost every problem people have with Cloud Monitoring comes from not being clear about this model. It deserves five minutes.
A time series is the fundamental unit, and it is identified by the combination of three things:
| Component | What it is | Example |
|---|---|---|
| Metric type | What is measured | loadbalancing.googleapis.com/https/request_count |
| Monitored resource | Which object emits it | https_lb_rule with its project and its name |
| Labels | Dimensions for filtering and grouping | response_code_class="500", matched_url_path_rule="/catalogo" |
| Points | (instant, value) pairs | (2026-08-05T10:00:00Z, 142) |
That design explains something that is surprising at first: the number of time series multiplies with the labels. If a metric has 3 response code classes, 5 paths and 2 regions, that is 30 series. If somebody adds a label with the customer identifier, and there are 50,000 customers, that is 1,500,000 series. That is called a cardinality explosion, you pay for it and it can end up making the metric unusable. We will come back to it in section 12, which is where the mistake gets made.
The three metric kinds
Telling them apart is essential because each one is queried differently:
| Kind | What it represents | Example | How it is queried |
|---|---|---|---|
| Gauge | An instantaneous value that goes up and down | CPU at 42 %, 87 open connections | Read directly |
| Cumulative counter | A total that only grows from an origin | Total requests served | It has to be derived: rate per second |
| Distribution | A histogram of values in an interval | Latencies of every request | Percentiles are extracted |
The counter is the one that causes the most confusion. The load balancer's request_count is cumulative: its value grows indefinitely. If you plot it as it is, you see an upward ramp that says nothing. What you want is the rate: how many requests per second. That is why counters always get a rate or delta aligner applied.
The distribution is the most valuable and the most under-used. A latency metric as a distribution does not store one number per interval: it stores a complete histogram. That lets you ask for the 50th, 95th and 99th percentile after the fact, without having decided in advance which one interested you. If instead you had stored only the average latency, you would have lost the information forever, and the average is precisely the least useful statistic here, as we will see in section 5.
- The metrics scope: seeing all three projects at once
AlpinaShop has its resources spread across three projects: the shop in alpinashop-prod, testing in alpinashop-dev and analytics in alpinashop-datos. By default, each project has its own metrics view, which forces you to jump between three consoles to understand an incident. Unsustainable.
A metrics scope solves this: one project acts as the scoping project and aggregates the metrics of several monitored projects.
flowchart TD
A[alpinashop-prod<br/>SCOPING PROJECT] --> D[Single metrics scope:<br/>dashboards and alerts for all three]
B[alpinashop-dev] --> D
C[alpinashop-datos] --> D
# Add the other two projects to the alpinashop-prod metrics scope
gcloud beta monitoring metrics-scopes create \
--monitored-project=alpinashop-dev \
--project=alpinashop-prod
gcloud beta monitoring metrics-scopes create \
--monitored-project=alpinashop-datos \
--project=alpinashop-prodA design detail worth thinking about before running that: which project should be the scoping one? Using alpinashop-prod is the most common and the simplest. The alternative — creating a dedicated project, for example alpinashop-observabilidad — is conceptually cleaner because it separates "operating the platform" from "being the platform", and it allows read access to the metrics to be granted without granting any access to production. For AlpinaShop, with three projects and three people, alpinashop-prod is enough; in an organization with twenty projects, the dedicated project clearly wins.
And an important warning: the metrics scope unifies the view, not the permissions. Anyone consulting a dashboard that shows metrics from alpinashop-datos needs monitoring read permissions on that project. It is consistent with 03-04, and it stops aggregating the view becoming an indirect way of exposing information.
- Platform metrics you already have for free
Here is good news a lot of people do not know about: GCP is already collecting hundreds of metrics from your resources, without you having done anything and at no extra cost. All the infrastructure from modules 2, 3 and 4 has been emitting data for months that nobody has looked at.
The ones that matter for AlpinaShop:
| Resource | Metric | What it indicates |
|---|---|---|
| Load balancer | https/request_count |
Requests per second, with the code class |
| Load balancer | https/total_latencies |
End-to-end latency (distribution) |
| Load balancer | https/backend_latencies |
Backend-only latency |
| Load balancer | https/backend_request_count |
Requests reaching the backend (as opposed to the CDN) |
| Compute Engine / MIG | instance/cpu/utilization |
CPU saturation |
| Compute Engine | instance/disk/read_ops_count |
Disk pressure |
| Cloud SQL | database/cpu/utilization |
CPU of alpinashop-pedidos |
| Cloud SQL | database/network/connections |
Open connections |
| Cloud SQL | database/disk/utilization |
Space used |
| Cloud SQL | database/replication/replica_lag |
Replica lag |
| Pub/Sub | subscription/num_undelivered_messages |
Unacknowledged messages |
| Pub/Sub | subscription/oldest_unacked_message_age |
Age of the oldest one |
| BigQuery | query/scanned_bytes |
Bytes processed: a direct cost |
| GKE | container/cpu/limit_utilization |
Usage against the limit |
| Cloud Functions | function/execution_count |
Invocations (remember 06-03) |
| Cloud Functions | function/execution_times |
Duration (distribution) |
| Cloud Run | request_count, request_latencies |
The same as the load balancer |
| Cloud CDN | https/request_count filtered by cache_result |
Hit ratio |
The table that really gets used in an incident is the inverse one: from the symptom to the metric.
| Observed symptom | Metric to look at | What it means if it is bad |
|---|---|---|
| "The website is slow" | https/total_latencies p95 and p99 |
Confirms or refutes the perception |
| "It is slow" and the backend is fine | Compare total_latencies with backend_latencies |
If they differ a lot, the problem is the network or the client |
| 500 errors | https/request_count filtered by response_code_class=500 |
How many and since when |
| Slowness + high CPU in the MIG | instance/cpu/utilization |
Not enough capacity: check the autoscaling |
| Slowness + low CPU everywhere | database/network/connections |
Probable exhaustion of the connection pool |
| Errors after a deployment | request_count 5xx + Error Reporting |
Regression introduced by the deployment |
| Stale analytics data | num_undelivered_messages |
The consumer cannot keep up or is down |
| High BigQuery bill | query/scanned_bytes per user |
Somebody is running SELECT * over the history |
| High network bill | https/request_count by cache_result |
The CDN hit ratio has dropped |
That last row connects directly with 03-03: if Cloud CDN's hit ratio drops, traffic goes to the backend, the egress bill goes up and latency gets worse. It is a perfect silent failure — nothing breaks, everything degrades — and it is only detected by looking at the metric.
- Exploring: aligners, aggregations and why the average lies
Metrics Explorer is the interactive query tool. Its model has four steps and it is worth understanding them because they are the same ones that later show up in dashboards and alerts.
flowchart LR
A[1. Select<br/>metric and resource] --> B[2. Filter<br/>by labels]
B --> C[3. ALIGN<br/>within each series]
C --> D[4. AGGREGATE<br/>across series]
D --> E[Chart]
Aligning turns the raw points of each series into regular points at a fixed interval. Aggregating combines several series into fewer series or into a single one. The order matters and confusing them produces meaningless numbers.
| Aligner | What it does | For which kind | Example |
|---|---|---|---|
rate |
Change per second | Counters | Requests per second |
delta |
Difference over the interval | Counters | Requests in 60 s |
mean |
Average over the interval | Gauges | Average CPU per minute |
max / min |
Extremes of the interval | Gauges | Peak connections |
percentile_95 |
Percentile within the interval | Distributions | p95 latency |
| Aggregation | What it does | Example |
|---|---|---|
sum |
Sum across series | Total requests from all the instances |
mean |
Average across series | Average CPU of the MIG |
max |
Maximum across series | The busiest instance |
count |
How many series there are | Number of live instances |
| Group by label | One series per value | Latency per path |
Why the average lies
This is the most important concept in the section, and probably the one most often explained badly in the industry.
Imagine 1,000 requests to the catalogue in a minute. 950 take 100 ms and 50 take 8 seconds. The average is:
A dashboard showing the average shows 495 ms. That looks acceptable. And yet 50 customers a minute are waiting eight seconds, which is more than enough time to go to another website. The average has hidden exactly the problem you were looking for.
The percentiles tell a different story:
| Statistic | Value | What it really means |
|---|---|---|
| Average | 495 ms | A number that happens to nobody |
| p50 (median) | 100 ms | Half the customers see this |
| p95 | ~100 ms | 95 % are fine |
| p99 | 8,000 ms | The worst 1 % suffer 8 seconds |
Three practical rules worth adopting without exception:
- For latency, always percentiles, never the average. The p50 tells you what the typical experience is like; the p95 and the p99 tell you how bad the bad is.
- Never average percentiles across series. The average of ten instances' p95s is not the p95 of the whole: it is a number with no statistical meaning. Cloud Monitoring computes the percentile over the combined distribution if you ask for
percentile_95as the aligner andsumas the aggregation of distributions; doing it the other way round produces rubbish. - The p99 matters more than its name suggests. With a million requests a day, the p99 is 10,000 requests. And they do not fall at random on 10,000 different customers: they tend to concentrate on those with large baskets or long histories, that is to say, the best customers.
- The autumn campaign dashboard
The autumn campaign is AlpinaShop's commercial peak. Marta wants a screen that answers "is the shop all right?" at a glance.
The design principle, imported from SRE practice: a dashboard must tell a story from top to bottom, starting with what the customer perceives and ending with the technical causes. If you have to think while looking at it, the dashboard is wrong.
| Row | Chart | Metric | Why it is there |
|---|---|---|---|
| 1. Is traffic arriving? | Requests/s | https/request_count with rate and sum |
A collapse is as serious as a spike |
| 1 | 5xx error rate | request_count filtered / total |
The most direct indicator of pain |
| 2. Is it fast? | Latency p50/p95/p99 | https/total_latencies |
Three lines on the same chart |
| 2 | Latency per path | The same, grouped by matched_url_path_rule |
Isolates which part is slow |
| 3. Is it holding up? | MIG CPU | instance/cpu/utilization, mean and max |
The average hides the hot instance |
| 3 | Active instances | count of MIG series |
Is it scaling? |
| 4. And the DB? | Cloud SQL connections | database/network/connections |
Common cause of slowness with no CPU |
| 4 | Cloud SQL CPU | database/cpu/utilization |
Heavy queries |
| 5. And the CDN? | Hit ratio | request_count by cache_result |
Cost and latency |
| 5 | Bytes served | https/response_bytes_count |
Detects pattern changes |
Two design decisions deserve an explanation.
The MIG's CPU is shown twice, mean and max. If ten instances are at 20 % and one is at 98 %, the average says 27 % — apparently healthy — while that instance is saturated and serving extremely slow requests. It is the same conceptual error as averaging latencies, applied to capacity. The average hides imbalances; the maximum gives them away.
Latency and errors are at the top, the infrastructure at the bottom. The order reflects the question that is really asked in an incident: first "are the customers suffering?" and only then "why?". A dashboard that starts with the CPU invites you to look at infrastructure metrics that may be perfectly fine while the shop is down.
- The dashboard as versionable JSON
A dashboard built by clicking in the console is a fragile asset: nobody knows who changed it, it cannot be replicated in another project and it disappears if somebody deletes it. Every Cloud Monitoring dashboard has a JSON representation that can be exported, versioned in Git — following 06-02 — and recreated.
# Export an existing dashboard
gcloud monitoring dashboards list --project=alpinashop-prod --format=json
gcloud monitoring dashboards describe DASHBOARD_ID \
--project=alpinashop-prod --format=json > paneles/campana-otono.json
# Create or update from the file
gcloud monitoring dashboards create --config-from-file=paneles/campana-otono.json \
--project=alpinashop-prodA fragment with the two charts from row 1, which shows the structure without overwhelming you:
{
"displayName": "AlpinaShop - Autumn campaign",
"mosaicLayout": {
"columns": 12,
"tiles": [
{
"width": 6, "height": 4, "xPos": 0, "yPos": 0,
"widget": {
"title": "Requests per second",
"xyChart": {
"dataSets": [{
"timeSeriesQuery": {
"timeSeriesFilter": {
"filter": "metric.type=\"loadbalancing.googleapis.com/https/request_count\" resource.type=\"https_lb_rule\" resource.label.\"url_map_name\"=\"alpinashop-url-map\"",
"aggregation": {
"alignmentPeriod": "60s",
"perSeriesAligner": "ALIGN_RATE",
"crossSeriesReducer": "REDUCE_SUM"
}
}
},
"plotType": "LINE"
}]
}
}
},
{
"width": 6, "height": 4, "xPos": 6, "yPos": 0,
"widget": {
"title": "Latency p50 / p95 / p99 (ms)",
"xyChart": {
"dataSets": [
{
"legendTemplate": "p50",
"timeSeriesQuery": {
"timeSeriesFilter": {
"filter": "metric.type=\"loadbalancing.googleapis.com/https/total_latencies\" resource.type=\"https_lb_rule\"",
"aggregation": {
"alignmentPeriod": "60s",
"perSeriesAligner": "ALIGN_PERCENTILE_50",
"crossSeriesReducer": "REDUCE_MEAN"
}
}
},
"plotType": "LINE"
},
{
"legendTemplate": "p99",
"timeSeriesQuery": {
"timeSeriesFilter": {
"filter": "metric.type=\"loadbalancing.googleapis.com/https/total_latencies\" resource.type=\"https_lb_rule\"",
"aggregation": {
"alignmentPeriod": "60s",
"perSeriesAligner": "ALIGN_PERCENTILE_99",
"crossSeriesReducer": "REDUCE_MEAN"
}
}
},
"plotType": "LINE"
}
]
}
}
}
]
}
}Notice how it maps onto section 5: perSeriesAligner is the aligner — ALIGN_RATE for the counter, ALIGN_PERCENTILE_99 for the distribution — and crossSeriesReducer is the aggregation. The JSON is nothing more than the written form of what you did in the explorer.
The practical way to work is to build the dashboard in the console until it looks right, export it, save it in alpinashop-infra/paneles/ and manage it as code from then on. And when Terraform arrives in 06-07, the google_monitoring_dashboard resource consumes exactly this JSON, so the dashboard comes to be created alongside the infrastructure it monitors.
- MQL and PromQL for advanced queries
The filter and aggregation interface covers most cases, but there are questions it cannot express: ratios between two metrics, joins, derived calculations. For that there are two languages.
MQL (Monitoring Query Language) is Cloud Monitoring's own language, with a pipeline syntax:
fetch https_lb_rule | metric 'loadbalancing.googleapis.com/https/request_count' | filter resource.url_map_name == 'alpinashop-url-map' | align rate(1m) | every 1m | group_by [metric.response_code_class], [requests: sum(value.request_count)]
And the case where MQL becomes essential, the error rate as a proportion:
{
fetch https_lb_rule :: loadbalancing.googleapis.com/https/request_count
| filter metric.response_code_class == '500'
| align rate(1m) | every 1m | group_by [], [errors: sum(value.request_count)]
;
fetch https_lb_rule :: loadbalancing.googleapis.com/https/request_count
| align rate(1m) | every 1m | group_by [], [total: sum(value.request_count)]
}
| join
| value [error_rate: errors / total * 100]That — dividing two different series — cannot be expressed with filters and aggregations. And it is exactly what you want in order to alert: 2 % of errors is what is relevant, not "200 errors", because 200 errors out of 100,000 requests is noise and out of 500 it is a catastrophe.
PromQL is Prometheus's language, supported in Cloud Monitoring through Managed Service for Prometheus. The same query:
sum(rate(loadbalancer_googleapis_com:https_request_count{response_code_class="500"}[1m]))
/
sum(rate(loadbalancer_googleapis_com:https_request_count[1m])) * 100| Criterion | MQL | PromQL |
|---|---|---|
| Reach | GCP only | Industry standard |
| Portability | None | High: Prometheus, Grafana, other clouds |
| GKE metrics | It works | Natural: it is what the applications emit |
| Community and examples | Limited | Enormous |
| 2026 recommendation | For occasional cases on GCP | Preferable if you come from Kubernetes |
For AlpinaShop, which has the shop on GKE and aspires not to tie itself too closely to one provider, PromQL is the more sensible bet in the medium term. MQL remains useful for one-off queries over platform metrics that have no Prometheus equivalent.
- Alerts: the anatomy of a policy
A dashboard is only useful if somebody looks at it. An alert works for you.
An alerting policy has four pieces, and each one answers a different question:
flowchart LR
A[CONDITION<br/>what to measure and what threshold] --> B[DURATION<br/>how long it has to be bad]
B --> C[AGGREGATION<br/>over which set]
C --> D[NOTIFICATION<br/>to whom and how]
D --> E[DOCUMENTATION<br/>what to do on receiving it]
The duration is the most underrated piece. Without it, any instantaneous spike fires an alert. A 15-second latency spike because the autoscaler started an instance is not an incident: it is the system working. An alert that fires for that teaches the team to ignore it, which is the worst thing that can happen to an alert.
The criterion for choosing the duration is honest and simple: how long does this have to be bad before it is worth waking somebody up? If the answer is "five minutes", set five minutes.
A complete policy in JSON, AlpinaShop's error rate one:
{
"displayName": "AlpinaShop - High 5xx error rate",
"combiner": "OR",
"conditions": [
{
"displayName": "5xx errors above 2% for 5 minutes",
"conditionMonitoringQueryLanguage": {
"query": "{ fetch https_lb_rule :: loadbalancing.googleapis.com/https/request_count | filter metric.response_code_class == '500' | align rate(1m) | every 1m | group_by [], [e: sum(value.request_count)] ; fetch https_lb_rule :: loadbalancing.googleapis.com/https/request_count | align rate(1m) | every 1m | group_by [], [t: sum(value.request_count)] } | join | value [err_rate: e / t * 100] | condition err_rate > 2 '%'",
"duration": "300s",
"trigger": { "count": 1 }
}
}
],
"alertStrategy": {
"autoClose": "1800s"
},
"notificationChannels": [
"projects/alpinashop-prod/notificationChannels/CANAL_SLACK_ALERTAS",
"projects/alpinashop-prod/notificationChannels/CANAL_CORREO_INFRA"
],
"documentation": {
"mimeType": "text/markdown",
"content": "## High 5xx error rate\n\n**What it means:** more than 2 % of requests to the shop are returning a server error.\n\n**Impact:** customers are seeing error pages. Orders are being lost.\n\n**First steps:**\n1. *Autumn campaign* dashboard: has the latency gone up too?\n2. Was there a deployment in the last 30 minutes? Check Cloud Build.\n3. Logs Explorer: `severity=ERROR` in the last 15 minutes (06-06).\n4. Check the Cloud SQL connections: a frequent cause.\n\n**Rollback:** promote the previous image with the pipeline from 06-01.\n\n**Escalation:** if there is no diagnosis within 15 minutes, notify Marta."
}
}gcloud alpha monitoring policies create \
--policy-from-file=alertas/tasa-error-5xx.json \
--project=alpinashop-prodThe documentation field is what separates a useful alert from an alert that generates anxiety. It arrives on somebody's phone at eleven at night; with no instructions, that person starts from scratch, sleepy and in a hurry. With the first four checks written down, they know what to look at. And the 30-minute autoClose stops an already-resolved incident leaving an incident open forever.
- Notification channels and AlpinaShop's three alerts
A channel defines how the alert arrives:
| Channel | Latency | Interrupts | Use at AlpinaShop |
|---|---|---|---|
| Minutes | No | Informational notices, summaries | |
| Slack | Seconds | A little | The #alpinashop-alertas channel, the main one |
| SMS | Seconds | Yes | Only for the critical stuff |
| PagerDuty / Opsgenie | Seconds | Yes, with escalation | When there are formal on-call rotas (07-06) |
| Webhook | Seconds | It depends | Your own automations |
| Pub/Sub | Seconds | No | Reacting with a function (06-03) |
gcloud beta monitoring channels create \
--display-name="AlpinaShop alerts Slack" \
--type=slack \
--channel-labels=channel_name="#alpinashop-alertas" \
--project=alpinashop-prodThe Pub/Sub channel deserves a mention for what it enables: an alert can publish to a topic and a Cloud Function from 06-03 can react automatically — create an issue, run a diagnostic, or even take a narrowly scoped corrective action. It is the door to automated response, with the obvious caveat that an automation acting on production must be very tightly scoped.
The three alerts AlpinaShop starts with
Deliberately three. Not thirty.
Alert 1 — 5xx error rate > 2 % for 5 minutes. Channel: Slack + SMS to Marta.
Reasoning behind the threshold: the shop's normal error rate is 0.1-0.3 %, dominated by requests to URLs that no longer exist. 2 % is an order of magnitude above the noise: it does not fire from natural variation, and at that level there are customers seeing errors. It is a proportion, not an absolute value, for the reason given in section 8. Five minutes filters out deployment spikes.
Alert 2 — p95 latency > 2 seconds for 10 minutes. Channel: Slack.
Reasoning: the shop's usual p95 is around 400 ms. Two seconds is the point where perception changes from "it works fine" to "it is slow" and basket abandonment starts to rise. p95 is chosen rather than p99 because the p99 is more volatile and would generate more noise; and a long duration, 10 minutes, is chosen because transient slowness is common and does not require immediate action. No SMS: it is a problem to be dealt with during working hours.
Alert 3 — Unacknowledged messages on pedidos-nuevos > 1,000 for 15 minutes. Channel: Slack + email.
Reasoning: in normal operation the queue is practically empty because the consumers keep up. A thousand messages piled up for fifteen minutes means the consumer is down or cannot keep up, and that means there are orders that are not being processed. It is the example of an alert that detects a failure that is invisible from the outside: the website works, customers buy, everything looks fine, and the orders are piling up unprocessed.
| Alert | Threshold | Duration | Channels | What it protects |
|---|---|---|---|---|
| 5xx errors | > 2 % | 5 min | Slack + SMS | Customers seeing errors |
| p95 latency | > 2 s | 10 min | Slack | Degraded experience |
| Order queue | > 1,000 msg | 15 min | Slack + email | Failure invisible from the outside |
Notice the pattern: each alert detects a different kind of failure — visible error, visible degradation and invisible failure. Ten more alerts about CPU would have added nothing, because high CPU already shows up in the latency.
- Alert fatigue: the uncomfortable conversation
This deserves a section of its own because it is the real problem of observability, and it is almost never said clearly.
Alert fatigue is the state in which a team receives so many alerts that it stops reacting to them. It is not a failure of discipline: it is a rational response. If nineteen out of twenty daily alerts are noise, ignoring them all is a strategy with a 95 % success rate and an enormous cost when it fails.
How you get there, almost always by the same route:
- Somebody sets up an alert for every available metric, "just in case".
- The thresholds are set by eye, without looking at historical values.
- There is no duration, so any spike fires.
- They all go to the same channel with the same urgency.
- Nobody reviews or retires the ones that add nothing.
The five antidotes, in order of importance:
| Antidote | Concrete rule |
|---|---|
| Alert on symptoms, not on causes | Alert if customers are suffering, not if the CPU is at 80 % |
| Thresholds from historical data | Look at the real value for the last 4 weeks before deciding |
| Always a duration | No alert without a time window |
| Levels of urgency | Does this justify waking somebody up? If not, email or Slack |
| Periodic review | Every quarter: did it fire? was it useful? If not, retire it |
The first is the most important and the most counter-intuitive. A CPU at 90 % is not a problem if the customers do not notice it: it may be the system making good use of its resources. Alerting on CPU produces false positives — high CPU with no impact — and false negatives — customers suffering with low CPU, for example because of a lock in the database. Causes are investigated with the dashboard after a symptom has raised the alarm.
And a control question that is worth more than any list, for every alert you are about to create:
If this alert fires at 3 in the morning, is there something somebody must do immediately?
If the answer is no, it is not an alert: it is a chart on a dashboard. Marta and Dani have only three alerts configured, and that is deliberately the sign of a healthy alerting system.
- Custom metrics from the application
Platform metrics tell you what the infrastructure is doing. They do not tell you what the business is doing. No GCP metric knows how many orders per minute come into AlpinaShop, and that is probably the most important metric of all: if orders drop to zero, something is wrong even if every technical indicator is green.
# catalogo/metricas.py
from google.cloud import monitoring_v3
import time, os
client = monitoring_v3.MetricServiceClient()
PROJECT = f"projects/{os.environ['GOOGLE_CLOUD_PROJECT']}"
def record_order(amount_eur: float, channel: str) -> None:
"""Writes a point to the custom orders metric."""
series = monitoring_v3.TimeSeries()
series.metric.type = "custom.googleapis.com/tienda/pedidos"
# LABELS: few and of LOW cardinality. See the warning below.
series.metric.labels["canal"] = channel # web | mobile | phone
series.metric.labels["entorno"] = os.environ.get("ENTORNO", "prod")
series.resource.type = "generic_task"
series.resource.labels.update({
"project_id": os.environ["GOOGLE_CLOUD_PROJECT"],
"location": "europe-west1",
"namespace": "tienda",
"job": "catalogo-web",
"task_id": os.environ.get("HOSTNAME", "unknown"),
})
now = time.time()
point = monitoring_v3.Point({
"interval": {"end_time": {"seconds": int(now)}},
"value": {"double_value": 1.0},
})
series.points = [point]
client.create_time_series(name=PROJECT, time_series=[series])The warning about labels is the critical part of this section. It is tempting to add series.metric.labels["id_cliente"] = customer_id. Do not. With 50,000 active customers you would create 50,000 time series for every combination of the other labels, and that is the cardinality explosion from section 2: high cost, slow queries and a useless metric.
| Label | Possible values | Acceptable? |
|---|---|---|
canal |
3 | Yes |
entorno |
2-3 | Yes |
categoria_producto |
~20 | Yes |
codigo_postal |
~11,000 | No |
id_cliente |
50,000+ | Never |
id_pedido |
Unlimited | Never |
The rule: a metric label must have tens of possible values, not thousands. The high-cardinality stuff — the customer identifier, the order identifier — goes into the logs, which are designed for it, and is queried from there. It is one of the reasons the pillars are three and not one.
There is an alternative route that in 2026 is usually better if you already live in Kubernetes: expose a /metrics endpoint in Prometheus format and let Managed Service for Prometheus scrape it. It avoids an API call per event, it is the industry standard and it fits with PromQL. For the catalogue on GKE, it is the option I would recommend.
There are also log-based metrics: counting log entries that match a filter and turning them into a metric, without touching the application's code. It is a very powerful technique for quickly instrumenting what is already being logged, and it is developed in 06-06.
- Uptime checks
Everything above measures the system from the inside. Uptime checks measure it from the outside: Google sends requests from several regions of the world and checks that the shop responds.
That difference in point of view detects entire classes of failure that no internal metric sees: an expired TLS certificate, a badly propagated DNS record, an over-aggressive Cloud Armor rule blocking legitimate customers, a misconfigured load balancer or a complete regional outage.
gcloud monitoring uptime create tienda-alpinashop \
--resource-type=uptime-url \
--resource-labels=host=www.alpinashop.example,project_id=alpinashop-prod \
--path=/salud \
--port=443 \
--protocol=https \
--period=1 \
--timeout=10 \
--regions=EUROPE,USA_OREGON,ASIA_PACIFIC,SOUTH_AMERICA \
--content-matchers-content='"estado":"ok"' \
--project=alpinashop-prodFour decisions worth understanding:
- The
/saludpath is the same endpoint as the load balancer'shc-catalogohealth check from 03-02. Reusing it makes sense, but with one important nuance: the load balancer's check decides whether an instance receives traffic and must be fast and shallow; the uptime one measures whether the service is genuinely healthy. A/saludthat just returns200 OKwithout checking anything will be green with the database down. The endpoint should verify at least connectivity to Cloud SQL, with a short timeout so it does not itself become a problem. --content-matchers-contentchecks that the response contains what is expected, not just that the code is 200. A server that returns an error page with code 200 — more common than it sounds — is detected this way.- Four spread-out regions distinguish a global problem from a regional one. If only the check from Asia fails, the problem is network or DNS, not the application.
- A 1-minute period is the reasonable balance between fast detection and cost.
And the associated alert, with one deliberate detail:
gcloud alpha monitoring policies create \
--notification-channels=CANAL_SMS_MARTA \
--display-name="AlpinaShop - Shop not responding" \
--condition-display-name="Uptime check failing in 2+ regions" \
--condition-filter='metric.type="monitoring.googleapis.com/uptime_check/check_passed" resource.type="uptime_url"' \
--duration=180s \
--project=alpinashop-prodThe condition requires failure from two regions or more. A failure from a single region is usually a problem with the intermediate network, not with the shop, and alerting on it generates exactly the noise from section 11. Requiring two regions makes this the most reliable alert in the whole system: if the shop does not respond from two continents, the shop is down. It is AlpinaShop's only alert that goes straight to SMS without passing through Slack.
- Error Reporting: grouped exceptions
When the Flask catalogue raises an uncaught exception, the stack trace ends up in the logs. With thousands of requests a day, finding a new error in the logs is looking for a needle in a haystack.
Error Reporting solves exactly that: it collects the exceptions, groups them by their signature — the exception type and the stack, not the literal message — and presents a list of distinct problems with their frequency, their first occurrence and their last.
The practical difference: instead of 4,000 log lines, you see "KeyError: 'talla' in catalogo/carrito.py:87, 340 occurrences, first seen 2 hours ago, affecting 89 users". And that "first seen 2 hours ago" is pure gold, because it usually coincides with a deployment.
Instrumenting Flask is straightforward:
# catalogo/app.py
import google.cloud.logging
from google.cloud.error_reporting import Client as ErrorClient
logging_client = google.cloud.logging.Client()
logging_client.setup_logging() # Python logs go to Cloud Logging
error_client = ErrorClient(service="catalogo-web", version=os.environ["VERSION_IMAGEN"])
@app.errorhandler(Exception)
def handle_error(e):
# report_exception captures the full traceback from the current context
error_client.report_exception()
app.logger.exception("Unhandled error while serving %s", request.path)
return render_template("error.html"), 500The version parameter is more important than it looks: by passing the SHA of the deployed image, Error Reporting knows in which version each error appeared. That connects directly with 06-01 and turns a difficult question — "is this error new?" — into a fact. With that information, Error Reporting notifies regressions: it warns when an error type that did not exist appears, and when one that had been marked as resolved comes back.
| Capability | What it provides |
|---|---|
| Grouping by signature | 4,000 log lines → 12 distinct problems |
| Count and trend | Is it getting worse? |
| First and last occurrence | Correlation with deployments |
| Affected users | Prioritises by real impact, not by volume |
| Notification of new errors | Detects regressions automatically |
| Link to the source code | Jumps to the exact line (06-02 integration) |
An important piece of advice about exception design: Error Reporting groups by the signature of the stack, so a generic error repeated in twenty different places groups badly. Specific exceptions and messages with context — but with no personal data — make the grouping useful.
- The Ops Agent on the MIG's VMs
There is a gap in the platform metrics that surprises everybody the first time: Google cannot see inside your virtual machines. It knows the CPU, the network and the disk operations, because the hypervisor measures those. It does not know the memory used, the free space on the disk or the processes running, because those can only be seen from inside the operating system.
The Ops Agent is a single agent that collects system metrics and sends logs from the VMs. It replaces the old separate Stackdriver monitoring and logging agents.
Installing it on the templates of the alpinashop-web-mig MIG, via the startup script from 02-01:
#!/bin/bash
# Fragment of the instance template's startup script
curl -sSO https://dl.google.com/cloudagents/add-google-cloud-ops-agent-repo.sh
sudo bash add-google-cloud-ops-agent-repo.sh --also-installAnd its configuration, in /etc/google-cloud-ops-agent/config.yaml:
logging:
receivers:
catalogo_app:
type: files
include_paths: [/var/log/alpinashop/catalogo.log]
service:
pipelines:
catalogo:
receivers: [catalogo_app]
metrics:
receivers:
hostmetrics:
type: hostmetrics
collection_interval: 60s
service:
pipelines:
default:
receivers: [hostmetrics]What shows up in Cloud Monitoring once it is installed:
| Metric | Why it matters |
|---|---|
agent.googleapis.com/memory/percent_used |
The most common cause of unexplained restarts |
agent.googleapis.com/disk/percent_used |
A full disk brings the application down silently |
agent.googleapis.com/processes/count_by_state |
Zombie processes, process leaks |
agent.googleapis.com/swap/percent_used |
If swap is active, memory is short |
Memory deserves emphasis: a memory leak in the application ends with the system killing the process. From the outside it looks like random restarts and lost requests, with no metric to explain it. Without the agent, that diagnosis is practically impossible.
And a note on the future: when AlpinaShop moves the catalogue to Cloud Run (DA-001, 07-02), the agent is no longer needed because the platform exposes those metrics itself. It is a real advantage of managed services that rarely gets mentioned.
- The cost of observability
It is easy for the observability bill to grow without anybody noticing. It is worth knowing the model — always with the current prices from the official documentation:
| Component | You pay for | Free tier | Risk |
|---|---|---|---|
| Platform metrics | Nothing | All of them | None |
| Custom metrics | Time series ingested | An initial volume | High: cardinality |
| Logs | GB ingested | A generous monthly volume | The biggest of all (06-06) |
| Traces | Spans ingested | A monthly volume | Medium: controlled with sampling |
| Uptime checks | Executions | Ample | Low |
| Dashboards and alerts | Nothing | — | None |
The fact that dashboards, alerts and platform metrics are free is the best news in this lesson: everything built in sections 4 to 11 adds no cost. What you pay for is what you ingest: custom metrics and, above all, logs.
The five tips for keeping it from running away:
- Watch the cardinality of your custom metrics. It is trap number one: a customer identifier as a label multiplies the cost by thousands.
- Do not invent metrics that already exist. Before instrumenting, check whether GCP already emits it. Many custom metrics duplicate free metrics.
- Exclude the noise from the logs. Health probes and static assets generate an enormous volume with no diagnostic value. Exclusion filters are the subject of 06-06.
- Tune the retention. Not all logs need keeping for 30 days; for some, 7 is enough, and what has to be kept for a long time is cheaper exported to Cloud Storage.
- Set a budget alert on the cost of observability itself, using what you learned in 01-04. It is ironic and it is exactly right.
And the perspective to keep hold of: observability is expensive compared with nothing and cheap compared with an incident. Twenty minutes of the shop being down during a campaign costs more than a month of logs. The decision is not "how much to spend on observability", but "what do I need to see so as not to be blind, and what am I paying to see that nobody ever looks at".
Common Mistakes and Tips
Using the average for latency. The most widespread conceptual error. The average hides exactly the problem you are looking for. Percentiles always: p50 for the typical, p95 and p99 for the bad.
Averaging percentiles across instances. The average of the p95s is not the p95 of the whole. It is a meaningless number. Configure the aligner and the aggregation correctly.
Alerting on causes instead of symptoms. An alert on CPU at 80 % produces false positives and false negatives. Alert when customers are suffering; investigate the cause afterwards with the dashboard.
Alerts with no duration. Any spike fires, the team learns to ignore them and the alerting system dies. No alert without a time window.
Absolute thresholds instead of proportions. "200 errors" means nothing without knowing how many requests there were. Alert on percentages.
High-cardinality labels on custom metrics. A customer or order identifier as a label: runaway cost and a useless metric. That goes into the logs.
Creating thirty alerts on day one. Fatigue appears quickly and is hard to reverse. Start with three, live with them for a month, and only add more when a real incident proves one was missing.
Alerts with no documentation. It reaches the phone at eleven at night and whoever receives it does not know what to look at. The documentation field with the first four steps is mandatory.
Forgetting the Ops Agent on the VMs. Without it there are no memory or disk metrics, and those are the causes of the hardest failures to diagnose.
A /salud endpoint that checks nothing. Returning 200 OK without verifying dependencies makes the check green with the database down. Have it check the essentials, with a short timeout.
Dashboards built by clicking and never exported. They get lost, they cannot be replicated and nobody knows who changed them. Export the JSON to Git.
A final tip: the best time to set up observability is before the incident. During an incident, with the shop down and the phone ringing, nobody has time to build a dashboard. And observability is not judged by how pretty it is on a normal day, but by how fast it takes you to the cause on the worst day of the year.
Exercises
Exercise 1: design an alert from historical data
AlpinaShop wants to alert on connections to alpinashop-pedidos. The historical data for the last 4 weeks: median 45 connections, p95 120, maximum observed 180 during a weekend campaign that worked correctly, and the instance's configured limit is 250. Design the complete alerting policy — metric, threshold, duration, channel and documentation — justifying each number, and explain what would happen with a threshold of 100 and with one of 240.
Exercise 2: choose what to measure for a problem nobody sees
Lucía notices that last week's sales report shows 40 fewer orders than expected on Tuesday mornings, consistently. No alert has fired, the dashboard is green and the uptime check has never failed. Propose which metrics you would look at, in what order, and what new instrumentation you would add so that this kind of problem is detected automatically in future.
Exercise 3: redesign an alerting system suffering from fatigue
A company similar to AlpinaShop has 34 alerts configured. Over the last month they fired 280 times; of those, 11 corresponded to real incidents. The team has muted the Slack channel and reviews the alerts "when it can". Diagnose the problem, propose a concrete method for reducing the number of alerts and define the minimum set you would start again with, explaining what kind of failure each one detects.
Solutions
Solution 1
The policy:
| Element | Value | Justification |
|---|---|---|
| Metric | cloudsql.googleapis.com/database/network/connections |
The direct one |
| Aligner | ALIGN_MAX over 60 s |
The peak is what matters, not the minute's average |
| Threshold | 200 connections | 80 % of the limit of 250 |
| Duration | 5 minutes | Filters out legitimate peaks |
| Channel | Slack + email to Marta | Serious, but not at 3 a.m. |
| Severity | Warning | It is an early warning, not an outage |
Why 200. It is 80 % of the hard limit of 250. It leaves room to act before the instance starts refusing connections, and it is well above the maximum historically observed in normal operation (180), so it does not fire on the most extreme known legitimate use. The general rule: alert at a percentage of the limit, not at a multiple of the median, because what does the damage is exhausting the resource.
With a threshold of 100: it would fire constantly. The historical p95 is 120, that is to say, 5 % of the time 120 connections are exceeded in completely normal operation. An alert at 100 would fire several times a day with nothing wrong. It is the direct route to the fatigue of section 11 and to the channel being muted.
With a threshold of 240: it would arrive too late. At 240 out of 250 there are ten connections of headroom; at the speed this counter grows during a peak, the time between the alert and exhaustion is measured in seconds. Marta would receive the notification at the same time as the customers' first connection errors, and the alert would have stopped being preventive and become a notice that it is already too late.
Associated documentation, which is half the value:
## High Cloud SQL connection count
**What it means:** alpinashop-pedidos is above 200 of its 250 maximum connections.
**Impact if it reaches 250:** the application starts receiving connection errors
and the shop returns 500s. It has not happened yet.
**First steps:**
1. Is there a legitimate traffic peak? Autumn campaign dashboard, requests/s.
2. If there is NO traffic peak: probable connection leak after a deployment.
Check Cloud Build: was there a deployment in the last 2 hours?
3. Check the number of MIG instances and pods: has it scaled a lot?
4. Check Cloud Functions with a high --max-instances connected to the DB (06-03).
**Immediate mitigation:** reduce --max-instances on the connected functions.
**Underlying mitigation:** review the application's connection pool.And one further recommendation that improves the design considerably: add a second condition on the trend, not just the level. A jump from 45 to 190 connections in five minutes is a far more reliable sign of a leak than the absolute value, and it would warn sooner. Most connection leaks show up as a ramp that goes up and never comes down, and that is detectable well before reaching 80 % of the limit.
Solution 2
The first thing is to recognise what kind of problem this is: a partial, silent failure. There is no outage, no mass errors, no generalised slowness. Something is failing for a subset of users or of operations, and all the aggregate indicators dilute it. These are the most expensive failures because they last for weeks.
Metrics to look at, in this order:
| Step | What to look at | What you are looking for |
|---|---|---|
| 1 | https/request_count on Tuesday mornings versus other days |
Is traffic dropping or only orders? |
| 2 | 5xx error rate grouped by matched_url_path_rule |
An error localised to a specific path |
| 3 | p95 latency per path, especially /carrito and /pago |
Localised slowness |
| 4 | request_count grouped by user_agent or country |
Is a segment affected? |
| 5 | Error Reporting filtered by time of day | New or increasing exceptions |
| 6 | Metrics of Tuesday's scheduled jobs | The most likely hypothesis |
The main hypothesis, and why. The regularity — Tuesday mornings, systematically — points strongly at something scheduled that runs on Tuesdays: a backup, a reindexing process, a heavy report over alpinashop_analitica, a maintenance task. That process competes for resources with the shop: it saturates Cloud SQL's CPU, occupies connections or locks tables, and during that window a fraction of the purchase attempts fail or are abandoned because of slowness.
Step 1 is especially informative and deserves spelling out: if traffic is normal but orders drop, the problem is in the purchase funnel; if traffic drops too, the problem is one of access — DNS, CDN, a marketing campaign that did not go out. They are completely different diagnoses and that single comparison separates them.
New instrumentation to add, which is the real goal of the exercise:
| Instrumentation | What it detects | Priority |
|---|---|---|
Business metric tienda/pedidos (section 12) |
The drop in orders, directly | High |
| Funnel metric per stage: views → basket → payment → confirmed | At which step people are lost | High |
| Alert on orders/hour compared with the same hour the week before | Anomalies, not fixed thresholds | High |
| p95 latency per path on the dashboard | Localised slowness | Medium |
| Duration metric for the scheduled jobs | Correlation with the degradation window | Medium |
The funnel metric is the key to the exercise. With counters per stage, the question "where are the orders being lost?" has an immediate answer: if product views are normal, baskets are normal and confirmed payments drop, the problem is in the gateway or in the final step. Without it, you have to reconstruct it by hand from the logs every time.
And the right alert is not a fixed threshold, but a comparative one. "Fewer than X orders per hour" is useless because the volume varies enormously between 4 in the morning and 8 in the evening, and between a Tuesday in July and a campaign Saturday. What works is comparing with the same time slot the week before and alerting on a deviation greater than 30 %. That comparison absorbs the daily and weekly seasonality.
And the underlying lesson: every technical indicator was green while 40 orders a week were being lost. Technical observability does not replace business observability. The most important metric of a shop is how many orders come in, and no platform metric knows it. This case is the best possible argument in favour of the custom metrics from section 12.
Solution 3
The numerical diagnosis, which is devastating. 280 alerts for 11 real incidents means a precision of 3.9 %: 96 out of every 100 alerts are noise. With that ratio, muting the channel is the rational decision, and the team has done nothing blameworthy. The alerting system is broken, not the people.
The real cost is not the 280 interruptions: it is that the 11 real incidents were buried. A system with that precision is not merely useless, it is worse than having no alerts, because it creates the illusion of being watched over.
The almost certain causes, deducible from the pattern:
| Cause | Symptom | Fix |
|---|---|---|
| Alerts on causes and not symptoms | Lots about CPU, memory, disk | Retire them: they are charts, not alerts |
| Thresholds set by eye | They fire during normal operation | Recalculate with 4 weeks of history |
| No duration | They fire on any spike | Minimum 5 minutes on all of them |
| Absolute thresholds | "N errors" with no context | Convert to proportions |
| Everything to the same channel | The urgent and the informational mixed together | Separate by level |
| Never reviewed | 34 accumulated alerts | Mandatory quarterly review |
A concrete reduction method, in five steps:
Step 1, measure each alert. For each of the 34: how many times it fired, how many corresponded to a real incident and what action it prompted. It is an afternoon's work with the data from Cloud Monitoring.
Step 2, apply the 3-in-the-morning question. For each alert: if this fires in the middle of the night, is there something somebody must do immediately? All the ones that answer "no" get retired. My bet: of 34, about 20 go.
Step 3, classify the survivors.
| Class | Criterion | Destination |
|---|---|---|
| High precision and a clear action | It fired rarely and was always real | Kept |
| Right concept, bad threshold | It detects the right thing but fires too much | Recalibrated |
| Duplicate | Another alert detects the same thing sooner | Retired |
| Never fired | Zero firings in a month | Review whether it still makes sense |
Step 4, recalibrate with data. For each survivor, look at the real value for the last four weeks and set the threshold above the historical p99 or at a percentage of the hard limit, with a minimum duration of 5 minutes.
Step 5, start from scratch with the minimum set. And here comes the part that is hard to accept: it is preferable to disable all 34 and start with 4 than to try to fix the 34. An alerting system that has already been discredited does not win back trust by being tweaked; it has to be refounded.
The minimum set to start with, each one detecting a different class of failure:
| Alert | Threshold | Duration | Failure it detects | Channel |
|---|---|---|---|---|
| Uptime from 2+ regions | Failure | 3 min | Total outage | SMS |
| 5xx error rate | > 2 % | 5 min | Visible functional failure | Slack + SMS |
| p95 latency | > 2 s | 10 min | Visible degradation | Slack |
| Orders/hour versus the previous week | −30 % | 30 min | Silent business failure | Slack |
Four alerts. And the criterion for adding a fifth is strict and very useful: an alert is only added when a real incident has happened and none of the existing ones detected it. Every new alert has to be justified by a specific incident, not by a "just in case". That way the system grows guided by reality, and every alert that exists has a story backing it up.
Closing the exercise, with the lesson to take away: an alerting system is judged by its precision, not by its coverage. Four alerts with 80 % precision are worth infinitely more than thirty-four with 4 %, because the first get attended to and the second get muted. And a muted alert protects nothing.
Conclusion
AlpinaShop no longer finds out about its problems from a customer's email.
You know what observability is — understanding what is going on inside a system from what it emits outwards, including the failures you did not anticipate — and its three pillars, with their different profiles: cheap aggregated metrics for trends and alerts, detailed logs for specific cases, traces for knowing where the time went. And you know that Stackdriver is just a historical name.
You have mastered the data model: a time series is metric type + monitored resource + labels, with the consequence that labels multiply series and produce a cardinality explosion. You can tell apart gauge, counter and distribution, with the practical implications: a counter has to be derived with rate or it means nothing, and a distribution keeps the complete histogram so that percentiles can be asked for after the fact.
You have the three projects in a single metrics scope, with the warning that it unifies the view but not the permissions. You know the free platform metrics that have been accumulating for months without anybody looking at them, and the table that really gets used in an incident: from the symptom to the metric.
You know how to explore with aligners and aggregations in the right order, and above all why the average lies: 495 ms on average while 50 customers a minute wait eight seconds. Percentiles always, never average percentiles across series, and the p99 matters more than its name suggests because it usually falls on the best customers.
You have the autumn campaign dashboard built to tell a story from top to bottom — first whether the customers are suffering, then why — with the MIG's CPU shown as average and maximum because the average hides the hot instance. And you have it as versionable JSON in Git, ready to move to Terraform in 06-07. You know MQL for what filters cannot express — the error rate as a proportion — and PromQL as the more portable bet if you live in Kubernetes.
You know how to set up an alerting policy with its four pieces, with the duration as the underrated piece and the documentation field as what separates a useful alert from one that generates anxiety. You have AlpinaShop's three alerts, each with its reasoned threshold and each detecting a different kind of failure: visible error, visible degradation and failure invisible from the outside. And you have had the uncomfortable conversation about alert fatigue, with the control question that is worth the whole section: if this fires at 3 in the morning, is there anything to be done immediately? If not, it is a chart, not an alert.
You know how to emit business custom metrics — the orders per minute no GCP metric knows about — with the golden rule about labels: tens of values, never thousands, and the high-cardinality stuff goes into the logs. You have uptime checks from four continents that detect what no internal metric sees, with the two-region condition that makes them the most reliable alert in the system. You have Error Reporting grouping Flask exceptions by signature and notifying regressions, with the version that links every error to the deployment that introduced it. And you have the Ops Agent on the MIG's VMs, without which there are no memory or disk metrics and unexplained restarts stay unexplained.
And you know the cost: dashboards, alerts and platform metrics are free; what you pay for is what you ingest, with cardinality and log volume as the two traps. With the right perspective: observability is expensive compared with nothing and cheap compared with twenty minutes of the shop being down during a campaign.
But look at what metrics cannot do. When the 5xx error alert fires at eleven at night, it tells you that 4 % of requests are failing. It does not tell you what is failing. It does not tell you which customer it happened to, or with which product, or on which line of code. And when the p99 latency shoots up, the metric tells you that something is taking three seconds, but not who is taking it: the application, the database, the Vision API call, the network?
For that you need the other two pillars. In 06-06 come Cloud Logging and Cloud Trace: the logs that answer what exactly happened in this case, and the traces that answer where the time went. And with them, the complete journey through an AlpinaShop incident from the alert to the line of code.
Before that, however, there is an outstanding debt that can no longer be postponed. All this infrastructure — the VPC, the load balancer, the cluster, the alerts you have just created — still exists only because somebody ran the right commands in the right order, and that somebody was Marta, and the record of what she did is in her terminal history. In 06-05 that problem is tackled head-on.
Google Cloud Platform (GCP) Course
Module 1: Introduction to Google Cloud Platform
- What is Google Cloud Platform?
- Setting Up Your GCP Account
- A Tour of the GCP Console
- Projects, Resource Hierarchy and Billing
- Regions, Zones and the Shared Responsibility Model
- Cloud Shell and the gcloud CLI
Module 2: Core GCP Services
- Compute Engine: Virtual Machines on Google Cloud
- Cloud Storage: Object Storage
- Cloud SQL: Managed Relational Databases
- App Engine: Platform as a Service
- Google Kubernetes Engine (GKE)
- NoSQL Databases: Firestore, Bigtable and Spanner
- How to Choose the Right Compute Service
Module 3: Networking and Security
- VPC Networks
- Cloud Load Balancing
- Cloud CDN
- Identity and Access Management (IAM)
- Cloud Armor
- Secrets and Encryption: Secret Manager and Cloud KMS
- Cloud DNS, TLS Certificates and Publishing Services Securely
Module 4: Data and Analytics
- BigQuery: The Analytical Data Warehouse
- Cloud Dataflow: Batch and Streaming Data Processing
- Cloud Dataproc: Managed Spark and Hadoop
- Cloud Pub/Sub: Asynchronous Messaging
- Cloud Data Fusion: Code-Free Data Integration
- Orchestrating Pipelines with Cloud Composer and Workflows
- Data Governance and Dashboards with Dataplex and Looker Studio
Module 5: Machine Learning and AI
- Vertex AI: The Machine Learning Platform on GCP
- AutoML: Custom Models Without Writing Code
- TensorFlow on GCP: Training and Serving Models
- Natural Language API
- Vision API
- Generative AI on Vertex AI: Gemini Models and Embeddings
- MLOps: From Model to Product with Vertex AI Pipelines
Module 6: DevOps and Monitoring
- Cloud Build: Continuous Integration on GCP
- Cloud Source Repositories and Source Code Management
- Cloud Functions: Serverless Functions
- Cloud Monitoring (formerly Stackdriver): Metrics, Dashboards and Alerts
- Cloud Deployment Manager and Native Infrastructure as Code
- Cloud Logging and Cloud Trace: Logs, Traces and Diagnostics
- Terraform on GCP: Infrastructure as Code in Practice
Module 7: Advanced GCP Topics
- Hybrid and Multicloud with Anthos
- Serverless Computing with Cloud Run
- Advanced Networking: Shared VPC, Peering and Hybrid Connectivity
- Security Best Practices
- Cost Management and Optimization
- Reliability: SLOs, High Availability and Disaster Recovery
- Governance at Scale: Organization, Policies and Auditing
