Contoso Airlines today has a platform spread across web applications, an API on a scale set, containers, ephemeral functions, queues, topics, databases and AI services. When something goes wrong, the question is no longer "is the server down?", it is "which part of this chain is failing, since when, for how many passengers and why". Answering that requires the system to emit information and for somebody to have collected it before the problem happened. That is the discipline this module opens with.

This lesson is the map. You are going to understand what observability is and its three signals, how Azure Monitor gathers them all under a single umbrella, how platform metrics are configured along with the piece almost everybody forgets — the diagnostic setting — and how to build an alerting system that wakes somebody up only when it is worth it. All of it on Contoso's real resources.

Contents

  1. Observability: three signals, not a CPU alert
  2. Azure Monitor as an umbrella
  3. Platform metrics and the metrics explorer
  4. Diagnostic settings: the forgotten piece
  5. Activity log, resource logs and operating system logs
  6. Alerts: types, conditions and thresholds
  7. Action groups and processing rules
  8. Contoso's alert set
  9. SLI, SLO and error budget
  10. Dashboards and workbooks
  11. The cost of monitoring
  12. Common Mistakes and Tips
  13. Exercises
  14. Conclusion

  1. Observability: three signals, not a CPU alert

Monitoring is checking that a set of known indicators stays within range. Observability is the property of a system that lets you answer questions nobody had anticipated, without deploying new code. The difference matters: an alert on CPU at 90% is monitoring, and it would have said absolutely nothing about the passenger who paid and never received their boarding pass.

The three classic signals:

Signal What it is Nature Cost In Azure Answers
Metrics Numbers aggregated over time Numeric time series Very low Azure Monitor metrics Is there a problem? Since when?
Logs Discrete events with context Text and structured fields High, grows with volume Log Analytics What exactly happened?
Traces A request's journey across components A tree of correlated calls Medium, with sampling Application Insights Where did it break, and why?

The practical rule: metrics detect, logs explain, traces locate. A system that only has metrics knows something is wrong and does not know what. One that only has logs pays a fortune and searches blind. Contoso needs all three, and each one has its own lesson in this module.

  1. Azure Monitor as an umbrella

Azure Monitor is not a product: it is the name of the whole collection. It pays to see the full map from the start, so you know where each lesson fits.

flowchart TB
  subgraph ORI["Sources"]
    A1["Azure resources"]
    A2["VM and VMSS operating system"]
    A3["Application code"]
    A4["Subscription"]
  end
  subgraph PLAT["Data platforms"]
    M["Metrics<br/>93 days"]
    L["Logs<br/>log-contoso-pro"]
  end
  subgraph EXP["Experiences"]
    AI["App Insights (07-03)"]
    CI["Container / VM Insights"]
    NW["Network Watcher"]
  end
  subgraph CON["Consumption"]
    AL["Alerts and<br/>action groups"]
    LI["Workbooks and dashboards"]
    EX["Export<br/>Event Hubs / Storage"]
  end
  A1 --> M
  A1 -- "diagnostic setting" --> L
  A2 -- "agent + DCR" --> L
  A3 --> AI
  A4 -- "activity log" --> L
  AI --> L
  CI --> L
  NW --> L
  M --> AL
  L --> AL
  M --> LI
  L --> LI
  L --> EX

Three ideas from the diagram: there are two data platforms — metrics and logs — with different prices, retentions and query languages; data does not arrive on its own in the logs, you need a diagnostic setting or an agent, and that is the most frequent oversight; and the Insights (Application, Container, VM) are not separate stores but visual experiences built on top of those very same data.

  1. Platform metrics and the metrics explorer

Platform metrics are emitted by Azure automatically, with nothing to configure and at no cost, for almost every resource. Their properties:

  • One-minute granularity on most resources, aggregated afterwards into larger intervals.
  • 93-day retention. After that they disappear, unless you export them to Log Analytics.
  • Dimensions: attributes that let you break a metric down. An App Service's Http5xx can be split by instance; Storage requests, by API type or by response code.
  • Aggregations: average, minimum, maximum, total and count. Picking the wrong aggregation is the classic mistake — an hour's average latency hides the spike that angered the passengers.

Query the response time of the bookings website from the CLI:

# Full resource identifier; almost every metrics command asks for it
APP_ID=$(az webapp show \
  --name app-contoso-reservas-pro \
  --resource-group rg-contoso-reservas-pro \
  --query id -o tsv)

# Last day's metrics, grouped into 5-minute intervals
az monitor metrics list \
  --resource "$APP_ID" \
  --metric "HttpResponseTime" "Http5xx" "Requests" \
  --start-time 2026-08-14T00:00:00Z \
  --end-time   2026-08-15T00:00:00Z \
  --interval PT5M \
  --aggregation Average Maximum Total \
  --output table
  • --metric accepts several metrics at once; the names are the internal ones, not the portal's. To discover them: az monitor metrics list-definitions --resource "$APP_ID".
  • --interval PT5M is ISO 8601 duration notation: 5 minutes. PT1H would be one hour.
  • --aggregation asks for all three aggregations so you can compare average and maximum in the same table.

In the portal, the metrics explorer does the same thing visually: you pick a resource, a metric and an aggregation, and add a split by dimension. Marta Ríos' typical flow when a slowness complaint comes in is: HttpResponseTime as an average and as a percentile, split by instance. If a single instance is slow, it is a problem with that instance; if all of them are, the problem is further back — usually in sql-contoso-reservas-pro.

  1. Diagnostic settings: the forgotten piece

Platform metrics exist on their own. Resource logs do not. An App Service sends its HTTP logs nowhere, nor does SQL Database send its slow queries, nor Key Vault who accessed what, until you create a diagnostic setting that says which categories are exported and where to. It is common to discover this on the day of the incident, when it is already too late: logs do not exist retroactively.

Each resource type publishes its own categories. Real examples:

Resource Useful categories What for
App Service AppServiceHTTPLogs, AppServiceConsoleLogs, AppServiceAppLogs, AppServiceAuditLogs Requests, application output, who published
SQL Database SQLInsights, QueryStoreRuntimeStatistics, Errors, Timeouts, Deadlocks Slow queries, blocking
Key Vault AuditEvent Who read each secret
Storage (blob) StorageRead, StorageWrite, StorageDelete Access to tarjetas-embarque
Front Door FrontDoorAccessLog, FrontDoorWebApplicationFirewallLog Traffic and WAF blocks
Service Bus OperationalLogs Deliveries and failed messages

And three possible destinations, which can be combined:

Destination Use Relative cost Querying
Log Analytics Investigation and alerts High (per GB ingested) KQL, immediate
Storage account Cheap archive and compliance Very low Requires download and processing
Event Hubs Forwarding to third parties (SIEM, Splunk) Medium Outside Azure

Create the diagnostic setting for the bookings website:

LOG_ID=$(az monitor log-analytics workspace show \
  --workspace-name log-contoso-pro \
  --resource-group rg-contoso-seguridad-pro \
  --query id -o tsv)

az monitor diagnostic-settings create \
  --name diag-reservas-a-log \
  --resource "$APP_ID" \
  --workspace "$LOG_ID" \
  --logs '[
    {"category":"AppServiceHTTPLogs","enabled":true},
    {"category":"AppServiceAppLogs","enabled":true},
    {"category":"AppServiceAuditLogs","enabled":true}
  ]' \
  --metrics '[{"category":"AllMetrics","enabled":true}]' \
  --export-to-resource-specific true

Two parameters deserve attention. --export-to-resource-specific true makes the data land in dedicated tables (AppServiceHTTPLogs) instead of in the generic AzureDiagnostics table; it is a better schema, better performance and it allows different table plans, and you will make use of it in 07-02. --metrics AllMetrics duplicates the platform metrics into the logs: useful for correlating and for keeping them beyond 93 days, but it is ingestion you pay for, so switch it on only where you need it.

Doing this resource by resource does not scale. That is why Contoso already has, inside the "Base de gobernanza de Contoso" initiative from module 4, the diagnostico-app-service assignment with a DeployIfNotExists effect: any new App Service automatically gets its setting pointing at log-contoso-pro. Compliance is checked with az policy state summarize --resource-group rg-contoso-reservas-pro. Policy is what turns "remember to switch diagnostics on" into a guarantee. Remember that DeployIfNotExists only acts on new resources, or when you run a remediation task over the existing ones.

  1. Activity log, resource logs and operating system logs

Three layers that get confused constantly:

Layer What it records Example at Contoso How it is collected
Activity log Control plane operations: who created, modified or deleted a resource "Diego Salas restarted app-contoso-api-disponibilidad-pro" Automatic, 90 days; exportable
Resource logs What happens inside the service An HTTP request, a slow SQL query Diagnostic setting
OS logs Operating system events and counters Event log and syslog from vm-motor-disponibilidad-dev Azure Monitor Agent + DCR

The activity log always answers the question "what changed just before it broke?", which resolves a notable share of incidents:

az monitor activity-log list --resource-group rg-contoso-reservas-pro \
  --start-time 2026-08-14T06:00:00Z \
  --query "[].{Time:eventTimestamp, Who:caller, What:operationName.localizedValue, Status:status.value}" -o table

For virtual machines, the Azure Monitor Agent replaces the old agents and does not decide on its own what to collect: that is determined by a data collection rule (DCR), an independent, reusable resource that states which sources are read, how they are transformed and which workspace they go to. That separation lets you change the collection on a hundred machines by editing a single object: you create the dcr-contoso-servidores DCR once and associate it with each machine using az monitor data-collection rule association create.

A cost detail with consequences: a badly filtered DCR that collects all the syslog from a chatty server can cost more than the virtual machine itself. DCRs support KQL transformations at ingestion time so you can discard the irrelevant before paying for it; we will come back to that in 07-02.

  1. Alerts: types, conditions and thresholds

An alert is a rule that evaluates a data source and, if the condition is met, fires an action group. The available types:

Type Source Typical latency Cost Example at Contoso
Metric Platform or custom metrics ~1 min Low, per time series db-reservas DTU > 80%
Log search A KQL query over log-contoso-pro 5-15 min Higher, depending on frequency 20 exceptions of the same type in 5 min
Activity log Control plane ~1 min Free Somebody deletes an NSG rule
Service health Azure incidents Immediate Free App Service degradation in West Europe
Resource health The health of a specific resource ~1 min Free vmss-api-disponibilidad-pro unavailable

The service health and resource health ones are free and almost nobody configures them, even though they are the ones that tell "we broke something" apart from "Azure has a problem in the region". Always set them up.

On thresholds:

  • Static: you set the number. Predictable and auditable. Right when there is a commitment — the SLO says 800 ms, you alert at 800 ms.
  • Dynamic: Azure learns the historical pattern, including daily and weekly seasonality, and alerts on deviations. Ideal for metrics with a natural cycle: traffic to app-contoso-reservas-pro falls overnight, and a static threshold of "fewer than 100 requests per minute" would fire every night.

Two parameters govern the noise: the evaluation frequency (how often it is checked) and the aggregation window (how much time is aggregated on each check). Evaluating every minute over a one-minute window fires on any transient spike. Contoso uses evaluation every 5 minutes with a 15-minute window for latency, and every minute with a 5-minute window for availability.

  1. Action groups and processing rules

An action group is the list of who gets notified and what gets run. You define it once and reuse it across every rule.

az monitor action-group create \
  --name ag-guardia-contoso \
  --resource-group rg-contoso-seguridad-pro \
  --short-name GuardiaCon \
  --action email marta   [email protected] \
  --action email diego   [email protected] \
  --action sms   guardia 34 600000000 \
  --action webhook automatizacion https://ejemplo.contosoairlines.example/hook-runbook useCommonAlertSchema
  • --short-name (12 characters maximum) is what appears as the sender on the SMS.
  • useCommonAlertSchema forces the common alert schema, a JSON payload with the same shape for every type. Without it, each alert type sends a different payload and the consumer becomes unmaintainable. Always switch it on.

The available action types: email, SMS, push notification, voice call, email to an Azure Resource Manager role, webhook, secure webhook with Entra ID, Logic App, Azure Function, Automation runbook and ITSM connector. Contoso uses three groups:

Group Actions When
ag-guardia-contoso SMS + email + webhook Severity 0 and 1: wakes somebody up
ag-equipo-contoso Email + Teams message via Logic App Severity 2: looked at during business hours
ag-registro-contoso Webhook to the incident system only Severity 3 and 4: it is on the record

Alert processing rules act on alerts that have already been generated, without touching the rules. They serve two purposes: suppressing during a maintenance window and adding an action group to a set of alerts in one go.

az monitor alert-processing-rule create \
  --name apr-mantenimiento-sabado \
  --resource-group rg-contoso-seguridad-pro \
  --rule-type RemoveAllActionGroups \
  --scopes "/subscriptions/<id>/resourceGroups/rg-contoso-reservas-dev" \
  --schedule-recurrence Weekly --schedule-recurrence-days Saturday \
  --schedule-start-time 02:00 --schedule-end-time 06:00 \
  --description "Development patching window"

Without this, on the night of the monthly deployment the team gets forty text messages, and by the third time somebody silences the on-call phone permanently. Noise is not an annoyance: it is how an alerting system dies.

  1. Contoso's alert set

This is the actual table Marta maintains, and its value lies as much in what it contains as in what it deliberately does not.

Alert Source Condition Severity Group Who it wakes
/salud health endpoint down Availability test 3 of 5 locations fail for 5 min 0 ag-guardia-contoso On-call, in the middle of the night
API p95 latency App Insights metric p95 > 800 ms for 15 min 1 ag-guardia-contoso On-call
5xx errors on bookings Http5xx metric > 25 in 5 min 1 ag-guardia-contoso On-call
cola-emision-tarjetas depth Storage metric > 500 messages for 10 min 2 ag-equipo-contoso Diego, in working hours
db-reservas DTU SQL metric > 85% for 20 min 2 ag-equipo-contoso Bookings DBA
Certificate about to expire Activity log / Key Vault < 30 days of validity 3 ag-registro-contoso An incident record, nobody
VMSS resource health Resource health Degraded or unavailable 1 ag-guardia-contoso On-call

Notice what is not there: no CPU alert, no memory alert, no application disk alert. Not because they do not matter, but because they are causes, not symptoms. CPU at 95% with latency inside the target and no errors is not a problem: it is a well-used machine. Alerting on causes produces pointless wake-ups and, worse, trains the team to ignore the phone. You alert on what the passenger notices; causes are investigated afterwards, with metrics and logs, which is what they are for.

A metric rule via the CLI, to set the pattern (the mandatory tags apply here too: an alert rule is a billable resource and must be attributed to CC-1042):

az monitor metrics alert create \
  --name alerta-5xx-reservas \
  --resource-group rg-contoso-seguridad-pro \
  --scopes "$APP_ID" \
  --condition "total Http5xx > 25" \
  --window-size 5m \
  --evaluation-frequency 1m \
  --severity 1 \
  --action ag-guardia-contoso \
  --description "Server errors on the bookings website" \
  --tags entorno=produccion proyecto=contoso-reservas centro-coste=CC-1042 propietario=marta.rios

  1. SLI, SLO and error budget

Without these three concepts, thresholds get picked by intuition and argued about forever.

  • SLI (service level indicator): the concrete measurement. "Percentage of requests to /api/disponibilidad answered successfully in under 800 ms."
  • SLO (service level objective): the value you commit to. "99.5% monthly."
  • Error budget: what the SLO allows you to fail. 99.5% monthly is 3 hours and 39 minutes of non-compliance.

The error budget turns an argument about opinions into one about arithmetic: if 70% has been consumed in the first week, risky changes get frozen and the effort goes to reliability; if there is budget left at the end of the month, you deploy more cheerfully. And it answers the awkward question Nuria Peña will ask in module 8: going from 99.5% to 99.99% does not cost a little more, it costs multiplying the infrastructure.

  1. Dashboards and workbooks

Two tools that look alike and are not the same thing:

Dashboard Workbook
What it is A mosaic of fixed tiles An interactive document with text, parameters and queries
Parameters No Yes: environment, time range, resource
Source Pinned metrics and queries KQL, metrics, ARG, external JSON
Typical use An operations screen, always visible Guided investigation, monthly report
Shared as An Azure resource with its own RBAC An Azure resource, exportable as a template

Contoso's operations dashboard lives on a screen in the control center and shows six tiles: health endpoint availability, requests per minute and p95 latency for the API, 5xx errors on the website, cola-emision-tarjetas depth, db-reservas DTU and active alerts by severity. Nothing more: nobody looks at a dashboard with thirty charts. The parameterized workbook "Bookings Health" is the investigation tool instead, with three parameters — time range, environment and booking reference — that fill in the KQL queries of the next lesson; the same template serves production and development by changing a dropdown, and it is versioned in the contoso-infra repository alongside the Bicep templates. The availability tests that feed the first alert in the table belong to Application Insights and are explained in 07-03.

  1. The cost of monitoring

A warning worth internalizing now, even though the detail arrives in 07-02: data ingestion and retention in Log Analytics is one of the Azure bills that surprises people most. Nothing warns you that you left debug logging switched on in production; it is simply that at the end of the month the Log Analytics line is the biggest one in the subscription. Contoso lived through it: switching on AppServiceConsoleLogs across all four applications during an investigation and forgetting to switch it off cost more than the App Service plan that generated it.

Item Cost
Platform metrics and their 93-day retention Free
Activity log, 90 days Free
Activity log, service health and resource health alerts Free
Metric alert rules Cents per monitored time series per month
Log search alert rules Per rule and per frequency: evaluating every minute costs more
Notifications Email free up to a limit; SMS and voice are charged per message
Log Analytics ingestion and retention The dominant line item, per GB
Export to a storage account Very cheap, ideal for archiving

Two practical consequences: a log alert evaluated every minute across fifty resources is a considerable bill, so tune the frequency to what you genuinely need to detect; and alerts targeting SMS are best reserved for severity 0 and 1, which is exactly what ag-guardia-contoso does.

Common Mistakes and Tips

  • Believing logs exist without a diagnostic setting. On the day of the incident you discover there is nothing, and logs do not appear retroactively. Switch them on with policy, not by hand.
  • Alerting on causes. CPU, memory and disk generate noise; alert on symptoms visible to the passenger.
  • Picking the wrong aggregation. The average hides the spikes. For latency, percentiles; for errors, totals.
  • Not configuring the free alerts. Service health and resource health cost nothing and save you an hour investigating a problem that is Azure's.
  • Evaluation windows that are too short, or maintenance without suppression. Both produce predictable noise, and noise kills the system's credibility: use alert processing rules.
  • Tip: define the SLO before the threshold. If you do not know what you promised, any number is debatable.
  • Tip: every alert must carry a link to its runbook in the description; an alert at 3 in the morning with no instructions is half an alert. And review quarterly which ones fired and what was done about them: the ones that never led to an action get deleted.

Exercises

Exercise 1. Contoso wants to keep an eye on the new public "Contoso Miles" API (centro-coste=CC-2077), deployed on an App Service in rg-contoso-reservas-dev. Define: which signals you would switch on, with what diagnostic setting, three alerts with their severity and action group, and one justified cost decision.

Exercise 2. Over two weeks the team has received 340 alerts and has acted on 6. Analyze what is going on, propose four concrete measures and explain the risk of doing nothing.

Exercise 3. The business asks for a 99.99% availability SLO on db-reservas. Calculate the monthly error budget, state what it implies technically and what counterproposal you would make.

Solutions

Solution 1:

  • Signals and diagnostics: platform metrics (free, they already exist); resource logs via a diagnostic setting with AppServiceHTTPLogs and AppServiceAppLogs — not AppServiceConsoleLogs, which is the one that blows up the bill — into log-contoso-pro with --export-to-resource-specific true and a short retention because this is a trial; and traces with Application Insights (07-03). If the diagnostico-app-service assignment reaches rg-contoso-reservas-dev, the setting deploys itself: it is worth verifying.
  • Alerts: (1) resource health, severity 1, ag-guardia-contoso, free; (2) Http5xx > 25 in 5 min, severity 2, ag-equipo-contoso, because a side project does not justify waking anybody up; (3) HttpResponseTime with a dynamic threshold, severity 3, ag-registro-contoso, because there is no known baseline yet.
  • Cost: email notifications, no SMS, and evaluation every 15 minutes instead of every minute. Tags entorno=desarrollo, proyecto=contoso-millas, centro-coste=CC-2077, propietario.

Solution 2: the ratio is 1 action for every 57 notifications: the system has stopped being informative and the team has already learned to ignore it. The usual causes: alerts on causes (CPU, memory), one-minute evaluation windows that capture transient spikes, no suppression during deployments and maintenance, and a single incident firing ten rules at once. Measures: (1) delete every alert that has not led to an action in three months; (2) widen aggregation windows to 15 minutes on noisy metrics and move to dynamic thresholds where there is seasonality; (3) create alert processing rules for the deployment and maintenance windows; (4) reassign severities so that only 0 and 1 use SMS, leaving the rest on email or an incident record. The risk of not acting is concrete: the next real outage will arrive as notification number 341 and nobody will look at it.

Solution 3: 99.99% monthly leaves 4 minutes and 23 seconds of error budget per month — less than a manual failover takes. Technically it implies automatic failover with fg-contoso-reservas, a service tier with the corresponding SLA, zone redundancy, retries with reconnection logic in the application and zero-downtime deployments, on top of the fact that Azure's SLA for a single logical server is no longer enough. A reasonable counterproposal: 99.95% for the database (21 minutes and 54 seconds of budget) combined with an experience SLO — "the passenger can look up their booking" — backed by a read-only cache during the failover. It costs a fraction and it protects what the business actually values. And, on top of that, the commitment has to be measured against an agreed SLI, not against the availability the provider declares.

Conclusion

You now have the full map. You know that observability rests on three signals — metrics to detect, logs to explain, traces to locate — and that Azure Monitor is the umbrella covering the two data platforms, the Insights experiences and consumption through alerts, workbooks and dashboards. You know platform metrics: free, with one-minute granularity and 93 days of retention, their dimensions and aggregations, and how to explore them on app-contoso-reservas-pro. And above all, you have mastered the piece almost everybody forgets: the diagnostic setting, its categories per resource type, its three destinations and its deployment at scale through the diagnostico-app-service assignment with DeployIfNotExists.

You can tell the activity log — who changed what — from resource logs and from operating system logs, collected by the Azure Monitor Agent with its data collection rules. You have seen the five alert types, including the free ones almost nobody switches on, static versus dynamic thresholds and the role of the evaluation frequency; you have created ag-guardia-contoso with the common alert schema and you know how to suppress maintenance with processing rules. Contoso's alert set has taught you the rule that holds everything else up — alert on symptoms, never on causes — with SLI, SLO and error budget as the criterion for choosing thresholds instead of arguing about them, and with dashboards and parameterized workbooks as the query surface. And you carry the cost warning: platform metrics and alerts are almost free, text messages are charged for, and log ingestion is the line item that catches you out.

That is exactly where the next lesson goes. Everything you sent with the diagnostic settings has landed in log-contoso-pro, and so far you have only stored it. In 07-02, Log Analytics and KQL Queries, you will learn to interrogate it: the tables you are going to find, the KQL language from scratch and explained line by line, the full journey to investigate the complaint from the passenger who paid and never got their boarding pass, and the three levers — what gets ingested, on which plan and for how long it is kept — that control the bill we have just announced.

Azure Course

Module 1: Introduction to Azure

Module 2: Core Azure Services

Module 3: Azure Databases

Module 4: Security in Azure

Module 5: Azure DevOps

Module 6: Advanced Azure Services

Module 7: Monitoring and Management

Module 8: Cost Management and Optimization

Module 9: Case Studies and Best Practices

© Copyright 2026. All rights reserved