Over seven modules we have built the pieces: boundaries and data (module 2), contracts and events (3), code (4), deployment (5), operations (6) and security (7). TechCorp showed up as the example in every one of them, but we have never told the whole story in order: what the team did first, what came next, what went sideways along the way and what the outcome was. This lesson is that account, the one the close of module 7 announced: how the migration was carried out step by step. It is a case study, not a technique lesson: when a phase uses the strangler fig, the outbox or a canary, we will point to the lesson where it is explained instead of repeating it, and we will not write code (the missing services are implemented in 08-02 and the complete system is deployed in 08-03).
We will walk through the roughly 13 fictional months of the migration (July 2025 to August 2026): phase 0 of prerequisites, the six extraction phases in the order decided in 02-02 (catalog, inventory, notifications, payments, orders, customers), the shutdown of the monolith, a chronological table, the metrics before and after, the real costs and three stumbles worth knowing about so you do not repeat them. All dates, figures and names are invented, but plausible for a store with 3,000 orders a day and 25 engineers.
Contents
- Starting point and plan (July 2025)
- Phase 0: the prerequisites (July–September 2025)
- Phase 1: catalog, the first extraction and the first Black Friday (September–November 2025)
- Phase 2: inventory, the first event consumer (December 2025–January 2026)
- Phase 3: notifications, email out of the process (February 2026)
- Phase 4: payments, money and caution (March–April 2026)
- Phase 5: orders, the saga replaces the transaction (May–July 2026)
- Phase 6: customers and the shutdown of the monolith (July–August 2026)
- Chronological table of the migration
- Metrics before and after
- Real costs and what went wrong
- The final architecture
- Starting point and plan (July 2025)
In July 2025 Marta presents to management the plan that came out of the analysis in 01-04 and 01-05: techcorp-shop is a Node.js/Express monolith on a single PostgreSQL with five measurable problems (Thursday deployments with 4 rollbacks out of 12, 50 minutes without selling on the previous Black Friday, 20% of the time spent on coordination, the email memory leak that took down charging, and the product_attributes table). The proposal is not "rewrite the store" but extract six services one at a time, with the store still selling, in the order from 02-02: catalog → inventory → notifications → payments → orders → customers.
Three conditions Marta sets that govern everything that follows:
| Condition | Practical consequence |
|---|---|
| The store does not stop. No planned outage longer than a few minutes; no campaign put at risk. | Everything is done with strangler fig and branch by abstraction (02-02) and with rollback by configuration. |
| Every phase delivers visible value in under a quarter. | The catalog goes first because it solves the most expensive problem (campaign outages) with the least risk. |
| At most two database technologies and a single language. | PostgreSQL + MongoDB (02-04); Node.js 20 in every service (04-01). |
And a rule of Luis's that will show up in every phase: "if it is done more than once per service, automate it before the second time" (04-01, 05-03). It is what turns a six-service migration into something a 25-person team can sustain.
- Phase 0: the prerequisites (July–September 2025)
The first important decision was not to extract anything for three months. The four prerequisites from 01-04 (CI/CD, containers, observability, ownership culture) did not exist, and extracting a service without them would have produced a distributed monolith operated by hand. What was done:
- Organization. The "back-end / front-end / systems" teams were reorganized into the four teams from 02-01: Shopping Experience (catalog and customers), Orders (orders and inventory, with Luis), Payments & Communications (payments and notifications) and Platform (gateway, Kubernetes, CI/CD, observability). Each team owned its areas inside the monolith from day one: Conway's law applied before touching the code.
- CI for the monolith. The
ci.ymlfrom 05-03 in its minimal version: tests on every PR, a Docker image of the monolith (the multi-stageDockerfilefrom 05-01) published toghcr.io/techcorp/techcorp-shop:sha-…. The 40-minute suite was parallelized down to 12. Deployments still happened on Thursdays, but with a reproducible image. - Kubernetes and the gateway. Platform stood up the cluster (05-02), deployed the monolith onto it (three replicas, the probes from 03-05) and put the gateway on port 8080 in front, with every route pointing at the monolith (03-04). For six weeks the gateway did nothing but forward and log: it is the strangler fig requirement from 02-02, "the facade exists before the first extraction".
- Minimal observability. Structured logs with pino shipped to Loki, RED metrics for the gateway and the monolith in Prometheus, one Grafana dashboard (06-01). No traces yet.
- The modular monolith. The table × area matrix from 02-02 had exposed that
createOrderwrote tostockand topaymentswithout owning them. Before moving anything, Luis's team created thewarehouse.reserve()/warehouse.release()andpayments.recordCharge()modules inside the monolith, andcreateOrderstarted calling them instead of writing someone else's SQL. It was the least glamorous work of the migration and the one that sped it up the most: when phases 2 and 4 arrived, the seam was already there. @techcorp/common-httpv0.1 (04-01): logger,X-Request-Id, RFC 7807 errors, health routes. It was adopted in the monolith first, so that the first service would be born looking the same.
Outcome of phase 0 (September 30, 2025): no service extracted, no business problem solved yet… and a platform the six following phases would lean on without arguing about tooling again.
- Phase 1: catalog, the first extraction and the first Black Friday (September–November 2025)
catalog-service (3001, MongoDB) was built exactly as shown in 04-02, with the freshly created node-service-template. The sequence was the one from 02-04 §6, and it deserves to be seen with dates because it is the mold for all the others:
| Week | Step | Technique (lesson) |
|---|---|---|
| 1–3 (Sep) | Service with GET /v1/products in its three forms and an OpenAPI contract. Component and contract tests. |
04-02, 04-05 |
| 4 | Initial load with scripts/migrate-catalog.js (idempotent, upsert): 41,000 products and their key-value attributes converted into documents. |
02-04 §6 |
| 5–6 (Oct) | Synchronization: the monolith publishes product.updated to RabbitMQ (its first event) and Catalog consumes it. RabbitMQ enters production here. |
03-02 |
| 7 | Shadow comparison: the monolith reads from both sources and logs discrepancies. 212 products with a different price were found (rounding in the script) and 3 events lost to a restart: the script was fixed and rerun (idempotent). |
02-04 §6 |
| 8 | Read cutover: the gateway routes /api/v1/products/* to Catalog (strangler) and createOrder switches to HttpProductsRepository with REMOTE_CATALOG=true at 10%, 50%, 100% over three days. |
02-02 §4, 03-04 |
| 9 (Nov) | Write cutover: the product administration panel moves to catalog-service; the monolith stops writing to products. |
02-04 |
| 10–13 | The products table becomes read-only; retired on December 15, after confirming that only one monthly report used it (it was rewritten on top of a read-only view exported by Catalog; in phase 6 it would move to the read database from section 8). |
02-04 §6 |
TechCorp's first canary deployment was Catalog's (05-04): 10% of traffic for two hours watching the RED dashboard. And the first HPA too (06-04): minReplicas: 2, maxReplicas: 20, with the 30 s Redis cache in front of MongoDB.
The exam came on November 28, 2025. With the catalog separated and autoscaled, Black Friday went like this: browsing ×22 over a normal day, Catalog went from 2 to 14 replicas in 12 minutes, p95 of GET /v1/products at 140 ms, and order creation and charging —still in the monolith— did not notice a thing: 0 minutes without selling versus 50 the year before. Marta took that number to management and the migration stopped being up for debate. It was also the moment it became clear that the catalog was 70% of the previous campaign's infrastructure cost: scaling only what saturates cost a third of replicating the monolith onto ten servers.
- Phase 2: inventory, the first event consumer (December 2025–January 2026)
inventory-service (3006, PostgreSQL) was the first service with delicate business state and the first one that consumes events. What changed compared with the catalog:
- Reservation model. The monolith only had
stock.reserved; the service is born with thestock,reservations(withexpires_at) andreservation_linestables from 02-04 §7.2. The reservation goes from being a counter to being an entity with ACTIVE/CONSUMED/RELEASED states (02-03) and a 900 s expiry. - Branch by abstraction for the reservation. The
warehousemodule created in phase 0 received its second implementation (HttpStockReservation→POST /v1/reservationsfrom 03-01) and theREMOTE_INVENTORYflag. For three weeks the monolith reserved over HTTP;OUT_OF_STOCKwas still a synchronous409increateOrder. - The first consumer. Since the monolith had been publishing events since phase 1, Inventory started consuming
order.confirmedandorder.cancelledfrom theinventory.ordersqueue to consume or release the reservation (03-02). With that came the first DLQ with messages (anorder.cancelledfor an order whose reservation had still been made through the old SQL path: a poison message, with no reservation to release) and the first version ofprocessOnceoutside Orders. - Luis's rule in action. The outbox and consumer idempotency had been written in Orders (still inside the monolith) to publish
order.confirmed; when Inventory needed them, they were moved to@techcorp/common-http(messaging/outbox,messaging/idempotency) instead of being copied. It is the version the four services in 08-02 use. - KEDA (06-04) showed up here: scaling Inventory by the length of
inventory.orderswas more useful than doing it by CPU.
The January sales (January 7, 2026) were the load test: orders ×3.5, Inventory from 2 to 6 replicas, no incidents. The stock data migration (12,000 rows) was trivial compared with the catalog's: initial load, shadow double-read for a week, cutover.
- Phase 3: notifications, email out of the process (February 2026)
The shortest extraction (four weeks) and one of the most rewarding. services/email.js and its templates became notifications-service (3005), with no database except the deliveries table to avoid resending (02-04 §7.4). It consumes order.confirmed and order.cancelled from notifications.orders; the monolith stopped calling email.send() in createOrder and started trusting the event.
Two immediate effects:
- The email memory leak stopped being the monolith's problem. The email provider failure happened again on March 3, 2026 (piled-up connections): this time the Notifications pod was restarted by its
livenessProbetwice, messages waited in the queue and innotifications.orders.retry(06-03), and no charge was affected. Incident 4 from 01-05, closed by design. - The
createOrderresponse dropped by 300 ms on average: the SMTP send was no longer in the request.
It was also the phase in which Notifications debuted the deferred retries with TTL and the DLQ as they ended up in 06-03, because an email provider returns 503 frequently: retrying immediately was useless.
- Phase 4: payments, money and caution (March–April 2026)
payments-service (3003, PostgreSQL) took services/paymentProvider.js and returnsController.js with it. It was the extraction with the most review and the least hurry, and the one that debuted the most security pieces:
payments.recordCharge()→payments-servicewith branch by abstraction and a flag; the monolith was still the one asking to charge (synchronously, insidecreateOrder), but the one talking to the payment provider was now the service.- Tokenization and PCI scope (07-03): the card form started talking directly to the payment provider; TechCorp only receives
tok_…. The PCI DSS scope was reduced to the service and itspayments-providerSecret. - Signed webhook
POST /v1/webhooks/payment-providerwith HMAC, a time window andreceived_webhooks(07-02 §9): the payment provider confirms charges and refunds asynchronously. - Audit (07-03 §9): insert-only
audit_logtable for manual payment status changes and refunds. UNIQUE (order_id)onpaymentsand idempotency key =orderIdtoward the payment provider (06-03 ex. 1): the "one order, one charge" guarantee moved from the monolith's transaction to two explicit constraints.- Mandatory canary deployment (05-04) with human review of the promotion PR (05-03 §8), as befits a service that moves money.
One data point: the team spent an entire week on the resilience test (06-03 §11) against the fake payment provider —503s, timeouts, duplicate responses— before moving a single real cent. Not one duplicate charge in production since.
- Phase 5: orders, the saga replaces the transaction (May–July 2026)
The central phase, the longest (eleven weeks) and the one that gives meaning to the previous ones. When it arrived, Catalog, Inventory, Notifications and Payments were already services and the monolith was already publishing and consuming events: createOrder was still a function that called the others over HTTP and coordinated with its local transaction. What changed:
orders-service(3002) as built in 04-04:Orderaggregate, the state machine from 02-05,POST /v1/orders→202, outbox with relay,orders.sagaandorders.customersconsumers.- The choreographed saga replaced the transaction: Inventory started reacting to
order.created(instead of receiving a synchronousPOST /v1/reservations), Payments tostock.reserved(instead of the monolith's call), Orders topayment.confirmed. TheASYNC_SAGAflag in the monolith allowed, for two weeks, a percentage of orders to follow the new path while the rest used the old one; the "Orders saga" dashboard (06-01) compared both. - The visible contract change:
POST /ordersstopped answering201 CONFIRMEDand started answering202 PENDING+ polling with ETag (03-01). It was the only change of the migration that required coordination with the front-end team and with the mobile app (throughbff-mobile, 03-04): two sprints of notice, aDeprecationheader on the old route (03-06). - The
PAYMENT_TIMEOUTwatchdog (06-03 §9) and theCronJob reconcile-reservationswere written before the cutover, not after an incident: a choreographed saga without a watchdog is a saga that gets stuck. - Strangler cutover:
/api/v1/orders/*to the service on June 15, 2026 at 10%, 100% on the 19th;ASYNC_SAGA=trueat 100% on June 26; migration of the 2.1 million historical orders to theordersDB over three nights in batches (idempotent initial load + shadow status-change events); removal of thecreateOrdercode from the monolith on July 10.
And on July 22, 2026, INC-2031 (06-05 §10): a stock.reserved with lines: null published by Inventory 2.3.0 jammed payments.stock for 40 minutes; 61 orders affected, 38 cancelled by the watchdog. It is the incident that gave operations their final shape: retry queues and straight-to-DLQ for permanent errors in Payments, DlqHasMessages and SagaLateBurn* alerts, validation of outgoing events against the AsyncAPI schema in Inventory, and the common/dlq runbook with scripts/reprocessDlq.js. Without INC-2031, half of lesson 06-05 would not exist; with it, the system ended up as it is today.
- Phase 6: customers and the shutdown of the monolith (July–August 2026)
customers-service (3004) was last for the reasons in 02-02: identity and personal data, and customer_id was everywhere. Being last, it benefited from everything before it:
- Keycloak (07-01) had entered production at the start of phase 5, federating the credentials from the monolith's
customerstable so that/api/v1/orders/*could be protected with JWT. In phase 6 it was completed: identities were migrated to thetechcorprealm, customer sign-up moved tocustomers-service, which writes thecustomerIdattribute in Keycloak, and thecustomerIdclaim travels in the token; the monolith stopped holding sessions. customers_refin Orders (02-04) had been fed bycustomer.updatedpublished by the monolith since phase 5; only the producer changed.- GDPR (07-03 §8):
DELETE /v1/customers/{id}→ anonymization +customer.deleted; Orders deletes the replica and anonymizes old addresses. GET /v1/customers/{id}with the rule "the customer themselves, an operator, or a service token withcustomers:read" (07-01 ex. 2), andPOST /v1/orderswith the check bodycustomerId= tokencustomerId.
The shutdown of the monolith happened on August 7, 2026. What was still inside at that point and where it went:
| Leftover in the monolith | Destination |
|---|---|
Sales reports and management dashboards (the four-table JOINs from 02-04 §8) |
Read database techcorp-analytics (PostgreSQL) fed by events (order.confirmed, product.updated, customer.updated) with its own consumer; the reports were rewritten against it. It is the "company-scale" CQRS from 02-05. |
| Internal administration panel (several areas) | Lightweight web application that consumes the APIs through the gateway with the operator role. |
| Nightly export to the ERP | Event consumer that generates the file. |
Historical read-only orders table |
Migrated to orders in phase 5; the monolith's copy was kept for a month and archived. |
# 2026-08-07 10:30 — the monolith's last commit: the gateway no longer has a catch-all route
kubectl -n techcorp scale deployment techcorp-shop --replicas=0 # nobody notices: it has not received traffic for three weeks
# 2026-08-21 — after two weeks at zero replicas with no request to the catch-all (metric gateway_monolith_routes_total = 0): deleted
kubectl -n techcorp delete deployment,service,configmap -l app=techcorp-shopThat scale --replicas=0 for two weeks before deleting is the last application of the phase 1 principle: no hurry in retiring the old.
- Chronological table of the migration
| Phase | Dates | Service | Main technique | Lessons | Main risk | Outcome |
|---|---|---|---|---|---|---|
| 0 | Jul–Sep 2025 | — (modular monolith, platform) | CI, containers, K8s + gateway in front, minimal observability, four teams, returning stock/payments writes to their owner |
01-04, 02-02, 03-04, 05-01/02/03, 06-01 | Three months "without results" | Foundation for everything else; suite 40 → 12 min |
| 1 | Sep–Nov 2025 | catalog-service |
Strangler of /products, HttpProductsRepository + REMOTE_CATALOG, initial load + events + shadow, HPA, Redis |
02-02, 02-04, 04-02, 05-04, 06-04 | Data discrepancies | Black Friday without an outage (0 min vs 50); campaign cost ÷3 |
| 2 | Dec 2025–Jan 2026 | inventory-service |
Reservations with expires_at, REMOTE_INVENTORY, first consumer, first DLQ, KEDA; outbox/idempotency into common-http |
02-04, 03-01, 03-02, 06-04 | Selling what is not there | January sales ×3.5 without incident |
| 3 | Feb 2026 | notifications-service |
Consumer of order.confirmed/cancelled, retries with TTL + DLQ |
03-02, 06-03 | Duplicate/lost emails | Memory leak isolated; createOrder −300 ms |
| 4 | Mar–Apr 2026 | payments-service |
Tokenization, HMAC webhook, audit, UNIQUE(order_id), canary with review |
05-04, 06-03, 07-02, 07-03 | Duplicate charges / PCI | Zero duplicate charges; minimal PCI scope |
| 5 | May–Jul 2026 | orders-service |
Choreographed saga, outbox, 202 + ETag, watchdog, reconciliation, ASYNC_SAGA |
02-05, 03-01, 04-04, 06-03, 06-05 | Stuck sagas | Orders deployable daily; INC-2031 (40 min) and its postmortem |
| 6 | Jul–Aug 2026 | customers-service + shutdown |
Keycloak + customerId, customers_ref, GDPR, read DB for BI |
02-04, 07-01, 07-03 | Personal data | Monolith at 0 replicas on 2026-08-07 |
- Metrics before and after
The four DORA metrics from 05-03 §12 measured in August 2026 (average across the six services), plus the business and organization ones Marta presented to management:
| Metric | Before (June 2025) | After (August 2026) | How it is measured |
|---|---|---|---|
| Deployment frequency | 1 every 1–2 weeks (Thursday night), the whole store | 11 a day in total (Catalog 3–4, Orders 2–3, the rest 1–2) | Argo CD syncs per service |
| Change lead time | 9 days median | 4 hours median (23 min for hotfixes) | Commit date → sync date |
| Change failure rate | 33% (4/12 rollbacks) | 6% (5 reverts/rollbacks out of 82 deployments in July) | Incidents or rollbacks / deployments |
| MTTR | 50 min (Black Friday); ~2 h median on failed deployments | 9 min median (argocd app rollback or canary-weight=0) |
From the alert to closure |
| Availability of "create order" | 99.5% monthly, estimated (not measured) | 99.93% (SLO 99.9%, 06-05) | SLI slo:orders_availability |
| Catalog p99 latency during a campaign | > 8 s (saturation) | 280 ms | Latency SLI |
| Coordination time (Orders team) | ~20% | ~7% (contract reviews and promotion PRs) | Quarterly survey + time on cross-team PRs |
| Test suite | 40 min (everything) | 4–7 min per service | CI |
| Campaign infrastructure cost | ×3.3 (ten monolith servers) | ×1.4 (only Catalog and Inventory scale) | Provider invoice |
Two honest caveats Marta added to the table: the availability "before" is an estimate (there was no SLI, which is a data point in itself), and the 6% failure rate includes INC-2031, which was the worst incident of the year and still lasted less than the best Thursday deployment.
- Real costs and what went wrong
What it cost. A year of migration is not free, and it is worth having the figures (fictional, but of a realistic order of magnitude) for the exercise in 08-04:
| Cost | Magnitude | Comment |
|---|---|---|
| Team time | ~40% of development capacity for 13 months (about 10 full-time equivalents) | Half of it in phase 5. The rest kept delivering features; the business did not stop. |
| Infrastructure | +35% the first year (cluster, RabbitMQ, observability, Keycloak, staging environments); −20% versus the previous year during campaigns | With the catalog scaled separately, the net balance at year end was slightly positive thanks to the campaign savings. The detailed monthly cost is in 08-03. |
| Learning curve | Kubernetes, RabbitMQ, OpenTelemetry, Keycloak, Pact: each cost someone 2–4 weeks | Mitigated by Platform (templates, reusable workflow) and by doing it in order. |
| Paid tools | Private registry, managed Pact Broker, secrets manager | Minor; each was justified on its own. |
What went wrong. Three stumbles the team tells without embarrassment because they are the ones that teach the most:
- The dual write that was attempted (phase 1, October 2025). The first version of the catalog synchronization wrote the
productstable and the MongoDB document in the same monolith function, "so we don't have to set up RabbitMQ yet". Within a week the shadow comparison found 40 divergent products: a MongoDB timeout, an exception between the two writes, and no transaction covering them (02-04 §6). The code was thrown away, RabbitMQ was brought forward andproduct.updatedwas published from the source of truth. The lesson became a written rule: one write, in the owner; the copy, through events. discounts-service, the nanoservice (phase 4, April 2026). Marketing asked for coupons for the spring sales and the Shopping Experience team, eager to try out the template, created a service with acouponstable and a single endpointPOST /v1/discounts/calculate, called synchronously bycreateOrderon every order. It added 40 ms to every order, one more deployment, one more dependency on the critical path and no autonomous business capability (02-01: "can it be described in one sentence without 'and'?" yes, but "does it have its own data and lifecycle?" no). In phase 5 it was reabsorbed intoorders-serviceas thedomain/discounts.jsmodule with the table inside theordersDB, exactly what exercise 2 of 02-02 had recommended. If promotions ever grow (campaigns, rules, segments), they will come out again; today they are twelve rules and one table.- The cardinality that took down Prometheus (phase 3, February 2026). A Notifications consumer labeled
http_requests_totalwith the untemplated route (/v1/orders/ord-3f9a1c2binstead of/v1/orders/{id}) and another one addedorderIdas a label on a counter. In five days Prometheus went from 60,000 to 2.4 million series and the pod died from memory; for 25 minutes there were no metrics and no alerts. The fix (06-01 §11): labels only with bounded values,routefromreq.route.path, a metrics linting rule incommon-httpand asample_limiton the scrape. It is why theorders_cancelled_totalmetric carriesreason(four values) and neverorderId.
And a fourth, minor but frequent: for two months the 80-line ci.yml lived copied in three repositories with three small differences; the reusable workflow from 05-03 §11 arrived "before the fourth", not "before the second". Luis's rule gets broken sometimes too.
- The final architecture
This is how TechCorp stands in August 2026, with the monolith shut down. Compared with the diagram in 01-05 §7 (the target) there are three differences: there is no Istio (05-05: evaluated and postponed; mTLS with Linkerd when the Payments audit asks for it), there is a read DB for analytics, and Redis, Keycloak and the observability stack have their own names.
flowchart TB
subgraph External
WEB[web-store / bff-mobile :3010]
PSP[Payment provider]
SMTP[Email provider]
end
WEB -->|HTTPS + JWT| ING[Ingress api.techcorp.example<br/>cert-manager]
ING --> GW[API Gateway :8080<br/>requireToken, rate limit]
KC[Keycloak realm techcorp] -. JWKS .-> GW
GW -->|/api/v1/products| CAT[catalog-service :3001<br/>HPA 2-20 + Redis]
GW -->|/api/v1/orders| ORD[orders-service :3002<br/>canary]
GW -->|/api/v1/customers| CUS[customers-service :3004]
PSP -->|HMAC webhook| PAY
CAT --> MDB[(MongoDB catalog)]
ORD --> PGO[(PG orders)]
CUS --> PGC[(PG customers)]
INV[inventory-service :3006<br/>KEDA 2-10] --> PGI[(PG inventory)]
PAY[payments-service :3003<br/>canary] --> PGP[(PG payments)]
NOT[notifications-service :3005] --> SMTP
PAY --> PSP
ORD -->|GET products / customers| CAT
ORD --> CUS
ORD <-.-> MQ[(RabbitMQ techcorp.events<br/>amqps, vhost techcorp)]
MQ <-.-> INV
MQ <-.-> PAY
MQ -.-> NOT
MQ -.-> ANA[analytics consumer] --> PGA[(PG techcorp-analytics<br/>reports / BI)]
CUS -.->|customer.updated / deleted| MQ
subgraph Platform
OBS[Prometheus · Grafana · Loki · otel-collector · Jaeger]
ARGO[Argo CD ← techcorp/platform]
VAULT[Vault + ESO]
end
What is not in the diagram and matters as much as what is: the deny-all NetworkPolicies from 07-04 (only Orders talks to Inventory; only Payments goes out to the Internet), the SLOs and the eleven alerts from 06-05, the runbooks in techcorp/platform/runbooks/, and four teams that deploy without asking anyone's permission.
Common Mistakes and Tips
- Skipping phase 0. It is the strongest temptation ("let's start with something visible") and the one that costs the most: without a gateway there is no strangler, without CI there is no confidence, without observability there is no diagnosis. Three months of prerequisites saved a year.
- Extracting the core first. Orders was fifth because it depended on everyone; extracting it first would have produced a distributed monolith with five calls to the monolith per order.
- Deleting the old thing on cutover day. The
productstable kept read-only for six weeks and the monolith at zero replicas for two avoided two guaranteed incidents (the monthly report and a forgotten integration). - Measuring only at the end. The availability "before" is an estimate because it was not measured. Instrument the monolith before touching it: it is your baseline and your argument to management.
- Confusing a stumble with a failure. The dual write, the nanoservice and the cardinality cost days, not months, because each was detected with a tool that already existed (shadow, RED dashboard, Prometheus-down alert). The goal is not to never be wrong: it is to be wrong cheaply and to learn in writing.
- Tip: keep a chronological table like the one in section 9 from day one, with the "main risk" column filled in before starting each phase. It is the best communication tool with the business and the best material for the project postmortem.
Exercises
Exercise 1: Reorder under a different constraint
Imagine that TechCorp had had, in July 2025, a mandatory PCI audit in January 2026. Would the extraction order change? Propose an alternative order, indicate which phase is brought forward and at what cost, and which part of phase 0 would become essential before that phase.
Exercise 2: Diagnose a stumble
A team reports that, during their migration of a catalog, "the new service's data drifted out of sync every now and then and nobody knew why". With what you learned in phase 1 and in section 11: list the three most likely causes, the tool that would have detected them and the technique that avoids them.
Exercise 3: Read the metrics table
Using the table in section 10: (a) which metric proves that phase 1 solved problem 2 from 01-05? (b) why is the 6% failure rate better even though in absolute number of failed deployments (5 in July) it is higher than before (4 in a quarter)? (c) Which metric in the table is the least reliable, and why?
Solutions
Exercise 1. Yes: Payments would move from fourth to second place (after Catalog), so that by January the payment provider, tokenization, the HMAC webhook, the audit and Payments' securityContext/NetworkPolicies were isolated in a service with minimal PCI scope. Cost: Payments would be extracted with branch by abstraction from the monolith (as in the real phase 4) but before Inventory existed, so it would keep charging through a synchronous call from createOrder for longer, and its stock.reserved consumer would be written later (double work in the adapter). Moreover, the team would debut canary and human review with the most delicate service, instead of practicing with Catalog. From phase 0, the following would become essential before Payments: removing the write to payments from createOrder (payments.recordCharge()), secrets management (07-04: the payments-provider Secret cannot live in a YAML) and the "only Payments goes out to the Internet" NetworkPolicies.
Exercise 2. (1) Dual write: the monolith writes to both sources with no shared transaction; an exception or timeout between them leaves divergent copies. Detected by the shadow comparison; avoided by "one write in the owner + events" (02-04 §6). (2) Lost events: the producer publishes without an outbox (or the consumer acks before writing) and a restart loses messages. Detected by the shadow and the outbox_pending/un-acked messages metric; avoided by the transactional outbox (02-05 §7) and ack after processing (03-02). (3) Out-of-order or duplicate events that overwrite a new value with an old one. Detected by the shadow (intermittent discrepancies that "fix themselves" on the next change); avoided by processOnce + comparison by updatedAt (04-04 ex. 2). Bonus: a non-idempotent initial load script that was rerun halfway (duplicates or stale values); avoided by upsert (02-04).
Exercise 3. (a) "Catalog p99 latency during a campaign" (> 8 s → 280 ms) and above all the business figure from section 3: 0 minutes without selling on 2025-11-28 versus 50; the campaign cost ×3.3 → ×1.4 confirms that only what saturates gets scaled. (b) Because the rate is measured per deployment and the number of deployments has multiplied by more than 20: 5 failures out of 82 (6%) versus 4 out of 12 (33%). Moreover, each failure affects one service, is detected in a 10% canary and is reverted in minutes (MTTR 9 min), versus reverting the whole store after two hours. It is the DORA relationship from 05-03: more, smaller deployments → fewer failures and lower MTTR. (c) The "before" availability (99.5%): it is a retrospective estimate, because the monolith had no SLI; and to a lesser extent the coordination time, which is based on surveys. The table itself warns about it: without prior measurement there is no reliable baseline, and that is a lesson in itself.
Conclusion
TechCorp's migration was not a rewrite but thirteen months of ordered extractions on top of a phase 0 of prerequisites that produced no service and made everything possible: four teams owning their areas, CI and containers, a cluster with the gateway in front as the strangler facade, minimal observability and a modular monolith in which createOrder stopped writing to stock and payments. After that, each phase solved a concrete problem from 01-05 with the techniques of the previous modules: Catalog (initial load, events, shadow, flag, HPA) and the first Black Friday without an outage; Inventory (reservations with expiry, first consumer, first DLQ, KEDA); Notifications (email out of the process, the end of the leak that took down charging); Payments (tokenization, HMAC webhook, audit, canary with review); Orders (the choreographed saga with outbox and watchdog instead of the transaction, the 202, and INC-2031 as the teacher); Customers (Keycloak, customerId, GDPR) and the shutdown of the monolith with analytics moved to a read database. The metrics tell the outcome —from a fortnightly deployment to eleven a day, from 9 days to 4 hours of lead time, from 33% to 6% failures, from 50 to 9 minutes of recovery, 99.93% availability for "create order"— and the stumbles (the dual write, discounts-service, the cardinality) tell what it cost to learn it.
This account has referred constantly to code that already exists (catalog-service from 04-02, orders-service from 04-04, the gateway from 03-04 and 07-01) and to four services we have described but never written: Inventory, Payments, Notifications and Customers. The next lesson implements them in a condensed but complete way, with the same template and the same library, closes the event map and follows order ord-88213 through the six real services, database by database and queue by queue.
Microservices Course
Module 1: Introduction to Microservices
- Basic Concepts of Microservices
- Advantages and Disadvantages of Microservices
- Comparison with the Monolithic Architecture
- When to Adopt Microservices: Decision Criteria
- The Course Case Study: TechCorp's Online Store
Module 2: Microservice Design
- Microservice Design Principles
- Decomposing Monolithic Applications
- Defining Bounded Contexts
- Data Management: One Database per Service
- Distributed Consistency: Sagas, CQRS and Event Sourcing
Module 3: Communication between Microservices
- RESTful APIs
- Asynchronous Messaging
- Communication Protocols: gRPC, GraphQL
- API Gateway and Backend for Frontend
- Service Discovery and Load Balancing
- API Contracts and Versioning
Module 4: Implementing Microservices
- Choosing Technologies and Tools
- Building a Simple Microservice
- Configuration Management
- Hands-On Integration: Consuming APIs and Publishing Events
- Testing Microservices: Unit, Integration and Contract Tests
Module 5: Deployment and Orchestration
- Containers and Docker
- Orchestration with Kubernetes
- CI/CD for Microservices
- Deployment Strategies: Rolling, Blue-Green and Canary
- Service Mesh: Istio and Linkerd
Module 6: Monitoring and Maintenance
- Monitoring and Logging
- Distributed Tracing with OpenTelemetry
- Error Handling and Recovery
- Scalability and Performance
- SLOs, Alerts and Incident Management
Module 7: Security in Microservices
- Authentication and Authorization
- Communication Security
- Security Practices
- Container and Kubernetes Security
