catalog-service already boots, but its src/config.js is a minimal version that reads four variables and trusts that they are right. In a system with seven services, three environments (development, staging, production) and several replicas per service, configuration stops being a detail: it is what makes the same artifact (the same image, the same code) behave differently in each place without recompiling anything, and it is also the most common route by which passwords leak. This lesson pins down how a TechCorp service obtains, validates and uses its configuration: what is configuration and what is not, where it comes from and with what precedence, a config.js module with fail-fast validation that every service will reuse (we write it for orders-service, which needs it in full in 04-04), how secrets are handled, when centralized configuration pays off, and what feature flags are and how they are read.
Contents
- Configuration vs. code: the twelve-factor principle
- Kinds of configuration: per environment, per instance and feature flags
- Sources and precedence
src/config.js: loading and validating at startup.envin development and secrets in every other environment- Centralized configuration: when it pays off
- Hot reload vs. restart
- Feature flags with a minimal example
- TechCorp's configuration per environment and how it hooks into Kubernetes
- Configuration vs. code: the twelve-factor principle
The rule from the third factor of the Twelve-Factor App methodology is the one we follow: configuration lives in the environment, not in the code. The practical test to tell them apart: does this value change between development, staging and production, or between two installations of the same service? If yes, it is configuration; if not, it is code.
| It is configuration (varies per environment) | It is code (does not vary) |
|---|---|
URL of the database, the broker, other services (ORDERS_DB_URL, RABBITMQ_URL, CATALOG_URL) |
The name of the techcorp.events exchange and the orders.saga queue (they are part of the 03-02 contract, identical in every environment) |
| Listening port, log level, timeouts | The order state machine, the validation rules |
| Credentials, API keys, certificates | The error codes (CUSTOMER_NOT_FOUND) and the routes (/v1/orders) |
Feature flags (REMOTE_CATALOG, PAYMENT_NEW_PROVIDER) |
The 100-id limit per batch (contract rule, 03-01) |
The corollary that is hardest to internalize: the same Docker image runs in all three environments. If going to production requires rebuilding with a different properties file inside, what gets deployed is not what was tested in staging. The image is built once (05-01, 05-03) and the configuration is injected at startup.
- Kinds of configuration: per environment, per instance and feature flags
| Kind | Examples at TechCorp | Who changes it and how often |
|---|---|---|
| Per environment | URLs, credentials, LOG_LEVEL, HTTP_TIMEOUT_MS |
Platform, when creating or changing the environment; rarely |
| Per instance | PORT (locally, if two services run at once), HOSTNAME/pod name for the logs, number of workers |
The orchestrator, on each replica; the service reads it, it does not decide it |
| Feature flags | REMOTE_CATALOG (read products from the new service or from the monolith, 02-04), PAYMENT_NEW_PROVIDER |
The product/development teams, frequently and without deploying |
All three kinds come in through the same door (environment variables) at first; the difference lies in the rate of change, and that is what later justifies a separate flags service (section 8).
- Sources and precedence
A service can usually receive its configuration from several places at once. For the behavior to be predictable you must set a precedence and stick to it:
flowchart LR
A[1. Default values<br/>in config.js] --> B[2. .env file<br/>development only]
B --> C[3. Process<br/>environment variables]
C --> D[4. Secrets mounted<br/>as file or variable]
D --> E[Validated, immutable<br/>config object]
style E fill:#dfd,stroke:#393
- Default values: only for what is safe in any environment (
PORT=3002,LOG_LEVEL=info,HTTP_TIMEOUT_MS=2000). Never a default database URL pointing at somewhere real, and never a credential. .envfile: development convenience (04-02). It does not exist in production.- Environment variables: the universal interface. Docker, Docker Compose, Kubernetes, systemd, GitHub Actions... they all know how to set them.
- Secrets: technically they also arrive as variables or as mounted files; we distinguish them because their lifecycle (who sees them, how they rotate) is different (section 5).
What is not on the list: command-line arguments (they mix configuration with startup) and per-environment configuration files inside the image (config/production.json), which violate section 1.
src/config.js: loading and validating at startup
src/config.js: loading and validating at startupThe goal is to fail fast and with a clear message if something is missing or has the wrong type, instead of starting up and discovering at three in the morning that HTTP_TIMEOUT_MS was the string "2s" and fetch interpreted it as NaN. We use zod, already present in the service since 04-02. This is the full config.js for orders-service, with the variables agreed in 03-05:
// src/config.js (orders-service)
const { z } = require('zod');
// 1. The schema: name, type, whether required and default value of EVERY variable the service uses.
// It is also its documentation: if it is not here, the service does not read it.
const schema = z.object({
NODE_ENV: z.enum(['development', 'test', 'production']).default('development'),
PORT: z.coerce.number().int().min(1).max(65535).default(3002), // coerce: environment variables are always text
LOG_LEVEL: z.enum(['fatal', 'error', 'warn', 'info', 'debug', 'trace']).default('info'),
// Own dependencies (no default: required)
ORDERS_DB_URL: z.string().url().startsWith('postgres'), // postgres://user:password@host:5432/orders
RABBITMQ_URL: z.string().url().startsWith('amqp'), // amqp://rabbitmq:5672
// Services Orders calls synchronously (03-01, 03-05)
CATALOG_URL: z.string().url(), // http://catalog-service:3001
CUSTOMERS_URL: z.string().url(), // http://customers-service:3004
HTTP_TIMEOUT_MS: z.coerce.number().int().min(100).max(30000).default(2000),
// Outbox relay (04-04)
OUTBOX_INTERVAL_MS: z.coerce.number().int().min(50).default(500),
// Feature flag inherited from the branch by abstraction of 02-04:
// true = read products from catalog-service; false = from the monolith (via CATALOG_URL pointing at it)
REMOTE_CATALOG: z.enum(['true', 'false']).default('true').transform((v) => v === 'true')
});
// 2. Load and validate. safeParse does not throw: it lets us build a readable message.
function loadConfig(env = process.env) {
const result = schema.safeParse(env);
if (!result.success) {
const problems = result.error.issues.map((i) => ` - ${i.path.join('.')}: ${i.message}`).join('\n');
// Fail-fast: without valid configuration there is no service. Kubernetes will show the pod in CrashLoopBackOff with this message.
throw new Error(`Invalid configuration:\n${problems}`);
}
// 3. Immutable: nobody can change config.PORT at runtime "just to try"
return Object.freeze(result.data);
}
// 4. Loaded ONCE when the module is imported. Tests call loadConfig() with their own object.
const config = loadConfig();
// 5. Safe version for logs and for /health: secrets masked
function configForLog(c = config) {
const mask = (url) => url.replace(/\/\/([^:]+):([^@]+)@/, '//$1:***@'); // postgres://svc:***@host/orders
return { ...c, ORDERS_DB_URL: mask(c.ORDERS_DB_URL), RABBITMQ_URL: mask(c.RABBITMQ_URL) };
}
module.exports = { config, loadConfig, configForLog };How it is used from server.js (and only from there and from the infrastructure modules: routes and domain never read process.env):
// src/server.js (orders-service, excerpt)
const { config, configForLog } = require('./config'); // if the config is invalid, this require throws and the process dies right here
const logger = createLogger({ service: 'orders-service', level: config.LOG_LEVEL });
logger.info({ config: configForLog() }, 'configuration loaded');
const pool = createPostgresPool({ url: config.ORDERS_DB_URL });
const catalogClient = createCatalogClient({ baseUrl: config.CATALOG_URL, timeoutMs: config.HTTP_TIMEOUT_MS });And what happens if someone deploys without RABBITMQ_URL and with an impossible port:
Error: Invalid configuration: - RABBITMQ_URL: Required - PORT: Number must be less than or equal to 65535
Five decisions worth understanding:
z.coercebecause all environment variables are strings:PORT="3002"must become the number3002, andREMOTE_CATALOG="true"the booleantrue. Without explicit conversion,if (process.env.REMOTE_CATALOG)is truthy for"false"too.- Required with no default for everything that points at something real. A default
ORDERS_DB_URLpointing atlocalhost"works" on the laptop and starts up in production pointing nowhere, with a confusing error minutes later. - Validation at startup, not on first use: a pod that does not start (and that Kubernetes marks as failed immediately) is preferable to one that starts, passes readiness and fails on the first order.
Object.freezeso the configuration is a value, not global mutable state.- Dependencies receive values, they do not read
process.env:createCatalogClient({ baseUrl, timeoutMs }). That is what allows, in 04-05, testing the client against a fake server on another port without manipulating the process environment.
.env in development and secrets in every other environment
.env in development and secrets in every other environmentIn development, the .env file (loaded by dotenv in npm run dev, 04-02) holds the local values. Two fixed rules:
.envis in.gitignore. Always. Even if it "only has development passwords": what isdev-orderstoday is a copy of the production one pasted in a hurry tomorrow..env.exampleis versioned: same keys, sample or empty values, comments. It is the living documentation of what the service needs and the first thing a new developer copies.
# .env.example (orders-service) — copy to .env and fill in. NEVER put real values here.
NODE_ENV=development
PORT=3002
LOG_LEVEL=debug
ORDERS_DB_URL=postgres://svc_orders:dev-orders@localhost:5432/orders
RABBITMQ_URL=amqp://localhost:5672
CATALOG_URL=http://localhost:3001
CUSTOMERS_URL=http://localhost:3004
HTTP_TIMEOUT_MS=2000
OUTBOX_INTERVAL_MS=500
REMOTE_CATALOG=trueIn staging and production there is no .env. Non-sensitive values are set by the orchestrator as environment variables; secrets (DB passwords, broker credentials, payment provider keys) follow three rules:
| Rule | What it means in practice |
|---|---|
| Never in the repository | Not in .env, not in plain-text Kubernetes YAML, not in git history (a secret that was pushed and deleted is still in the history: it must be rotated) |
| Never in the image | No COPY .env, no ENV PASSWORD=... in the Dockerfile; anyone with access to the image registry would read it |
| Injected by the environment at runtime | The service receives them as an environment variable or as a file mounted at a known path; it neither knows nor cares where they come from |
That last point is the interface between the service and secrets management, and it is all an Orders developer needs to know: ORDERS_DB_URL will arrive as an environment variable. Whoever sets it may be a Kubernetes Secret (05-02), Vault with pod injection, or the cloud's secrets manager (07-04 goes deeper). If in some case the secret arrives as a file (common practice with Vault: /vault/secrets/orders-db), config.js is extended with a simple rule: if ORDERS_DB_URL_FILE exists, the file is read and its contents become ORDERS_DB_URL. That <VARIABLE>_FILE convention is the one used by the official PostgreSQL and RabbitMQ images, and TechCorp's template adopts it.
- Centralized configuration: when it pays off
When there are dozens of services and many shared values (for example, the broker URL, which is the same for everyone), some organizations set up a configuration server from which each service reads at startup (or to which it subscribes):
| Tool | Model | Strength | Cost |
|---|---|---|---|
| Spring Cloud Config | HTTP server serving properties from git | Perfect integration with Spring; history in git | Only makes sense in Java/Spring ecosystems |
| Consul KV | Distributed key-value store with watch | Hot changes; combines with the discovery from 03-05 | One more piece to operate and secure |
| etcd | Strongly consistent key-value | It is what Kubernetes itself uses | Rarely used directly from applications |
| AWS Parameter Store / Secrets Manager, Azure App Configuration, GCP Secret Manager | Managed service | Nothing to operate; auditing and rotation included | Cloud coupling; latency and read quotas |
| Kubernetes ConfigMap + Secret | Cluster objects injected as variables or files | Already there; declarative; per namespace | No reload (section 7); not a "real" secrets manager without additional encryption |
When a configuration server pays off: when values change frequently and in many services at once (thresholds, feature flags, rate limits), or when the same value must stay in sync in dozens of places and editing twenty ConfigMaps is a risk. When it does not: with seven services, three environments and values that change a few times a year, it is one more piece that can go down and that every service needs at startup (if the configuration server does not respond, nothing starts).
TechCorp's decision: environment variables as the service's single interface; in Kubernetes, ConfigMap for non-sensitive values and Secret for secrets (05-02). Feature flags, if they grow, will go to a flags service (section 8), not to a general configuration server.
- Hot reload vs. restart
A configuration server or a ConfigMap mounted as a file lets a service detect changes and apply them without restarting. It sounds attractive, but it has a cost:
| Aspect | Hot reload | Restart (new configuration = new deployment) |
|---|---|---|
| Code complexity | High: you must manage which parts to reload (can ORDERS_DB_URL change with an open pool?), atomicity, half-applied values |
None: the process starts with the new config, period |
| Traceability | Hard to know which configuration each replica has at any moment | Every deployment is recorded; all replicas converge |
| Validation | You must validate on the fly and decide what to do if the new value is invalid | The fail-fast from section 4: if invalid, the new pod does not start and the old one keeps running |
| Disruption | None | None, if the deployment is rolling (05-04): replicas are replaced one at a time |
That is why in Kubernetes the usual practice is to restart: a configuration change is just another rollout, with the same safety and the same rollback as a code change. The exception is feature flags, whose value must change in seconds and without a deployment: for them, and only for them, dynamic reads are accepted.
- Feature flags with a minimal example
A feature flag is a configuration value that enables or disables a behavior at runtime. It lets you deploy code switched off, turn it on gradually and turn it off in seconds if something goes wrong (05-04 will relate it to canary deployments). We already used one in 02-04, REMOTE_CATALOG, and here we add the one Payments will need: PAYMENT_NEW_PROVIDER.
Minimal version, read from configuration (changing it requires a restart):
// src/flags.js (payments-service) — version 1: flags as configuration
function createFlags(config) {
return {
// Returns a function so the rest of the code is not coupled to "where the flag comes from"
isActive: (name) => config.FLAGS[name] === true
};
}
// Usage in the charge use case (Payments): new and old code coexist; the flag decides
async function charge(order, { flags, currentProvider, newProvider }) {
const provider = flags.isActive('PAYMENT_NEW_PROVIDER') ? newProvider : currentProvider;
return provider.charge({ amount: order.total, idempotencyKey: order.orderId }); // 02-05
}When flags multiply, change several times a day or need to be enabled only for a percentage of users, you move to a flags service (Unleash, Flagsmith, LaunchDarkly, or GitLab's flags module). The service queries the state via SDK, with a local cache and a default value if the flags service does not respond. The important thing is that the isActive(name) interface does not change: only the implementation of createFlags changes, and the charge use case never notices:
// src/flags.js — version 2: flags from Unleash (concept; real SDK: unleash-client)
function createFlags({ unleashClient, defaults = {} }) {
return { isActive: (name, context) => unleashClient.isEnabled(name, context, defaults[name] ?? false) };
}Two warnings: flags must have an expiration date (a flag that has been true for a year is dead code in disguise), and their default value in production must be the safe behavior (the current payment provider, not the new one).
- TechCorp's configuration per environment and how it hooks into Kubernetes
With all of the above, the table the Platform team maintains for orders-service (fictitious values; in production the secrets are not even in this table, only the name of the Secret that holds them):
| Variable | Development (.env) |
Staging | Production | Source in Kubernetes |
|---|---|---|---|---|
NODE_ENV |
development |
production |
production |
ConfigMap |
PORT |
3002 |
3002 |
3002 |
ConfigMap |
LOG_LEVEL |
debug |
info |
info |
ConfigMap |
ORDERS_DB_URL |
postgres://svc_orders:dev-orders@localhost:5432/orders |
postgres://svc_orders:stg-Xk3…@pg-staging:5432/orders |
(in the orders-db Secret) |
Secret |
RABBITMQ_URL |
amqp://localhost:5672 |
amqp://orders:stg-Rq7…@rabbitmq:5672 |
(in the orders-rabbitmq Secret) |
Secret |
CATALOG_URL |
http://localhost:3001 |
http://catalog-service:3001 |
http://catalog-service:3001 |
ConfigMap |
CUSTOMERS_URL |
http://localhost:3004 |
http://customers-service:3004 |
http://customers-service:3004 |
ConfigMap |
HTTP_TIMEOUT_MS |
2000 |
2000 |
2000 |
ConfigMap |
OUTBOX_INTERVAL_MS |
500 |
500 |
250 |
ConfigMap |
REMOTE_CATALOG |
true |
true |
false → true during the migration |
ConfigMap (or flags service) |
Notice that CATALOG_URL has the same value in staging and production: it is the DNS name of the Kubernetes Service (03-05), resolved within each namespace. Only the credentials and a performance tweak or two change.
How this connects to Kubernetes, without writing the YAML yet (05-02): a ConfigMap called orders-service-config holds the non-sensitive keys; a Secret called orders-db holds ORDERS_DB_URL; the Orders Deployment declares "inject every key of that ConfigMap and that Secret as environment variables" (envFrom). From inside the container, process.env.ORDERS_DB_URL exists and config.js validates it exactly as it does on the laptop. The service does not distinguish the origin, and that indifference is the goal of this whole lesson.
Common Mistakes and Tips
- Configuration in the code (
const CATALOG_URL = 'http://catalog-service:3001'). It works until you have to point somewhere else in tests or in a second environment. Everything that varies, through an environment variable; everything, throughconfig.js. - Dangerous defaults:
ORDERS_DB_URLdefaulting tolocalhost,NODE_ENVdefaulting todevelopmentwith verbose logs and stack traces sent to the client in production. Only harmless defaults. process.envscattered across the code. Everyprocess.env.Xoutsideconfig.jsis a variable that is unvalidated, undocumented and without a clear default. A single entry point.- Secrets in the logs.
logger.info({ config })with the full PostgreSQL URL sends the password to Loki.configForLog()always. - Secrets in git "just this once." Mandatory rotation; deleting them from the last commit is not enough.
- Booleans as strings.
if (process.env.REMOTE_CATALOG)is always truthy.transform((v) => v === 'true'). - Hot-reloading configuration just in case. Complexity without need; a
rolloutis safer and auditable. Dynamic only for flags. - Eternal flags. Every flag with an owner, a purpose and a retirement date in its own definition.
Exercises
Exercise 1. Write the zod schema for src/config.js of catalog-service (04-02) with the variables PORT (default 3001), MONGO_URL (required, must start with mongodb), MONGO_DB (default catalog), LOG_LEVEL and NODE_ENV. Add the variable CACHE_MAX_AGE_S (seconds of the Cache-Control, integer between 0 and 3600, default 30) and explain in one sentence why this one is configuration and not code.
Exercise 2. An Orders developer proposes that, if CATALOG_URL is not defined, the service use http://catalog-service:3001 by default "because it is always that." Give one argument for, two against, and decide.
Exercise 3. Implement the <VARIABLE>_FILE convention from section 5 in loadConfig: before validating, for each variable in the schema, if NAME_FILE exists and NAME does not, read the file (synchronously, UTF-8, without the trailing newline) and use its contents as the value.
Solutions
Solution 1.
const schema = z.object({
NODE_ENV: z.enum(['development', 'test', 'production']).default('development'),
PORT: z.coerce.number().int().min(1).max(65535).default(3001),
LOG_LEVEL: z.enum(['fatal', 'error', 'warn', 'info', 'debug', 'trace']).default('info'),
MONGO_URL: z.string().url().startsWith('mongodb'), // required: no default
MONGO_DB: z.string().min(1).default('catalog'),
CACHE_MAX_AGE_S: z.coerce.number().int().min(0).max(3600).default(30)
});CACHE_MAX_AGE_S is configuration because its optimal value depends on the environment and the moment (during a campaign with frequent price changes it would be lowered to 5 s without deploying code); the fact that the batch carries Cache-Control is contract and therefore code.
Solution 2. For: convenience; in Kubernetes the name is stable (03-05) and it avoids one variable in every ConfigMap. Against: (1) in local development that name does not resolve, and the error ("ENOTFOUND catalog-service") shows up on the first request, not at startup; (2) during the migration with REMOTE_CATALOG=false, CATALOG_URL must point at the monolith, and a "correct" default hides that someone forgot to configure it in an environment. Decision: required, no default; the ConfigMap carries it explicitly. One more line of YAML in exchange for a clear failure at startup.
Solution 3.
const fs = require('node:fs');
function resolveSecretFiles(env) {
const result = { ...env };
for (const name of Object.keys(schema.shape)) { // only the variables the schema knows
const filePath = env[`${name}_FILE`];
if (filePath && result[name] === undefined) {
result[name] = fs.readFileSync(filePath, 'utf8').replace(/\r?\n$/, '');
}
}
return result;
}
function loadConfig(env = process.env) {
const result = schema.safeParse(resolveSecretFiles(env));
// ... same as before
}With this, mounting the Vault secret at /vault/secrets/orders-db and defining ORDERS_DB_URL_FILE=/vault/secrets/orders-db works without touching the rest of the service; and if the file does not exist, readFileSync throws at startup, which is what we want.
Conclusion
The configuration of a TechCorp microservice ends up like this: it lives in the environment, not in the code or the image; it arrives through environment variables (or mounted files, with the _FILE convention) with a clear precedence (harmless defaults → .env only in development → environment → secrets); it enters through a single door, src/config.js, which validates it with zod at startup and fails fast with a readable message, freezes it and hands it out to dependencies as values; secrets never touch the repository or the image and the service does not know where they come from (Kubernetes Secret, Vault: 05-02 and 07-04); centralized configuration is reserved for when the volume of changes justifies it; changes are applied by restarting except for feature flags, which have their isActive(name) interface and, if they grow, their own service; and the per-environment table for orders-service fixes the concrete values of PORT, ORDERS_DB_URL, RABBITMQ_URL, CATALOG_URL, CUSTOMERS_URL, HTTP_TIMEOUT_MS, OUTBOX_INTERVAL_MS and REMOTE_CATALOG.
That config.js is not a theoretical exercise: it is the first file of orders-service, TechCorp's most complex service. In the next lesson we build the rest on top of the 04-02 template: the PostgreSQL pool and the migrations, the HTTP clients to Catalog and Customers with their timeout and their ACL, the createOrder use case with Idempotency-Key and outbox in the same transaction, the relay that publishes to RabbitMQ and the saga consumers that move the order from PENDING to CONFIRMED. It is the lesson in which everything designed since module 2 becomes code that runs end to end.
Microservices Course
Module 1: Introduction to Microservices
- Basic Concepts of Microservices
- Advantages and Disadvantages of Microservices
- Comparison with the Monolithic Architecture
- When to Adopt Microservices: Decision Criteria
- The Course Case Study: TechCorp's Online Store
Module 2: Microservice Design
- Microservice Design Principles
- Decomposing Monolithic Applications
- Defining Bounded Contexts
- Data Management: One Database per Service
- Distributed Consistency: Sagas, CQRS and Event Sourcing
Module 3: Communication between Microservices
- RESTful APIs
- Asynchronous Messaging
- Communication Protocols: gRPC, GraphQL
- API Gateway and Backend for Frontend
- Service Discovery and Load Balancing
- API Contracts and Versioning
Module 4: Implementing Microservices
- Choosing Technologies and Tools
- Building a Simple Microservice
- Configuration Management
- Hands-On Integration: Consuming APIs and Publishing Events
- Testing Microservices: Unit, Integration and Contract Tests
Module 5: Deployment and Orchestration
- Containers and Docker
- Orchestration with Kubernetes
- CI/CD for Microservices
- Deployment Strategies: Rolling, Blue-Green and Canary
- Service Mesh: Istio and Linkerd
Module 6: Monitoring and Maintenance
- Monitoring and Logging
- Distributed Tracing with OpenTelemetry
- Error Handling and Recovery
- Scalability and Performance
- SLOs, Alerts and Incident Management
Module 7: Security in Microservices
- Authentication and Authorization
- Communication Security
- Security Practices
- Container and Kubernetes Security
