An operating system is not a single category: the firmware governing a weather station and the Linux kernel of a data center server solve the same underlying problem, but with opposite priorities and a difference in size of four orders of magnitude. In this lesson you are going to learn to classify operating systems using independent criteria — processing mode, number of users and tasks, field of use and licensing model — and, above all, to use that classification for what it is really good for: choosing. At the end we will apply the criteria to the three pieces of Meteora's infrastructure and you will see that the right answer is different in each case.
Contents
- Why classify: the criteria are not mutually exclusive
- Classification by processing mode
- Classification by number of users and tasks
- Classification by field of use
- Classification by license and development model
- Comparison table and selection criteria
- Applied case: what Meteora chooses for each piece
Why classify: the criteria are not mutually exclusive
The first mistake when studying this topic is treating the classification as a list of boxes where each system falls into exactly one. It does not work like that. The criteria are independent dimensions, and any real system occupies a position on all of them at once.
Take Linux on meteo-01:
- By processing mode: interactive/time-sharing, although it also runs batch jobs overnight.
- By users: multiuser.
- By tasks: preemptive multitasking.
- By field: server, although the same kernel is used on desktops, on phones and in embedded systems.
- By license: free (GPL v2).
The very same Linux kernel, with a different configuration and different tools around it, can be an 8 MB single-user embedded system. That is why a useful classification does not describe "what" a system is, but what priorities it has and what it is tuned for.
Classification by processing mode
This criterion answers: how do jobs reach the system and what is optimized when serving them?
Batch systems
Jobs pile up and run with no user interaction. What is optimized is throughput: jobs completed per hour. It does not matter if one particular job takes longer, as long as more get processed overall.
They sound old-fashioned, but they are more alive than ever: a bank's monthly billing run, the overnight retraining of a model, report generation or Meteora's own daily average computation are all batch processing. What changed is the wrapper: today they are called jobs, cron, pipelines or batch workloads.
# /etc/cron.d/meteora-summary
30 2 * * * meteora /opt/meteora/bin/aggregator --day=yesterday --mode=fullAn explanation of the block, which is a textbook batch job:
30 2 * * *arecron's five fields: minute 30, hour 2, any day of the month, any month, any day of the week. That is, every day at 02:30.meteorais the user it runs as.rootis never used for this: if the program has a bug, the damage is limited to whatmeteoracan touch.- The rest is the program and its arguments.
- Nobody is watching: the job runs, writes its output and finishes. If it fails, we will find out from the logs. That absence of interactivity is exactly what defines batch mode.
Interactive or time-sharing systems
The system serves users (or clients) who are waiting for a response. What is optimized is not total throughput but response time and its predictability.
The technical consequence is that the scheduler must favor processes that do a lot of input/output and little computation, because those tend to be the interactive ones. When aggregator (heavy CPU) and meteo-api (lots of network waiting) coexist on meteo-01, the scheduler gives preference to the latter precisely for this reason. You will study it in CPU Scheduling.
Real-time systems
Here what matters is not being fast but meeting deadlines. A real-time system guarantees that a task completes before a limit instant (its deadline). The key distinction:
| Hard real time | Soft real time | |
|---|---|---|
| Consequence of missing the deadline | System failure, possible physical harm | Quality degradation |
| Examples | Airbag, flight control, pacemaker, industrial robot | Video call, video playback, digital audio |
| Guarantee required | Mathematically provable | Statistical ("99.9% of the time") |
| Typical scheduler | Fixed priorities or EDF, no heuristics | Priorities with adjustments |
A counterintuitive but fundamental idea: a real-time system is not a fast system, it is a predictable system. An RTOS usually has worse average performance than Linux, because it gives up optimizations (aggressive caches, adaptive schedulers) whose worst-case behavior cannot be bounded. It prefers to be equally slow all the time rather than normally very fast and occasionally unpredictable.
You can verify on meteo-01 that Linux offers soft real-time scheduling policies:
SCHED_OTHER min/max priority : 0/0 SCHED_FIFO min/max priority : 1/99 SCHED_RR min/max priority : 1/99 SCHED_BATCH min/max priority : 0/0 SCHED_IDLE min/max priority : 0/0 SCHED_DEADLINE min/max priority : 0/0
What this output means:
chrtqueries and modifies real-time scheduling policies. The-moption shows the available priority ranges.SCHED_OTHERis the normal policy, the one used byingestor,aggregatorandmeteo-api. It has no real-time priorities (range 0/0).SCHED_FIFOandSCHED_RRare soft real-time policies with 99 priority levels. ASCHED_FIFOprocess runs until it blocks or voluntarily yields: it can monopolize a core.SCHED_DEADLINElets you declare an explicit deadline, and the kernel checks that the set of tasks is admissible.
Even with these policies, standard Linux is not a hard real-time system: it does not guarantee an upper bound on latency in all cases. For that there are systems such as QNX, VxWorks, FreeRTOS or Zephyr, or patches like PREEMPT_RT. The full topic is covered in Mobile and Real-Time Operating Systems.
Classification by number of users and tasks
These are two different dimensions that are often confused.
Single-user and multiuser
A multiuser system keeps separate identities with their own resources and permissions, and guarantees that one user cannot access what belongs to another. It does not require several people to be connected at once: meteo-01 is multiuser even though nobody logs in, because each service runs with a different identity.
In fact, that is the most important use of the multiuser model today: not separating people, but separating services.
USER COMMAND meteora aggregator meteora ingestor meteora meteo-api root sshd root systemd systemd+ systemd-resolved www-data nginx
Interpretation:
- Meteora's three processes run as
meteora, not asroot. If an attacker compromisesmeteo-api, all they get aremeteora's permissions: they can read/var/lib/meteora/, but not modify the system. nginxruns aswww-data, another separate identity, with its own permissions.systemd-resolveduses a dedicated system identity.- Only what is strictly necessary (
systemd,sshd) runs asroot.
This is the principle of least privilege applied through the multiuser model, and we will develop it in Protection Principles and Access Control.
Single-tasking and multitasking
A single-tasking system runs one program at a time (MS-DOS, the simplest firmware of a microcontroller). A multitasking system keeps several running concurrently. And within multitasking there is a decisive distinction:
| Cooperative multitasking | Preemptive multitasking | |
|---|---|---|
| Who yields the CPU | The program itself, voluntarily | The operating system, through a clock interrupt |
| If a program enters an infinite loop | The whole machine hangs | Only that process is affected |
| Requires hardware with a timer | No | Yes |
| Examples | Windows 3.x, classic Mac OS | Linux, Windows NT and later, macOS, Android |
History has already passed judgment on this: every general-purpose system today is preemptive, because a system cannot depend on the good will of the programs it runs. If aggregator entered an infinite loop on a cooperative system, meteo-01 would stop responding entirely and would have to be rebooted physically.
Classification by field of use
This criterion is the most useful in practice, because it describes what the system is tuned for.
| Field | Main priority | Usual interface | Examples | Typical resources |
|---|---|---|---|---|
| Desktop | Perceived latency and ease of use | GUI | Windows 11, macOS, Ubuntu Desktop | 8-32 GB RAM |
| Server | Stability, sustained performance, security | Command line, remote | Debian, RHEL, Windows Server | 8 GB - 1 TB RAM |
| Embedded | Power draw, size, boot time, reliability | None or minimal | FreeRTOS, Zephyr, Yocto Linux | 64 KB - 512 MB RAM |
| Mobile | Battery, app isolation, touch | Touch GUI | Android, iOS | 4-16 GB RAM |
| Distributed | Location transparency, fault tolerance | API and orchestrator | Plan 9, Kubernetes-style layers | Many machines |
| Network | Serving shared resources | Remote administration | NetWare (historical), today's NAS | Variable |
Three of them are worth pausing on:
- Server. It carries no graphical interface, prioritizes sustained performance over the latency of a click, and accepts long, predictable update cycles. A server distribution such as Debian stable or RHEL offers 5 to 10 years of support with no incompatible changes, which is exactly the opposite of what a desktop user wants.
- Embedded. The system lives inside a device that is not perceived as a computer. The requirements change radically: boot in milliseconds, operation for years without rebooting, microamp consumption at rest and often the complete absence of a disk and of virtual memory management.
- Distributed. A genuine distributed operating system presents several machines as if they were one. It is an academically beautiful idea that is barely used in its pure form; in practice what is done is to put an orchestration layer on top of ordinary operating systems. You will see it in The Operating System in the Cloud.
Classification by license and development model
| Proprietary | Free / open source | |
|---|---|---|
| Access to the code | No, or very restricted | Yes, complete |
| License cost | Per machine, core or user | Zero (you pay for support if you want it) |
| Who fixes a bug | Only the vendor | Anyone, although in practice the community |
| Security auditing | Trust in the vendor | Verifiable by third parties |
| Dependency risk | High (vendor lock-in) | Low |
| Support | Contractual, with guarantees | Community-based, or commercially contracted |
| Examples | Windows, macOS, VxWorks, QNX | Linux, FreeBSD, OpenBSD, Zephyr, FreeRTOS |
Two clarifications that are almost always overlooked:
- Free does not mean free of charge in the relevant sense. Red Hat Enterprise Linux is free software and costs money: what you pay for is not the license but the support, the certification and the guaranteed updates. For a company, that difference matters more than the price.
- Within free software there are two license families. The copyleft ones (GPL, the Linux kernel's license) require derived works to be distributed as free software too. The permissive ones (BSD, MIT, Apache) allow proprietary derivatives; that is why parts of FreeBSD are inside macOS and the PlayStation console.
Comparison table and selection criteria
Let's bring it all together into a decision table. The columns are the questions people really ask when choosing a system:
| Need | Suitable type | Concrete example | Decisive criterion |
|---|---|---|---|
| Serve requests 24/7 with predictable updates | Server, multiuser, preemptive multitasking, free | Debian stable, RHEL | Stability and long-term support |
| Read a sensor every second with 64 KB of RAM | Embedded, hard real time, single-tasking or minimal multitasking | FreeRTOS, Zephyr | Determinism and memory footprint |
| Query application for phones | Mobile, multitasking, per-app isolation | Android, iOS | Ecosystem and battery management |
| Workstation for office work and design | Desktop, single-user in practice | Windows, macOS, Ubuntu | Application compatibility |
| Process nightly batches of millions of records | Server with throughput-oriented scheduling | Linux with SCHED_BATCH, queueing systems |
Throughput over latency |
| Control an industrial robotic arm | Hard real time, embedded | QNX, VxWorks, Linux with PREEMPT_RT | Guaranteed maximum latency |
| Firewall or network gateway | Minimal, hardened server | OpenBSD, minimalist Linux | Reduced attack surface |
And the order in which the questions should be asked:
- Are there deadlines whose breach would be unacceptable? If so, hard real time; the remaining criteria become secondary.
- How many resources are there? With kilobytes of RAM, a general-purpose system is ruled out.
- Who is going to operate it, and how? A team that administers over SSH does not need a graphical environment.
- How long must the deployment last untouched? This determines whether you need a version with extended support.
- What ecosystem do you need? Sometimes the decision is imposed by a library or a client that only exists for one system.
- How much vendor dependency risk do you accept?
Applied case: what Meteora chooses for each piece
Meteora has to decide on three different operating systems. Let's look at the full reasoning.
The meteo-01 server
Choice: server Linux, a stable distribution (Debian stable or RHEL).
| Criterion | Required value | Consequence |
|---|---|---|
| Hard deadlines | There are none. If a request takes 300 ms instead of 100 ms, nothing serious happens | Rules out the need for an RTOS |
| Users | Several services with separate identities | Requires multiuser |
| Tasks | ingestor, aggregator and meteo-api running concurrently |
Requires preemptive multitasking |
| Interface | Remote administration over SSH | Do not install a graphical environment: less RAM and less attack surface |
| Stability | Must run for months untouched | A distribution with a long support cycle, not a rolling release |
| License | No per-core cost, auditable | Free software |
An important nuance: choosing "Linux" does not close the decision, you still have to choose a distribution. A rolling-release distribution such as Arch would give more modern software, but frequent changes are a risk in production. Meteora prefers older versions with security patches guaranteed for years.
The weather stations
Choice: an embedded real-time system, FreeRTOS or Zephyr style, on a microcontroller.
The reasoning changes completely:
- Resources: a typical microcontroller has 64-256 KB of RAM and no disk. Linux, even in a stripped-down version, needs at least several megabytes and an MMU. It is ruled out on size.
- Power: the station runs on a battery and a solar panel. The system must spend most of its time asleep and wake up only to measure and transmit. An RTOS offers fine control over low-power modes.
- Determinism: the sensor must be sampled at regular intervals. It is not hard real time in the "someone dies if it fails" sense, but it does require predictability: a reading shifted by 300 ms corrupts the time series.
- Maintenance-free reliability: the station may sit on a mountain, inaccessible for months. It must recover on its own from any failure (a watchdog timer) and must not have cumulative memory leaks.
- No interface: there is no screen and no keyboard. All interaction happens over radio.
One reasonable exception: if a station had to do heavy local processing (for example, image analysis from a camera), the choice would shift to an embedded Linux built with Yocto or Buildroot on a more powerful SoC. The resource criterion overrides all the rest.
The mobile query app
Choice: Android and iOS, that is, no real choice at all.
Here the technical analysis is almost irrelevant because the decision is imposed by the market: customers already own the device and the system. What is interesting is understanding what that system contributes to Meteora's app:
- Per-application isolation: each app runs with its own user identifier and can only access its private directory. It is the classic multiuser model reused to separate applications instead of people.
- Explicit permissions: accessing the location to show the nearest station requires the user's consent, managed by the system.
- Aggressive power management: the system can suspend or kill the app in the background. Meteora cannot assume its app stays alive; it must use the system's notification mechanisms instead of polling every minute.
- System-controlled life cycle: the app does not decide when it ends.
The conclusion of this case is the most valuable one in the lesson: the same company, with the same product, needs three operating systems of completely different types, and the reason is not anyone's taste but the constraints of each environment.
Common Mistakes and Tips
- Treating the criteria as mutually exclusive. "Is Linux multiuser or multitasking?" is a badly posed question: it is both, because these are different dimensions.
- Believing that real time means fast. It means predictable. An RTOS usually has worse average performance than Linux; what it guarantees is the worst case.
- Thinking multiuser implies several people. The dominant use today is separating services, not people.
meteo-01takes advantage of the multiuser model even though only one person administers it. - Choosing a system out of familiarity rather than requirements. Installing Ubuntu Desktop on a server because "it is the one I know" adds hundreds of unnecessary packages, memory consumption and attack surface.
- Forgetting that "Linux" is not a complete decision. Choosing between Debian, Alpine, RHEL or an image built with Yocto changes more things in practice than choosing between Linux and FreeBSD.
- Tip: when you have to justify a choice, write down the non-functional requirements first (latency, availability, memory, service life, who operates it) and only then look at which systems meet them. Doing it the other way round always ends up justifying the option that had already been decided.
Exercises
Exercise 1
Classify each system along the four dimensions we have seen (processing mode, users, tasks, field of use and license). If any dimension does not apply or is ambiguous, explain why:
- The firmware of a microwave oven.
- Android on a phone.
- Debian on
meteo-01. - MS-DOS in 1985.
Exercise 2
Meteora wants to add an early warning system: if a station detects a pressure drop greater than 5 hPa in 10 minutes, it must issue an alert. Two designs are proposed:
- (A) The station detects the condition locally and transmits a priority alert.
- (B) The station transmits everything normally and
aggregatoronmeteo-01detects the condition.
Analyze what type of operating system each design requires and what trade-offs each one implies in terms of latency, power consumption, reliability and complexity.
Exercise 3
A colleague proposes installing a rolling-release distribution on meteo-01 "so we always have the latest and don't fall behind on security". Write a reasoned reply with at least four arguments, also pointing out in which case their proposal would indeed be the right one.
Solutions
Solution 1
1. Firmware of a microwave oven
- Processing mode: soft real time. It has to react to the keypad and the timer within deadlines, but a 100 ms delay causes no harm. Some safety aspects (cutting the magnetron when the door opens) are normally solved in hardware, not in software.
- Users: single-user, or rather "no concept of a user": there are no identities and no permissions.
- Tasks: single-tasking, or very simple multitasking with an event loop.
- Field: embedded.
- License: proprietary almost always, although it may use a free kernel such as FreeRTOS underneath.
2. Android on a phone
- Mode: interactive, with soft real-time components (audio, video).
- Users: multiuser in its implementation (each app has its own UID) although single-user as far as the person is concerned on most devices. It is the perfect example of the "users" dimension being used today to isolate software, not people.
- Tasks: preemptive multitasking, with the peculiarity that the system can kill background apps to save battery.
- Field: mobile.
- License: mixed. The Linux kernel is GPL, the rest of AOSP is Apache 2.0 (free), but Google's layers (Play Services) are proprietary. A good example of how the license dimension is rarely binary.
3. Debian on meteo-01
- Mode: interactive/time-sharing, with nightly batch jobs. It is not real time.
- Users: multiuser.
- Tasks: preemptive multitasking.
- Field: server.
- License: free, mostly GPL.
4. MS-DOS in 1985
- Mode: interactive, in the limited sense that it responded to a user at a terminal.
- Users: single-user, with no concept of identity or permissions whatsoever.
- Tasks: single-tasking. There were tricks (TSR programs, memory-resident ones) that simulated concurrency, but with no preemption and no protection.
- Field: desktop/personal.
- License: proprietary.
Solution 2
Design A: detection at the station
- System required: the same embedded RTOS, but with more logic. A sliding window of the latest readings has to be kept in RAM, which consumes memory (10 minutes at one reading per minute means 10 pressure values, about 40 bytes: perfectly affordable even in 64 KB).
- Latency: excellent. The alert is issued at the very instant of detection, without depending on the network or on the aggregation cycle. If the normal transmission cycle is 15 minutes, this saves up to 15 minutes of delay.
- Power: worse in two senses. The microcontroller has to be woken up to evaluate the condition on every reading (although the computation is trivial and costs microseconds) and, above all, the radio has to be switched on outside the scheduled cycle, which is what really drains power.
- Reliability: better against network or server failures, because detection does not depend on the reading arriving. Worse against failures of the sensor itself: a station with a faulty sensor can generate false alerts with nobody cross-checking them.
- Complexity: high in the medium term. Changing the threshold from 5 hPa to 4 hPa requires updating the firmware of hundreds of stations deployed in the field, which requires a reliable remote update mechanism, one of the trickiest things in an embedded system.
Design B: detection at the server
- System required: none that is new. The station stays dumb and
aggregatoron Linux takes care of it. - Latency: worse. It depends on the transmission cycle plus the aggregation cycle. If
aggregatorprocesses every 5 minutes and the station transmits every 15, the delay can reach 20 minutes. - Power: unchanged, which is a considerable advantage.
- Reliability: worse against network outages, but better at anomaly detection, because the server can cross-check with neighboring stations and dismiss a pressure drop seen by only one station as a probable sensor failure.
- Complexity: far lower. Changing the threshold means editing
/etc/meteora/meteora.confand restarting a service.
Reasoned recommendation: start with B, because it is reversible and cheap, and move to A (or to a hybrid design: local detection with a conservative threshold plus server-side confirmation) only if the 20-minute latency is shown to have a real cost for customers. It is the general principle of not pushing logic to the edge of the network before a requirement demands it.
Solution 3
Arguments against the proposal:
- Security is not the same as the latest version. A stable distribution such as Debian backports security patches onto older versions. You receive the fix without receiving the behavior changes. The premise that "not updating the version = being insecure" is false.
- Every update is a risk of downtime. In a rolling release, hundreds of packages change every week, including libraries that
ingestor,aggregatorandmeteo-apidepend on. An incompatible change in a library can take the service down at a moment not of your choosing. - Changes are neither atomic nor easily reversible. Rolling back in a rolling release is notoriously difficult, whereas a stable distribution offers a tested upgrade path between specific versions.
- Operational cost multiplies. Updating continuously requires testing continuously. Meteora would have to maintain an equivalent test environment and validate every week, when its value lies in weather data, not in administering the system.
- The surface of change adds no business value. Having a more recent version of a graphics library on a server with no graphical environment contributes nothing.
When the colleague would be right:
- If the service depended on a very recent kernel or library feature (for example, a driver for new hardware or a network function not available in the stable version).
- In a local development environment, where a failure does not affect customers and it is useful to detect future incompatibilities early.
- If the application were deployed in containers with immutable images, where the host's base system matters less and updates are reversible: it is enough to redeploy the previous image. You will see this in Containers: Namespaces and cgroups.
Conclusion
Operating systems are classified along independent and simultaneous dimensions: processing mode (batch, interactive, hard or soft real time), number of users and tasks, field of use (desktop, server, embedded, mobile, distributed, network) and licensing model. A real system occupies a position on all of them at once, and the classification is only useful if it is used to choose with judgment, putting the non-functional requirements first and the candidates second.
The Meteora case leaves us with the central idea: a single company needs a stable server Linux for meteo-01, a kilobyte-sized embedded RTOS for its stations and the mobile systems the market imposes for its app. The constraints of each environment — determinism, memory, power, ecosystem — override any preference.
But notice that, underneath that diversity, they all do the same thing: they manage processes, memory, storage, devices, network and security, and they offer some interface. The priorities and the scale change, not the functions. In the next lesson, Main Functions of an Operating System, we will go through those functions one by one: they are the map of the rest of the course and we will follow them through the complete life cycle of an HTTP request to meteo-api.
Operating Systems Fundamentals
Module 1: Introduction to Operating Systems
- Basic Concepts of Operating Systems
- History and Evolution of Operating Systems
- Types of Operating Systems
- Main Functions of an Operating System
- Kernel Architecture: Monolithic, Microkernel and Hybrid
- User Mode, Kernel Mode and System Calls
Module 2: Resource Management
- Process Management
- CPU Scheduling
- Memory Management
- Virtual Memory and Paging
- Storage Management
- Device Management
- Drivers, Interrupts and I/O Operations
Module 3: Concurrency
- Concurrency Concepts
- Threads and Processes
- Inter-Process Communication (IPC)
- Synchronization and Mutual Exclusion
- Classic Concurrency Problems
- Deadlocks: Prevention, Detection and Recovery
Module 4: File Structures
- File Systems
- Directory Structures
- Partitions, Mounting and the Virtual File System
- File Management
- Space Allocation, Journaling and Integrity
- File Security and Permissions
Module 5: System Protection and Security
- Protection Principles and Access Control
- Users, Authentication and Privilege Escalation
- Common Threats and System Hardening
- Auditing, Logging and Incident Response
Module 6: Virtualization and Containers
- Virtualization: Hypervisors and Virtual Machines
- Containers: Namespaces and cgroups
- The Operating System in the Cloud
- Mobile and Real-Time Operating Systems
