The history of operating systems is not a list of dates to memorize: it is the history of a sequence of problems and solutions. Every generation of operating systems was born to solve a specific limitation of the previous one, and almost everything that looks obvious to you today — that you can have several windows open, that one program does not take down the others, that you can connect to a remote machine — was at the time a new and controversial idea. Understanding that journey will give you something valuable: when in the coming modules you meet a complicated mechanism (virtual memory, semaphores, a journaling file system), you will know what specific disaster it came to prevent. And you will see that the design ideas that survived are surprisingly few.

Contents

  1. Before the operating system: the bare machine and resident monitors
  2. Batch processing: making the most of every minute of a very expensive machine
  3. Multiprogramming: keeping the CPU from waiting on the disk
  4. Time sharing: giving interactivity back to the programmer
  5. UNIX and a design philosophy
  6. The personal computer: CP/M, MS-DOS, Windows and Macintosh
  7. Free software and the birth of Linux
  8. The era of networks and the Internet
  9. The current era: mobile, cloud and containers
  10. Chronological table of milestones
  11. The ideas that survived, and why

Before the operating system: the bare machine and resident monitors

The 1940s and 1950s. The first computers (ENIAC, EDSAC, IBM 701) had no operating system. The programmer booked the machine by the hour, physically walked into the room, loaded their program using switches or punched cards, ran it and read the results off lights or a printer.

The problem was brutal and very measurable: a machine that cost millions of dollars spent most of its time idle, waiting for a human to change a tape or a deck of cards. It is estimated that useful computing time did not reach 20% of the time rented.

The first solution was the resident monitor: a small program that stayed permanently in memory and took care of loading the next job automatically when the previous one finished. It did not schedule, did not protect, did not share anything out: it merely avoided human intervention between one job and the next. Even so, it is the direct ancestor of the kernel, because it introduced the idea of code that is always there and controls the programs.

Batch processing: making the most of every minute of a very expensive machine

The 1950s and early 1960s. The next step was to group similar jobs into batches: punched cards from many programmers were collected, transferred to magnetic tape on a cheap machine, and the expensive machine ran the whole tape in one go.

Here an idea appears that is still with us: the job control language (JCL on IBM systems), with which the programmer described what resources their job needed. It was, in essence, the first deployment script.

//JOB1      JOB  (ACCOUNT),'METEO AVERAGES',CLASS=A,TIME=(0,30)
//STEP1     EXEC PGM=AGGREGATOR
//INPUT     DD   DSN=METEORA.READINGS.D260831,DISP=SHR
//OUTPUT    DD   DSN=METEORA.AVERAGES.D260831,DISP=(NEW,CATLG)

Even though the syntax looks strange, the content will sound very familiar:

  • JOB declares the job, who gets billed for it (ACCOUNT) and how much time it may consume at most (TIME=(0,30), 30 seconds). It is exactly the same concept as a resource limit on a modern container.
  • EXEC PGM=AGGREGATOR says which program to run.
  • The DD (Data Definition) lines associate logical names with specific data sets. It is the same principle as a file descriptor or an environment variable holding a path: the program uses a logical name and the system decides what sits behind it.

The limitation still outstanding: while the job read from tape, the CPU sat idle. And reading a tape was thousands of times slower than computing.

Multiprogramming: keeping the CPU from waiting on the disk

The 1960s. The solution was to keep several jobs in memory at once. When job A blocks waiting for the tape, the system hands the CPU to job B. When A gets its data, it goes back into the queue. This is multiprogramming, and with it the operating system as a resource manager is truly born.

The impact is easy to quantify. Suppose a job spends 80% of its time waiting for input/output. With a single job in memory, the CPU is used 20% of the time. With n independent jobs in memory, the probability that all of them are waiting at once is 0.8ⁿ, so utilization is 1 − 0.8ⁿ:

Jobs in memory CPU utilization
1 20%
2 36%
3 49%
5 67%
10 89%

Going from 1 to 5 jobs triples throughput without changing the hardware. That calculation on its own justified all the added complexity.

But multiprogramming brought new problems, and three whole modules of this course come out of them:

  • If there are several programs in memory, they have to be protected from each other → memory management and protection.
  • Someone has to decide whose turn it is to use the CPU → scheduling.
  • Programs can interfere with each other when using shared resources → concurrency and synchronization.

The emblematic systems of this stage were IBM's OS/360 (1964) and Multics (which we will get to shortly). OS/360 is also famous for its disastrous development, which Fred Brooks recounted in The Mythical Man-Month: it was the first project to demonstrate that adding programmers to a late project makes it later.

Time sharing: giving interactivity back to the programmer

Batch processing was efficient for the machine but horrible for the human: you handed in your cards and got the result the next day. If you had left out a comma, you lost an entire day.

The idea of time sharing was to give each user a slice of CPU time so short and so frequent that everyone would believe they had the machine to themselves. Since humans think slowly and type slowly, a single machine could serve dozens of simultaneous users.

  • CTSS (Compatible Time-Sharing System, MIT, 1961) was the first one that really worked. It proved the idea was viable.
  • Multics (MIT, Bell Labs and General Electric, from 1965) was the ambitious project that wanted to take the idea to the extreme: a "computing utility" comparable to the electricity grid, with security through privilege rings, segmented memory, a hierarchical file system and high availability. Multics was too ambitious and arrived late, and Bell Labs abandoned the project in 1969.

Multics is usually told as a failure, but it is an extraordinarily influential failure: privilege rings, the hierarchical file system with nested directories, dynamic linking and the idea of treating memory and files uniformly all come from it. And its abandonment directly produced the most influential operating system in history.

UNIX and a design philosophy

1969, Bell Labs. Ken Thompson, frustrated by the disappearance of Multics, wrote a much smaller and simpler system on a disused PDP-7. Brian Kernighan jokingly christened it UNICS (a pun on Multics), and it ended up being called UNIX.

Two decisions changed everything:

  1. Rewriting it in C (1973, with Dennis Ritchie). Until then operating systems were written in assembly, tied to a specific machine. A system written in a high-level language could be ported to other hardware by recompiling. It was an idea considered risky and inefficient at the time; today it is the absolute norm.
  2. Distributing it with the source code to universities for a token price. That created an entire generation of programmers who learned by reading a real operating system, and it gave rise to BSD (Berkeley Software Distribution), which produced the most widely used TCP/IP implementation in the world, sockets and the vi editor.

The UNIX philosophy, still the best software design guide there is, boils down to a few principles:

  • Write programs that do one thing and do it well.
  • Write programs that work together, communicating through text streams.
  • Plain text is the universal interface.
  • Everything is a file.

An example with Meteora's data illustrates the power of those principles better than any explanation:

grep 'station=118' /var/log/meteora/meteo-api.log \
  | awk '{print $4}' \
  | sort \
  | uniq -c \
  | sort -rn \
  | head -5

Let's go part by part:

  • grep 'station=118' ... filters out of the log only the lines for station 118. grep knows nothing about meteorology: it just searches for text.
  • | is the pipe, UNIX's star contribution: it connects one program's output to the next one's input, with no temporary files and without either of them knowing the other exists.
  • awk '{print $4}' extracts the fourth field of each line (let's assume it is the time of the request). awk has no idea what that field means.
  • sort sorts, a requirement for the next step.
  • uniq -c collapses repeated lines and prefixes each one with how many times it appeared.
  • sort -rn re-sorts numerically (-n) and from highest to lowest (-r).
  • head -5 keeps the first five.

Result: the five hours with the most queries to station 118. None of those six programs was written with Meteora or with the other five in mind, and yet they cooperate. That is UNIX's design achievement, and it is the reason the command line is still a first-class tool fifty years later (we will come back to it in The Command Line as the System Interface).

The personal computer: CP/M, MS-DOS, Windows and Macintosh

The 1970s and 1980s. The microprocessor made hardware cheap enough that a single person could own a computer. But that computer was far more limited than the big systems: no protected memory, very little RAM and a single user. Curiously, that meant a technical step backwards: the first PC operating systems forgot almost everything learned about multiprogramming and protection, simply because the hardware did not allow it.

  • CP/M (Gary Kildall, 1974) was the first OS widely used on 8-bit microcomputers. Its great idea was the BIOS: separating the part of the system that depends on the specific hardware from the rest, so that the same CP/M would work on machines from different manufacturers. It is an early case of a hardware abstraction layer.
  • MS-DOS (Microsoft, 1981) arrived on the IBM PC. Technically it was poor — single-tasking, no memory protection, 8+3 character file names — but it was the system of the computer that became the industry standard.
  • Macintosh (Apple, 1984) popularized the graphical interface with windows, icons, menus and a mouse, ideas born at Xerox PARC with the Alto and the Smalltalk system. It changed forever who could use a computer.
  • Windows started out as a graphical environment on top of MS-DOS (1985). Windows 95 mixed 16-bit and 32-bit code with cooperative multitasking, which explains its legendary instability. The real break came with Windows NT (1993), designed from scratch by a team led by Dave Cutler, who came from VMS: a hybrid kernel, protected memory, preemptive multitasking and multiuser support. Every modern Windows descends from NT, not from MS-DOS.
Comparison MS-DOS (1981) Windows NT (1993) UNIX (1970s)
Multitasking No Yes, preemptive Yes, preemptive
Memory protection No Yes Yes
Multiuser No Yes Yes from the start
Portability x86 only Designed to be portable Portable thanks to C

Free software and the birth of Linux

1983-1991. Richard Stallman launched the GNU project with the goal of building a complete, free operating system compatible with UNIX. GNU produced fundamental pieces — the gcc compiler, the gdb debugger, bash, the coreutils — and, above all, the GPL license, which guarantees that derived code stays free. What GNU lacked was precisely the kernel: its own kernel project, Hurd, never became ready.

1991. Linus Torvalds, a student in Helsinki, released a kernel of his own as a personal project. His original message said it would be "just a hobby, won't be big and professional like GNU". Combined with the GNU tools, that kernel formed a complete and free operating system.

The key to its success was not only technical but organizational: the open, distributed development model, with thousands of contributors and a fast integration cycle, turned out to be more effective than anyone expected. Today the Linux kernel has more than 30 million lines of code and receives contributions from thousands of people and from almost every large technology company. On meteo-01, as on the vast majority of the world's servers, it is that kernel doing the work:

uname -a
Linux meteo-01 6.1.0-18-amd64 #1 SMP PREEMPT_DYNAMIC Debian 6.1.76-1 x86_64 GNU/Linux

What each field tells you:

  • Linux is the kernel; meteo-01 the machine's name.
  • 6.1.0-18-amd64 is the kernel version and the architecture.
  • SMP stands for Symmetric MultiProcessing: support for several CPU cores, a direct inheritance from multiprogramming.
  • PREEMPT_DYNAMIC indicates that the kernel can be preempted (interrupted) even while running kernel code itself, which reduces latencies.
  • GNU/Linux is a reminder of precisely the above: the system is the Linux kernel plus GNU tools.

The era of networks and the Internet

The 1980s and 1990s. With permanent connectivity, new requirements appeared that the OS had to absorb:

  • The network stack as part of the kernel. TCP/IP stopped being an add-on and became a core subsystem. The BSD implementation was the worldwide reference and its sockets API is still the standard.
  • The socket abstraction, which extended "everything is a file" to remote communication: a socket is read and written like a file.
  • Network file systems (NFS, SMB), which allowed a directory on one machine to appear inside another machine's tree.
  • Security as a real problem. With an isolated machine, protection was almost an academic matter. With a connected machine, any flaw is exploitable from the other side of the world. The 1988 Morris worm demonstrated this spectacularly by infecting a significant portion of the Internet of the time, and it is the reason module 5 of this course exists.

The current era: mobile, cloud and containers

Since the 2000s. Three shifts have defined the current stage:

  • Mobile. Android (on the Linux kernel) and iOS (on a kernel derived from Mach and BSD) brought operating systems to billions of devices, with two new priorities: battery (the system aggressively shuts down whatever is not in use) and per-application isolation (each app runs with its own user and explicit permissions, instead of inheriting those of the human user).
  • Cloud. Virtualization let many virtual machines share a physical server, turning computing into a service rented by the minute. Curiously, it is the old Multics idea — computing as a public utility — fulfilled sixty years later.
  • Containers. Instead of virtualizing the whole hardware, a set of processes is isolated within the same kernel using namespaces and cgroups. It is a further twist on the classic OS protection, not a technology foreign to it.

These three topics are developed in module 6: Virtualization: Hypervisors and Virtual Machines, Containers: Namespaces and cgroups and Mobile and Real-Time Operating Systems. Here we are only interested in placing them on the timeline.

Chronological table of milestones

Year Milestone Limitation it solved Lasting contribution
~1950 Resident monitors Human intervention between jobs Control code always in memory
1956 GM-NAA I/O (first recognized OS) Chaining jobs automatically Batch processing
1961 CTSS (MIT) A day's wait for a result Interactive time sharing
1964 IBM OS/360 One OS per machine model A family of machines with a common OS
1965 Multics Secure sharing between users Privilege rings, hierarchical files
1969 UNIX (Bell Labs) Multics's complexity Simplicity, "everything is a file", pipes
1973 UNIX rewritten in C Systems tied to one piece of hardware Operating system portability
1974 CP/M Each micro with its own software BIOS/system separation
1977 BSD Distribution and evolution of UNIX Sockets, TCP/IP, vi
1981 MS-DOS Lack of an OS for the IBM PC Standardization of the PC
1984 Macintosh An interface only for experts GUI with windows, icons and a mouse
1983 GNU project Closed proprietary software The GPL license and free tools
1991 Linux kernel GNU without a usable kernel A free kernel and distributed development
1993 Windows NT Instability of Windows on DOS Hybrid kernel, protected memory on the PC
1995-2000 Mass Internet Isolated machines Network stack in the kernel, security
2007-2008 iOS and Android OSes not adapted to mobile Power management, per-app isolation
2006-2010 Cloud (EC2 and similar) Underused servers Computing as a service
2013 Docker Heavy VMs for deployment Containers on namespaces and cgroups

The ideas that survived, and why

If you compare a 1970 system with meteo-01, almost all the code changes but the conceptual skeleton holds. These are the ideas that endured, and the reason they did:

  1. The process as an isolated unit of execution. It survives because it solves two problems at once (accounting and protection) with a single concept.
  2. The separation between user mode and kernel mode. It is the only known way for the system to be able to trust itself even though it does not trust the programs it runs. No alternative has been better. It is the subject of User Mode, Kernel Mode and System Calls.
  3. The file as a named sequence of bytes. It survives because of what it does not impose: by dictating no format, it works equally well for a .dat file of readings, an image or an executable.
  4. The directory hierarchy. A single tree is simple enough to understand and flexible enough to organize millions of files.
  5. "Everything is a file". Reusing a familiar interface for new things (devices, sockets, kernel information) hugely reduces what you have to learn and what you have to program.
  6. Pipes and the composition of small programs. They survive because they scale to problems nobody foresaw.
  7. Virtual memory. It gives each process a space of its own, allows using more memory than physically exists and simplifies loading programs. Three benefits from one mechanism.
  8. Time sharing with preemption. Without the OS's ability to take the CPU away from a process, a single badly written program paralyzes the machine. The era of cooperative multitasking (Windows 3.x, classic Mac OS) demonstrated in practice just how badly the alternative works.

And one idea that did not survive but is worth knowing about: systems that relied on the good will of programs (cooperative multitasking, unprotected memory) always failed. The practical lesson, which also applies to the software you write yourself, is that a system must not depend on its users behaving well.

Common Mistakes and Tips

  • Studying the dates instead of the problems. Nobody is going to ask you what year Multics came out. What does matter is that you can explain what problem time sharing solved and why multiprogramming demands memory protection.
  • Believing evolution always moved forward. It did not: the first PC systems were a technical step backwards compared with the big systems of the 1960s. Hardware and market constraints carry as much weight as good ideas.
  • Confusing Linux with GNU/Linux, or with a distribution. Linux is only the kernel. What you install is a set made up of that kernel plus GNU tools and many other programs, packaged by a distribution.
  • Thinking UNIX is obsolete. macOS is a certified UNIX, Android uses the Linux kernel, iOS derives from BSD and practically every server in the world is UNIX or Linux. The model is not merely far from obsolete: it won.
  • Tip: when a mechanism in the coming modules strikes you as needlessly complicated, ask yourself what specific disaster it prevents. There is almost always a historical anecdote behind it, and remembering it fixes the concept far better than the definition does.

Exercises

Exercise 1

For each limitation on the left, state which historical advance solved it and what new problem that advance introduced:

  1. The CPU is idle while the human changes the cards.
  2. The CPU is idle while the job reads from tape.
  3. The programmer waits an entire day to find out whether their program compiles.
  4. The operating system has to be rewritten from scratch when the machine changes.

Exercise 2

On meteo-01, ingestor spends 90% of its time waiting for network data and only 10% computing. Using multiprogramming's CPU utilization formula (1 − p^n, with p the fraction of time spent waiting), compute the utilization with 1, 3, 6 and 12 processes of that profile. Is it worth going from 6 to 12? Reason about what other factor of the real system limits this formula.

Exercise 3

Write a single command line, in the style of the UNIX philosophy, that from /var/log/meteora/meteo-api.log obtains the three IP addresses that have made the most requests. Assume the IP is the first space-separated field of each line. Explain which principle of the UNIX philosophy your solution exemplifies and why no Meteora-specific program is needed to solve it.

Solutions

Solution 1

  1. Human changing cards → solved by batch processing with a resident monitor, which chained jobs automatically. New problem: the user loses all interactivity and can no longer intervene while their job runs.
  2. CPU idle waiting for the tape → solved by multiprogramming, keeping several jobs in memory and switching when one blocks. New problems: the memory of some jobs has to be protected from others, someone has to decide who gets the CPU (scheduling) and race conditions appear when resources are shared.
  3. A day's wait for the result → solved by time sharing, giving very short CPU slices to many interactive users. New problems: how to guarantee reasonable response times when there are many users, how to stop one of them from monopolizing the system and how to isolate users who now genuinely coexist (the need for accounts, permissions and passwords is born).
  4. Rewriting the OS when the machine changes → solved by writing the operating system in C (UNIX, 1973), confining the hardware-dependent code to a few parts. New problem: a small loss of performance compared with hand-written assembly, and the need for a reliable compiler for each target architecture. The trade-off turned out to be so favorable that today nobody considers the opposite.

Solution 2

With p = 0.9 (the fraction of time a process spends waiting):

n Calculation Utilization
1 1 − 0.9¹ = 1 − 0.900 10.0%
3 1 − 0.9³ = 1 − 0.729 27.1%
6 1 − 0.9⁶ = 1 − 0.531 46.9%
12 1 − 0.9¹² = 1 − 0.282 71.8%

Is it worth going from 6 to 12? In terms of this formula, yes: you gain almost 25 percentage points, slightly more than was gained going from 3 to 6 (19.8 points). But notice that the marginal return is already falling: each added process contributes less than the previous one, because the curve flattens as it approaches 100%.

What limits the formula in the real world:

  • Memory. Each additional process takes up RAM. If adding processes makes the system start swapping pages to disk (swapping), performance collapses instead of improving. You will see this in Virtual Memory and Paging.
  • The statistical independence the formula assumes. If all the processes wait on the same resource (for example, the same disk or the same network card), their waits are correlated and real utilization is far worse than predicted.
  • The cost of context switching. Switching between processes costs CPU time that the formula ignores. With too many processes, a growing share of the CPU goes into managing the switching rather than into useful work.

Solution 3

awk '{print $1}' /var/log/meteora/meteo-api.log | sort | uniq -c | sort -rn | head -3

Step by step:

  • awk '{print $1}' prints only the first field of each line, that is, the IP. We could also use cut -d' ' -f1, which is equivalent and somewhat faster.
  • sort groups identical lines together, making them consecutive. It is essential because uniq only detects adjacent repetitions.
  • uniq -c replaces each group of identical lines with a single one preceded by the number of repetitions.
  • sort -rn sorts by that number (-n, numeric) from highest to lowest (-r).
  • head -3 keeps the first three.

The principle it exemplifies: the composition of small, specialized programs through pipes, with plain text as the universal interface. None of the five programs knows what an IP is, what Meteora is or what the others do; each performs a minimal transformation on lines of text.

Why no specific program is needed: because the log format is text and the operations required (filter, extract, group, count, sort) are generic. Writing a meteora-analyzer in Python would give the same result, but it would take twenty times as long, would have to be maintained and would only be useful for this one case. The common mistake here is the opposite one: using the command line for tasks that already require complex state or business logic, where a real program is indeed the right answer.

Conclusion

Operating systems evolved by solving one problem after another: first human intervention, then the idle CPU, then the lack of interactivity, later portability and finally isolation in a connected world. Each solution brought new problems, and that chain explains why a modern system has the structure it has.

There are three things worth taking away from this journey. First, that multiprogramming is the origin of almost all of an OS's complexity: without it there would be no need for memory protection, scheduling or synchronization. Second, that UNIX won through simplicity and portability, not through power, and that its philosophy remains the best design guide available. And third, that the ideas that survived are few and very general: process, dual mode, file, hierarchy, virtual memory.

All this evolution produced an enormous variety of present-day systems, from the firmware of a weather station that takes up a few kilobytes to a server kernel with millions of lines of code. In the next lesson, Types of Operating Systems, we will bring order to that variety by classifying it using clear criteria, and we will decide which kind of system suits each piece of Meteora's infrastructure.

Operating Systems Fundamentals

Module 1: Introduction to Operating Systems

Module 2: Resource Management

Module 3: Concurrency

Module 4: File Structures

Module 5: System Protection and Security

Module 6: Virtualization and Containers

Module 7: Administration and Troubleshooting in Practice

© Copyright 2026. All rights reserved