The previous lesson kept talking about "flows of execution" without committing to what they were. Now it is time to be specific. In module 2 you met the process: its memory image, its task_struct, its private address space protected by the MMU. A process is a unit of two things at once: resource ownership (memory, open files, credentials) and execution (a program counter advancing through the code).
The idea of the thread consists of separating those two things. A process is still the owner of the resources, but it can have several flows of execution inside it, sharing everything the process owns. That makes creation and communication brutally cheaper — and in exchange it removes the MMU's safety net between them, which is exactly why the rest of this module exists.
By the end you will know precisely what sibling threads share and what they do not, how much each operation costs in real microseconds, how Linux implements threads (a surprising answer: it does not implement them; it implements clone()), how to program with POSIX threads in C and with threading in Python, what the GIL really is with no myths attached, and how to choose between process, thread and event loop for each piece of Meteora.
Contents
- What a thread is and what its minimum private unit is
- What a thread shares and does not share with its siblings
- The thread control block
- Why threads exist: the numbers
- Implementation models: N:1, 1:1 and M:N
- How Linux really does it:
clone() - Inspecting threads from outside:
ps -eLfand/proc/<pid>/task/ - POSIX threads in C: the
ingestorwith four threads - Threads in Python, the GIL and
multiprocessing - Thread pools: why threads are almost never created by hand
- Process per connection, thread per connection and async in
meteo-api - Criteria for choosing between threads and processes
- Thread termination and cancellation
What a thread is and what its minimum private unit is
A thread (or thread of execution) is the basic unit of CPU usage: a sequential flow of instructions inside a process. A traditional process has exactly one thread; a multithreaded process has several running over the same address space.
The key question for understanding it is: what is the minimum amount of state a flow of execution needs in order to be independent of another? The answer is short:
- A program counter (
%rip): where it is executing right now. - A set of registers: its in-flight variables.
- A stack of its own: its local variables, its parameters and its call chain.
Nothing else. Everything else — the code, the global data, the heap, the open files, the page table — can be shared without the flows ceasing to be independent.
That private stack deserves a moment of attention, because it is the part most often forgotten. If two threads shared a stack, one's function call would trample the other's frame and the program would destroy itself within microseconds. That is why each thread gets its own stack region inside the shared address space: on Linux, 8 MB of reserved virtual space per thread by default (ulimit -s), although only the physical pages that are actually used ever materialize, thanks to the on-demand allocation you saw in module 2.
graph TB
subgraph P["meteo-api process (one address space)"]
COD["Code + global data + heap<br/>cache, counters (SHARED)"]
FD["Descriptor table: sockets,<br/>meteo-api.log (SHARED)"]
H1["Thread 1<br/>%rip + registers<br/>Own stack (8 MB)"]
H2["Thread 2<br/>%rip + registers<br/>Own stack"]
H3["Thread 3<br/>%rip + registers<br/>Own stack"]
end
The diagram holds the whole lesson in one image: three flows with the bare minimum kept private, floating on an ocean of common memory. That ocean is what makes threads fast and dangerous in equal measure.
What a thread shares and does not share with its siblings
This table is the reference you will come back to again and again. It is worth reading in full and carefully, because every row has a practical consequence.
| Element | Shared between threads? | Practical consequence |
|---|---|---|
| Address space (code, data, heap) | Yes | A pointer is valid for all of them. Any global is a potential race |
Page table / mm_struct |
Yes | Switching between threads does not flush the TLB: that is why it is 5-10 times cheaper |
| File descriptor table | Yes | One thread opens /var/log/meteora/meteo-api.log, all of them can write. And a close() affects everyone |
Working directory (cwd) |
Yes | One thread's chdir() changes everyone's relative paths |
Credentials (UID/GID meteora:meteora) |
Yes | You cannot drop privileges for a single thread |
Signal handlers (sigaction) |
Yes | There is one handler table per process, not per thread |
| PID | Yes | They all share the same visible PID; each one has its own TID |
mmap segments, including /dev/shm/meteora-cache |
Yes | One thread's mmap() is seen by all of them immediately |
Resource limits (ulimit) |
Yes | The descriptor limit belongs to the process, not the thread |
Program counter (%rip) |
No | Each thread is at its own place in the code |
| General-purpose registers | No | Saved and restored on the context switch |
| Stack | No | Private local variables. It is the natural "do not share" mechanism |
| TID (thread identifier) | No | gettid(); it is what you see in /proc/<pid>/task/ |
errno |
No (local storage) | Since 1995 it has been a macro expanding to a per-thread variable |
| Blocked signal mask | No | pthread_sigmask(): this one is per thread, even though the handler is common |
| Priority and scheduling policy | No | Each thread is scheduled separately: chrt -p <TID> |
errno, strtok(), TLS declared with __thread |
No | Thread-Local Storage |
| Cancellation state | No | Each thread decides whether it accepts cancellation and when |
Four rows deserve elaboration because they are a constant source of bugs.
errno is thread-local, and it has to be. If it were global, two threads making simultaneous system calls would trample each other's error code and no multithreaded program could check errors reliably. The solution was to turn it into a macro: in glibc, #define errno (*__errno_location()), where __errno_location() returns a different pointer per thread, obtained from the TLS block pointed at by the %fs segment register. That is why errno "just works" in multithreaded programs without you doing anything, and why you cannot do int *p = &errno; in one thread and use it from another.
Signals belong to the process, but the mask belongs to the thread. This asymmetry causes a lot of confusion. There is a single handler table: if one thread installs a handler for SIGHUP, it installs it for all of them. But each thread has its own mask of blocked signals, and when a signal aimed at the process arrives, the kernel picks any thread that does not have it blocked. That means you do not know which thread will handle it. The professional pattern is to block the signal in every thread and dedicate one to receiving them with sigwait(). We will see it applied to reloading /etc/meteora/meteora.conf in Inter-Process Communication (IPC).
The shared descriptor table is a double-edged sword. It is convenient: thread 1 accepts a socket and thread 2 serves it. But if thread 1 does close(fd) while thread 2 is in a read(fd), and the kernel reassigns that number to a new file, thread 2 reads from the wrong file. It is one of the hardest races to find in real servers.
Thread-local storage (TLS) is the escape hatch when you want a "global" that is not shared. In C, __thread is enough (or _Thread_local in C11): __thread unsigned long requests_of_this_thread = 0; gives one copy per thread, with no synchronization. It is the technique that underpins the previous lesson's advice: not sharing beats synchronizing well. If each meteo-api worker keeps its counter in TLS and someone adds them up once a minute, the critical section disappears.
The thread control block
Just as each process has its control block (module 2's task_struct), each thread needs its own: the TCB (Thread Control Block). It contains exactly what the previous table marked as "not shared": the TID; the state (running, ready, blocked, the same ones from module 2); the register context (%rip, %rsp, general-purpose and SIMD); the base and size of its stack region; the scheduler data (vruntime, policy, nice); its signal mask; the pointer to its TLS block; and a pointer to the process's PCB, which is how it reaches everything shared.
The relationship is hierarchical: one PCB, several TCBs pointing at it. And this is where the design becomes interesting on Linux, because Linux does not do exactly this, as we will see in section 6.
Why threads exist: the numbers
The justification for threads is purely quantitative. These are measurements on meteo-01 (Linux 6.1, Xeon at 3.0 GHz), obtained with microbenchmarks of 100,000 repetitions:
| Operation | Cost | Ratio |
|---|---|---|
Create and destroy a process (fork + exit + wait) |
~180 µs | baseline |
Create and destroy a thread (pthread_create + join) |
~22 µs | 8× cheaper |
| Context switch between processes | ~3.5 µs | baseline |
| Context switch between threads of the same process | ~1.2 µs | 3× cheaper |
| Passing 1 MB between processes through a pipe | ~180 µs | baseline |
| Passing 1 MB between threads (the same pointer) | ~0 µs | immediate |
| Memory per idle process (minimum RSS) | ~1,400 KB | baseline |
| Memory per additional thread (materialized stack) | ~12 KB | 100× less |
The reasons for each difference are concrete and rest on what you already know from module 2:
Creation. fork() has to duplicate the mm_struct, walk all the process's memory areas (VMAs), copy the page table marking everything copy-on-write, and duplicate the descriptor table. A clone() that shares mm skips all of that: it increments a reference count and that is that. The bigger the process, the bigger the difference: a process with 2 GB mapped takes far longer to fork() than one with 10 MB, whereas creating a thread costs the same in both cases.
Context switch. The key here is the TLB. Switching between processes requires loading %cr3 with a different page table, which invalidates the TLB entries (mitigated, but not eliminated, by the PCID identifiers of modern CPUs). Then come tens or hundreds of TLB misses while the new process warms back up. Between threads of the same process, %cr3 does not change: the TLB and the caches stay valid. That is most of the factor of 3.
Communication. Data shared between threads is not "sent": it is already there. Passing a pointer to a 1 MB buffer costs 8 bytes. Between processes it has to be copied through the kernel (two copies: user→kernel→user) or explicit shared memory has to be set up.
Those numbers explain why Meteora's ingestor uses threads and not processes to process batches: it creates 4 flows in 88 µs instead of 720 µs, and above all it passes the Reading array around without copying 17 MB.
Implementation models: N:1, 1:1 and M:N
A thread can be managed by the library in user space, by the kernel, or by both. Historically there have been three models.
N:1, user threads. All the management happens in a user-space library. The kernel sees a single process with a single thread and knows nothing. The library saves registers, switches the stack and jumps: a "context" switch costs ~100 nanoseconds, with no system call.
1:1, kernel threads. Each user thread corresponds to a task the kernel can schedule. The kernel knows about them, schedules them individually and spreads them across cores.
M:N, hybrid. M user threads are multiplexed over N kernel threads, with two levels of scheduling.
| N:1 (user) | 1:1 (kernel) | M:N (hybrid) | |
|---|---|---|---|
| Creation cost | ~1 µs | ~22 µs | ~1 µs (the lightweight ones) |
| Switch cost | ~0.1 µs | ~1.2 µs | ~0.1 µs within the same carrier |
| Real parallelism across cores | No | Yes | Yes |
| A blocking call blocks... | Every thread | Only that thread | Only the carrier |
| Kernel scheduling | Ignores the threads | Fair between threads | Two levels, hard to tune |
| Implementation complexity | Low | Medium | Very high |
| Examples | Java 1.1 green threads, GNU Pth | Linux NPTL, Windows, macOS | Old Solaris, Go (goroutines), Java 21 (virtual threads) |
The lethal flaw of N:1 is the third row from the bottom: if one thread does a read() on a socket and blocks, the kernel blocks the whole process, because it only sees one thread. The other 99 threads, even with work to do, sit idle. Combined with the impossibility of using several cores, it doomed the model for general use.
M:N solves both problems on paper, but it is devilishly complex: every blocking call has to be intercepted so the lightweight thread can be migrated to another carrier, and the two schedulers make decisions that contradict each other. Solaris implemented it in the 1990s and ended up abandoning it for 1:1. Linux tried it (IBM's NGPT project) and dropped it too.
The interesting thing is that M:N has come back through the door of languages, not of the operating system: Go's goroutines and Java 21's virtual threads are M:N implemented in the language runtime, which does control every blocking point because it controls the whole standard library. It is the same idea with the missing piece supplied.
Linux chose 1:1 with NPTL (Native POSIX Thread Library, 2003), betting on making kernel threads so cheap that complicating matters would not pay off. The 22 µs in the previous table are the result of that bet.
How Linux really does it: clone()
Here comes the part that surprises almost everyone: the Linux kernel has no concept of a "thread". It has tasks (task_struct), and each task decides what it shares with its creator. A "thread" is simply a task that shares the address space; a "process" is a task that does not. The same system call creates both: clone().
/* Simplified: the real call is clone3() on modern kernels */
long clone(unsigned long flags, void *stack, int *ptid, int *ctid, unsigned long tls);The flags are what decide what gets shared:
| Flag | What it shares with the parent |
|---|---|
CLONE_VM |
The address space (mm_struct) — the flag that makes a "thread" |
CLONE_FS |
Working directory, root, umask |
CLONE_FILES |
The file descriptor table |
CLONE_SIGHAND |
The signal handler table |
CLONE_THREAD |
The thread group: same visible PID, same signal delivery |
CLONE_SYSVSEM |
The System V semaphores |
CLONE_SETTLS |
Installs the given TLS block |
CLONE_NEWNS, CLONE_NEWPID, CLONE_NEWNET... |
It does not share: it creates new namespaces (the basis of containers, module 6) |
With those flags, the two classic operations are just two combinations: fork() is clone(SIGCHLD, ...) — sharing nothing — and pthread_create() is clone(CLONE_VM|CLONE_FS|CLONE_FILES|CLONE_SIGHAND|CLONE_THREAD|CLONE_SYSVSEM|CLONE_SETTLS|..., stack, ...). You can check it with strace:
$ strace -f -e trace=clone,clone3 ./ingestor_4threads 2>&1 | head -6
clone3({flags=CLONE_VM|CLONE_FS|CLONE_FILES|CLONE_SIGHAND|CLONE_THREAD
|CLONE_SYSVSEM|CLONE_SETTLS|CLONE_PARENT_SETTID|CLONE_CHILD_CLEARTID,
child_tid=0x7f2a4c1f8990, parent_tid=0x7f2a4c1f8990,
stack=0x7f2a4b9f8000, stack_size=0x7ffa80}, 88) = 4312It is all there: the list of flags, the 8 MB stack (0x7ffa80 bytes) reserved by the library, and the returned TID 4312.
This unification has three very practical consequences:
- The CFS scheduler schedules threads, not processes. When in module 2 we talked about
vruntimeand sharing out the CPU, the unit was the task. A process with 8 threads receives, by default, 8 times more CPU than a single-threaded process competing with it. - The boundary is a continuum, not a wall. You can create a task that shares memory but not descriptors, or shares descriptors but not memory. Containers exploit exactly that flexibility with the
CLONE_NEW*flags (module 6). - In
/procthreads exist and are visible. Each one has its own directory, as we will see now.
Inspecting threads from outside: ps -eLf and /proc/<pid>/task/
A plain ps -ef hides the threads: it shows one process, even if it has twenty flows inside. The -L option reveals them:
$ ps -eLf | head -1; ps -eLf | grep meteo-api | grep -v grep UID PID PPID LWP NLWP C STIME TTY TIME CMD meteora 2841 1 2841 5 0 08:12 ? 00:00:03 /usr/bin/meteo-api meteora 2841 1 2843 5 3 08:12 ? 00:04:11 /usr/bin/meteo-api meteora 2841 1 2844 5 3 08:12 ? 00:04:08 /usr/bin/meteo-api meteora 2841 1 2845 5 3 08:12 ? 00:04:15 /usr/bin/meteo-api meteora 2841 1 2846 5 3 08:12 ? 00:04:09 /usr/bin/meteo-api
How to read it:
PID2841 on all five rows: it is a single process.LWP(Light Weight Process) is each thread's TID. They are different: 2841, 2843, 2844, 2845, 2846.NLWP= 5: five threads in total.- The thread whose TID matches the PID (2841) is the main thread, the one that started in
main(). Its CPU time is 3 seconds against the others' 4 minutes: it is the thread that accepts connections and hands them out, while the four workers do the real work.
In /proc the structure is just as explicit:
$ ls /proc/2841/task/
2841 2843 2844 2845 2846
$ cat /proc/2841/task/2844/comm
api-worker-2
$ cat /proc/2841/task/2844/stat | awk '{print "utime:", $14, "stime:", $15, "core:", $39}'
utime: 24831 stime: 3102 core: 5One directory per thread, with its own stat, its own status, its own stack. Note that each thread has its own time counters and its current core. What you will not find is a different maps per thread: /proc/2841/task/2844/maps is identical to /proc/2841/maps, because the memory map belongs to the process. It is the practical proof of the first row of the sharing table.
A very useful trick for diagnosis: top -H shows threads instead of processes, and it is the quick way to discover that "the process is at 400 %" actually means "one thread is at 100 % and three at 100 %" or else "one thread is at 400 %... impossible, so there must be four". We will come back to it in Performance Monitoring and Troubleshooting.
POSIX threads in C: the ingestor with four threads
Let us get to real code. The ingestor receives batches of readings and has to validate and convert them before writing them to /var/lib/meteora/readings/. It is perfectly divisible work: each reading is independent of the rest.
/* ingestor_threads.c — processing a batch of Reading with 4 threads */
#include <stdio.h>
#include <stdlib.h>
#include <pthread.h>
#define N_THREADS 4
struct Reading { /* 24 bytes, the course glossary */
unsigned int station_id;
unsigned long timestamp;
float temperature, humidity, pressure;
};
/* Each thread gets ITS OWN chunk. No shared datum is written by two threads. */
struct Chunk {
struct Reading *readings; /* pointer to the shared array */
size_t start, end; /* range [start, end), exclusive to this thread */
size_t valid; /* RESULT: only this thread writes it */
int id;
};
static int reading_is_valid(const struct Reading *r) {
return r->temperature > -90.0f && r->temperature < 60.0f
&& r->humidity >= 0.0f && r->humidity <= 100.0f
&& r->pressure > 800.0f && r->pressure < 1100.0f;
}
void *process_chunk(void *arg) {
struct Chunk *c = (struct Chunk *)arg;
c->valid = 0;
for (size_t i = c->start; i < c->end; i++) {
if (reading_is_valid(&c->readings[i])) c->valid++;
else c->readings[i].station_id = 0; /* mark as discarded */
}
printf("[thread %d] range [%zu,%zu) valid=%zu\n",
c->id, c->start, c->end, c->valid);
return NULL;
}
int main(void) {
size_t n = 700000;
struct Reading *batch = malloc(n * sizeof(struct Reading));
/* ... here the batch would be filled from the socket ... */
pthread_t threads[N_THREADS];
struct Chunk chunks[N_THREADS];
size_t per_thread = n / N_THREADS;
for (int i = 0; i < N_THREADS; i++) {
chunks[i].readings = batch; /* THE SAME pointer */
chunks[i].start = i * per_thread;
chunks[i].end = (i == N_THREADS - 1) ? n : (i + 1) * per_thread;
chunks[i].id = i;
if (pthread_create(&threads[i], NULL, process_chunk, &chunks[i]) != 0) {
perror("pthread_create");
exit(1);
}
}
size_t total = 0;
for (int i = 0; i < N_THREADS; i++) {
pthread_join(threads[i], NULL); /* waits and releases resources */
total += chunks[i].valid; /* safe: the threads have finished */
}
printf("Total valid: %zu of %zu (%.2f%%)\n", total, n, 100.0*total/n);
free(batch);
return 0;
}It is compiled with gcc -O2 -pthread ingestor_threads.c -o ingestor_threads. The -pthread flag is mandatory: it defines _REENTRANT and links the library, and forgetting it produces incomprehensible failures.
Design points worth understanding well, because they are the correct pattern:
All the threads receive the same batch pointer. Nothing is copied. 17 MB of readings shared for the cost of four 8-byte pointers. This is exactly what makes threads cheap.
Each thread has a disjoint range. Thread 0 touches [0, 175000), thread 1 touches [175000, 350000), and so on. No two threads write the same position, so there is no race condition despite sharing the array. It is the most important technique in this lesson: partition the data instead of protecting the data.
Each thread's result goes into its own struct Chunk. If all four did global_total++, we would have exactly the race from the previous lesson. By having each one write to its own structure and summing in the main thread after the pthread_joins, the sum is safe with no synchronization primitive at all.
pthread_join serves two purposes: it waits for the thread to finish and it releases its resources (stack and TCB). A thread that finishes and that nobody joins stays around as a thread zombie consuming memory, just like the zombie processes of module 2. If you are not going to wait for a thread, create it detached (pthread_detach or the PTHREAD_CREATE_DETACHED attribute) so it frees itself.
pthread_create is checked by return value, not by errno. It is a peculiarity of the pthreads API: the functions return the error code directly and do not touch errno. Writing if (pthread_create(...) < 0) perror(...) is a classic mistake that detects nothing.
Performance measured over the 700,000 readings of the file 2026-08-31.dat:
$ time ./ingestor_1thread → real 0m0.412s $ time ./ingestor_threads → real 0m0.118s (3.49× with 4 threads)
3.49× out of a possible 4: 87 % efficiency. It is an almost ideal case precisely because there is no shared state being written. As soon as synchronization is needed, the number will drop.
Threads in Python, the GIL and multiprocessing
Python has real operating-system threads: threading.Thread ends up calling pthread_create. But the CPython interpreter has the GIL (Global Interpreter Lock), a global lock guaranteeing that only one thread executes Python bytecode at a time.
Before the myths, the reason it exists. CPython's memory management uses reference counting: every object carries a counter that is incremented and decremented continuously. That counter is a read-modify-write, exactly the race from the previous lesson. Without protection, two threads manipulating the same object would corrupt the counter and cause premature frees or leaks. The options were a per-object lock (slow because of the cost of atomic operations, and risky because of deadlocks) or a single global lock (simple and blisteringly fast for single-threaded code). CPython chose the latter in 1992 and has been carrying the decision ever since.
The essential point, and the one almost nobody states correctly: the GIL is released during I/O operations and during calls to C code that explicitly drop it. That splits the world in two:
# gil_demo.py — the same structure, two different workloads
import threading, time, multiprocessing
def cpu_bound(n): # compute: holds the GIL the whole time
return sum(i * i for i in range(n))
def io_bound(_): # wait: releases the GIL while waiting
time.sleep(0.5) # simulates reading from a station's socket
def measure(func, arg, n_threads, label):
t0 = time.perf_counter()
threads = [threading.Thread(target=func, args=(arg,)) for _ in range(n_threads)]
for t in threads: t.start()
for t in threads: t.join()
print(f"{label:28} {n_threads} threads: {time.perf_counter()-t0:.2f} s")
if __name__ == "__main__":
for n in (1, 4): measure(cpu_bound, 20_000_000, n, "CPU-bound (threading)")
for n in (1, 4): measure(io_bound, None, n, "I/O-bound (threading)")
t0 = time.perf_counter()
with multiprocessing.Pool(4) as p:
p.map(cpu_bound, [20_000_000] * 4)
print(f"{'CPU-bound (multiprocessing)':28} 4 procs: {time.perf_counter()-t0:.2f} s")Results on meteo-01 (8 cores, CPython 3.11):
CPU-bound (threading) 1 threads: 1.42 s CPU-bound (threading) 4 threads: 5.88 s ← WORSE than 4× a single one! I/O-bound (threading) 1 threads: 0.50 s I/O-bound (threading) 4 threads: 0.50 s ← perfect scaling CPU-bound (multiprocessing) 4 procs: 1.55 s ← 3.8× speedup
Read it slowly, because every line says something:
| Case | Result | Why |
|---|---|---|
| CPU with 1 thread | 1.42 s | Baseline |
| CPU with 4 threads | 5.88 s (≈ 4.1× the single one) | There is no parallelism: they take turns with the GIL. And on top of that there is a 4 % overhead from the lock changing hands every 5 ms |
| I/O with 1 thread | 0.50 s | Baseline |
| I/O with 4 threads | 0.50 s | Perfect parallelism: each thread drops the GIL when it enters sleep/read, so all four wait at the same time |
| CPU with 4 processes | 1.55 s | Each process has its own GIL: real parallelism, 3.8× |
The practical conclusions, without the myths:
- "Python threads are useless" is false. For I/O-dominated work — which is most server code — they scale perfectly. The
ingestorwaiting on 800 sockets is an ideal case forthreading. - "Python threads give CPU parallelism" is false too. For computing, you have to use
multiprocessing, or libraries that drop the GIL in their C code (NumPy, pandas,hashlib,zlibcompression), or your own extensions. - Adding threads to CPU-bound code makes it slower, not merely no faster. GIL contention has a cost.
The cost of multiprocessing is that arguments and results are serialized with pickle and travel through a pipe. For 17 MB of readings that is about 180 ms out and as much again back: if the computation takes less than that, you come out losing. The alternative is multiprocessing.shared_memory, which is POSIX shared memory, and we will see it in Inter-Process Communication (IPC).
A note on the future worth knowing: Python 3.13 introduced an experimental build without the GIL (PEP 703, free-threading), which replaces the global lock with biased reference counting and per-object locks. When it stabilizes, the previous table will change. Until then, the criterion is the one above.
Thread pools: why threads are almost never created by hand
The previous examples create threads, do the work and destroy them. In a real server that is a mistake, for three reasons:
- Repeated creation cost. 22 µs per thread looks like little, but at 1,200 requests per second that is 26 ms per second, 2.6 % of a core thrown away on housekeeping.
- No concurrency limit. One thread per request means a spike of 5,000 simultaneous requests creates 5,000 threads, 40 GB of virtual stack space and a scheduler that spends more time switching context than working. That is thread explosion, and it takes servers down.
- No resource control. Each thread can open descriptors, allocate memory and contact the database. With no ceiling, no capacity planning is possible.
The solution is the thread pool: a fixed number of threads created at startup that consume tasks from a shared queue.
# pool.py — the correct pattern for meteo-api
from concurrent.futures import ThreadPoolExecutor
import urllib.request
STATIONS = [f"http://station-{i:03d}.meteora.local/status" for i in range(1, 81)]
def query(url):
with urllib.request.urlopen(url, timeout=2) as r:
return url, r.status
# 8 threads serve 80 tasks: there are never more than 8 simultaneous connections
with ThreadPoolExecutor(max_workers=8, thread_name_prefix="meteo") as pool:
for url, status in pool.map(query, STATIONS):
print(f"{url} -> {status}")What this code gains over creating 80 threads:
- 8 threads are created once, not 80. The creation cost is amortized across all the tasks.
- Parallelism is capped at 8. The stations do not get 80 connections all at once, and the process does not blow up its consumption.
- The queue acts as a buffer. If tasks arrive faster than they are processed, they queue up instead of creating threads. It is a natural backpressure mechanism.
- The
withguarantees every thread is joined on exit, even if an exception occurs.
Sizing the pool has a useful rule. For CPU work, the optimal size is the number of cores (more threads only add context switches). For I/O work, the classic formula is:
For meteo-api: 8 cores, each request takes 0.4 ms of CPU and waits 12 ms on the database. That gives 8 × (1 + 12/0.4) = 8 × 31 = 248 threads. That number, and the fact that it is so large, is precisely what motivates the next section.
Process per connection, thread per connection and async in meteo-api
meteo-api has to serve HTTP queries. There are three classic architectures and the choice conditions everything.
One process per connection. The server does accept() and fork() — while(1){ int cli=accept(...); if(fork()==0){close(listen_fd); handle(cli); _exit(0);} close(cli); }. It is the model of inetd and of Apache prefork. Maximum isolation: if serving a request causes a segmentation fault, only that process dies. Cost: 180 µs and ~1.4 MB per connection; with 1,000 simultaneous connections, 1.4 GB in processes alone.
One thread per connection. The same, but with pthread_create. 22 µs and about 12 KB of real memory per connection (though 8 MB of virtual). It is the model of Apache in worker mode and of most traditional Java servers. It holds up well to a few thousand connections; beyond that, memory and context switching eat the machine.
Async with an event loop. A single thread (or one per core) that watches thousands of descriptors with epoll and serves whichever is ready. It is the model of nginx, Node.js and asyncio.
# api_async.py — one thread, thousands of connections
import asyncio
async def handle(reader, writer):
data = await reader.read(1024) # yields control while it waits
row = await query_cache(data) # yields again
writer.write(http_response(row))
await writer.drain()
writer.close()
async def main():
server = await asyncio.start_server(handle, "0.0.0.0", 8080)
async with server:
await server.serve_forever()
asyncio.run(main())Every await is a point where the coroutine yields control to the loop, which serves another connection while this one waits. The cost per connection drops to about 2-5 KB and there are no kernel context switches.
| Process/connection | Thread/connection | Async | |
|---|---|---|---|
| Cost per connection | ~1.4 MB, 180 µs | ~12 KB, 22 µs | ~3 KB, ~1 µs |
| Practical connections | hundreds | thousands | hundreds of thousands |
| Exploits several cores | Yes | Yes | Only with one process per core |
| Isolation against failures | Total | None | None |
| Risk of race conditions | Low | High | Very low |
| A slow operation affects... | Only that connection | Only that thread | All of them |
| Programming difficulty | Low | Medium | High (async contagion) |
| Examples | Apache prefork, PostgreSQL | Apache worker, Tomcat | nginx, Node.js, Redis |
The row that determines the most decisions is the second to last: in the async model, a blocking call or a long computation freezes the entire server. A time.sleep(1) instead of await asyncio.sleep(1) inside a coroutine stops all 10,000 connections for a second. It is the number one bug in async code, and it produces no error: just inexplicable latency.
Meteora's decision, now justified: meteo-api uses a pool of 8 threads because each request makes cache queries that take milliseconds and the number of simultaneous clients is in the hundreds, not the tens of thousands; the isolation is not worth the cost of processes. The ingestor, by contrast, keeps 800 permanent connections that are almost always idle — each station sends once a minute — which is the perfect case for async epoll, because 800 sleeping threads would be 800 stacks and 800 tasks in the scheduler for nothing.
Criteria for choosing between threads and processes
| Criterion | Choose processes if... | Choose threads if... |
|---|---|---|
| Isolation against failures | A fault must not take the service down (browsers, plugins, untrusted code) | A fault taking everything down is acceptable |
| Volume of shared data | Little and well delimited | A lot (17 MB of readings) or complex structures |
| Creation frequency | Low (at startup, or a few per minute) | High (thousands per second) |
| Security | You need to separate privileges or apply a different seccomp per piece |
All the code is equally trusted |
| Language | Python with a CPU load (the GIL) | C, Rust, Java, Go; or Python with I/O |
| Debugging | You would rather be able to isolate and debug each piece | You are willing to deal with races |
| Future horizontal scaling | You want to be able to move a piece to another machine | The service will always be local |
At Meteora, ingestor, aggregator and meteo-api are separate processes for three cumulative reasons: if the aggregator fails processing a corrupt file, the API keeps serving; the ingestor needs write permission on /var/lib/meteora/readings/ that the API must not have; and tomorrow the aggregator could be moved to another machine without changing anything in the design. Inside each process, threads, because there the work shares data and isolation adds nothing.
Thread termination and cancellation
A thread can end in four ways: by returning from its function (the value is collected by pthread_join); by calling pthread_exit(value), the same but from any depth of the stack; by receiving a pthread_cancel(tid) from another thread; or because any thread in the process calls exit().
That last possibility is critical and surprising: exit() does not end the thread, it ends the whole process. If a worker thread calls exit(1) on detecting an error, it takes the other three down with it along with the requests they were serving. In a worker thread you do return NULL or pthread_exit(), never exit().
Cancellation is the most delicate mechanism in pthreads. pthread_cancel(tid) does not kill the thread: it sends a request the thread acts on according to its configuration. With type PTHREAD_CANCEL_DEFERRED (the default), it only takes effect on reaching a cancellation point: read, write, sleep, pthread_cond_wait, accept and a few dozen more functions that may block. With PTHREAD_CANCEL_ASYNCHRONOUS, at any instruction, which is almost always a bad idea: the thread can die holding a mutex or with memory half allocated. And a thread can refuse it temporarily with pthread_setcancelstate(PTHREAD_CANCEL_DISABLE, &old) while it does something indivisible.
The underlying problem is cleanup: if the thread dies in the middle of a function, who frees what it had allocated? There are pthread_cleanup_push/pop for registering handlers, but it is a fragile mechanism and easy to forget. That is why professional practice avoids pthread_cancel and uses cooperative termination: a flag the thread checks and a mechanism to wake it up if it is blocked.
volatile sig_atomic_t stop = 0; /* written by the main thread */
void *worker(void *arg) {
while (!stop) {
struct Request *r = take_from_queue(); /* with a timeout */
if (r) handle(r);
}
free_own_resources(); /* guaranteed cleanup */
return NULL;
}The thread decides when to stop, at a point where it knows its state is consistent. It is more code, but it is the only approach that works reliably. And the flag, to be fully correct, should be an atomic_int, not a volatile: we will justify that in Synchronization and Mutual Exclusion.
Common Mistakes and Tips
Forgetting -pthread when compiling. Without it, glibc may link non-reentrant versions of some functions and errno may not be thread-local. The symptoms are random and inexplicable. It is -pthread (not -lpthread), and it goes on both the compile and the link step.
Checking pthreads errors with errno. The pthread_* functions return the error code as their return value and do not touch errno. The right way is int rc = pthread_create(...); if (rc != 0) fprintf(stderr, "%s\n", strerror(rc));.
Passing the address of a loop variable to the thread. The classic mistake: for (int i=0;i<4;i++) pthread_create(&t[i], NULL, f, &i);. All four threads get the same pointer and read the value of i when they happen to run, which may be 4 for all of them. Pass a pointer to an element of an array that survives (such as the example's chunks[i]) or the value cast to void *.
Neither joining nor detaching. A finished thread keeps its TCB and its stack until someone collects its status. In a server that creates threads continuously, that is a memory leak that grows until it exhausts the machine.
Calling exit() from a worker thread. It kills the whole process. Use return or pthread_exit().
Using non-reentrant functions. strtok, asctime, getpwnam, gmtime or rand keep state in shared static variables and return rubbish if two threads call them at once. Always use the _r variants: strtok_r, gmtime_r, getpwnam_r, rand_r.
Tip: partition the data before protecting the data. The ingestor example reaches 3.49× out of a possible 4 without a single synchronization primitive, because each thread works on a disjoint range and writes its result into its own structure. When you can partition, partition: it is faster and it cannot go wrong.
Tip: name your threads. pthread_setname_np(pthread_self(), "api-worker-2") (15 characters maximum) makes top -H, ps -eLf and gdb show readable names instead of repeating the executable's. When you are debugging a hang at three in the morning, you will be grateful.
Tip: do not create threads per unit of work, create a pool. The creation cost, the risk of thread explosion and the lack of backpressure make creating threads on demand almost always a mistake in a service.
Exercises
Exercise 1: demonstrating what is shared and what is not
Write a C program with two threads that experimentally checks three claims from the sharing table: (a) a global variable is shared, (b) a local variable of the thread function is private, (c) a __thread variable is private even though it is global. In addition, print from each thread its TID (gettid()) and its PID (getpid()) to confirm that the PID matches. Verify the result with ls /proc/<pid>/task/.
Exercise 2: measuring the GIL
Write a Python program that measures the time of a CPU task (for example, counting how many of a million simulated Readings are valid) with 1, 2, 4 and 8 threads using threading, and with 1, 2, 4 and 8 processes using multiprocessing. Build the comparison table and explain each column. Then repeat the thread experiment replacing the computation with time.sleep(0.25) and compare.
Exercise 3: choosing an architecture
Meteora wants to add a new component, meteo-alerts, which keeps a WebSocket connection open with every subscribed client in order to send it extreme-temperature warnings. 15,000 clients are expected to be connected simultaneously, with very little traffic (one message every few minutes), and each message requires 0.2 ms of CPU. Choose between process per connection, thread per connection and async. Justify it with memory and context-switch figures, and compute how much CPU in total the service would consume.
Solutions
Solution 1
/* what_is_shared.c */
#define _GNU_SOURCE
#include <stdio.h>
#include <unistd.h>
#include <pthread.h>
#include <sys/syscall.h>
int global = 0; /* (a) shared */
__thread int tls = 0; /* (c) one copy per thread */
void *thread_fn(void *arg) {
int local = 0; /* (b) on the stack, private */
int id = *(int *)arg;
for (int i = 0; i < 5; i++) { global++; local++; tls++; }
printf("thread %d: PID=%d TID=%ld | global=%d local=%d tls=%d | &local=%p\n",
id, getpid(), syscall(SYS_gettid), global, local, tls, (void *)&local);
sleep(2); /* so we can look at /proc while it is alive */
return NULL;
}
int main(void) {
pthread_t t1, t2;
int id1 = 1, id2 = 2;
printf("main: PID=%d TID=%ld\n", getpid(), syscall(SYS_gettid));
pthread_create(&t1, NULL, thread_fn, &id1);
pthread_create(&t2, NULL, thread_fn, &id2);
pthread_join(t1, NULL); pthread_join(t2, NULL);
printf("main: global=%d tls=%d\n", global, tls);
return 0;
}Typical output:
main: PID=5120 TID=5120 thread 1: PID=5120 TID=5121 | global=5 local=5 tls=5 | &local=0x7f3a4bff8e5c thread 2: PID=5120 TID=5122 | global=10 local=5 tls=5 | &local=0x7f3a4b7f7e5c main: global=10 tls=0
Line-by-line analysis:
global: thread 1 sees 5, thread 2 sees 10 (its own plus the other's), andmainsees 10. It is shared. The intermediate values vary from run to run, becauseglobal++is the race from the previous lesson; with more iterations the final total would be less than 10.local: both see 5, and their addresses differ by about 8 MB (0x7f3a4bff8e5cagainst0x7f3a4b7f7e5c, exactly the default stack size). Each thread has its own stack, and the gap between them confirms the 8 MB reservation per thread.tls: both threads see 5, andmainsees 0, even thoughtlsis declared global.__threadcreates one copy per thread, and the main thread has its own.- Identical PID (5120), different TIDs (5120, 5121, 5122): one process, three tasks; the main thread's TID matches the PID. And while they sleep,
ls /proc/5120/task/lists exactly those three TIDs.
Solution 2
# measure_gil.py
import threading, multiprocessing, time, random
def count_valid(n):
valid = 0
for i in range(n):
t = -90 + (i * 37 % 150) # simulated temperature
if -90 < t < 60: valid += 1
return valid
def wait(_):
time.sleep(0.25)
WORK = 4_000_000
def with_threads(func, arg, n):
t0 = time.perf_counter()
ts = [threading.Thread(target=func, args=(arg,)) for _ in range(n)]
[t.start() for t in ts]; [t.join() for t in ts]
return time.perf_counter() - t0
def with_processes(func, arg, n):
t0 = time.perf_counter()
with multiprocessing.Pool(n) as p:
p.map(func, [arg] * n)
return time.perf_counter() - t0
if __name__ == "__main__":
print(f"{'N':>3} {'CPU threads':>12} {'CPU procs':>10} {'I/O threads':>12}")
for n in (1, 2, 4, 8):
print(f"{n:>3} {with_threads(count_valid,WORK,n):>12.2f}"
f" {with_processes(count_valid,WORK,n):>10.2f}"
f" {with_threads(wait,None,n):>12.2f}")Results on meteo-01 (8 cores):
| N | CPU threads | CPU procs | I/O threads |
|---|---|---|---|
| 1 | 0.31 s | 0.36 s | 0.25 s |
| 2 | 0.64 s | 0.38 s | 0.25 s |
| 4 | 1.33 s | 0.41 s | 0.25 s |
| 8 | 2.81 s | 0.52 s | 0.25 s |
"CPU threads" column: the time grows linearly with the number of threads, and a bit more than linearly (2.81 s against the 2.48 s that 8×0.31 would be). There is no parallelism at all: the threads take turns with the GIL, which is handed over every 5 ms by default (sys.setswitchinterval()), and that shuffling adds 13 % of overhead. Eight threads do the work of eight, one after another, and pay a toll on top.
"CPU procs" column: the time barely rises (0.36 → 0.52 s) because the 8 processes really do run on the 8 cores, each with its own GIL. The rise that does happen is due to process startup and serialization, not to the computation. With 8 processes eight times more work is done in 1.4 times more time: an effective speedup of 5.5×.
"I/O threads" column: constant at 0.25 s with 1 thread and with 8. time.sleep() releases the GIL, so the eight waits elapse simultaneously. It is the scenario in which threading is exactly the right tool.
The rule that follows: in CPython, threading for waiting and multiprocessing for computing.
Solution 3
The data: 15,000 simultaneous connections, one message every 5 minutes per client, 0.2 ms of CPU per message.
Real CPU load. 15,000 clients / 300 s = 50 messages/s. At 0.2 ms each: 10 ms of CPU per second, 1 % of a core. The service has no computational problem at all: its problem is keeping 15,000 idle connections open.
Process per connection. 15,000 × 1.4 MB = 21 GB of RSS, impossible on meteo-01, plus 15,000 processes in the scheduler queue. Ruled out.
Thread per connection. 15,000 × 12 KB of materialized stack = 180 MB of RSS, plus about 10 KB of task_struct and kernel stack per thread: another 150 MB. Total ~330 MB to use 1 % of a core. It works, but it keeps 15,000 tasks in the scheduler, nearly all blocked in read(), and every wake-up drags migrations between cores and cache misses along with it.
Async. One descriptor and about 3 KB of state per connection: 45 MB, a single thread, zero context switches between connections. epoll_wait() returns only the ready descriptors, so the cost is O(events), not O(connections). At 50 events per second, the loop is practically asleep.
| Processes | Threads | Async | |
|---|---|---|---|
| Estimated RSS | 21 GB | 330 MB | 45 MB |
| Tasks in the scheduler | 15,000 | 15,000 | 1 |
| CPU usage | 1 % + overhead | 1 % + overhead | 1 % |
| Viable | No | Yes, at a high cost | Yes, comfortably |
Choice: async, and nothing argues against it. It is the textbook case for an event loop: an enormous number of connections, nearly all idle, very little computation per message. It is exactly the same profile as the ingestor with its 800 stations, and the same reason nginx serves hundreds of thousands of connections on a modest machine.
The only precaution, and it has to be taken seriously: no operation in the handler may block. If detecting an alert means writing to the database, that write must be asynchronous or delegated to a small thread pool with run_in_executor(). A synchronous 40 ms INSERT would freeze all 15,000 connections for that long.
Conclusion
A thread is the unit of execution inside a process, and its minimum private state is surprisingly small: program counter, registers and its own stack. Everything else it shares with its siblings: the address space, the descriptor table, the working directory, the meteora:meteora credentials, the signal handlers and mappings such as /dev/shm/meteora-cache. Left out, besides the stack and the registers, are the TID, the signal mask, the priority and errno, which since 1995 has been a macro over thread-local storage precisely because a global errno would make programming with threads impossible.
Threads exist for quantitative reasons we have measured: creating one costs 22 µs against a process's 180 µs, switching between threads costs 1.2 µs against 3.5 µs because there is no need to reload %cr3 or invalidate the TLB, and sharing 1 MB between threads costs zero against the 180 µs of copying it between processes. Of the three implementation models, N:1 died from blocking the whole process on every blocking call, M:N from its complexity — although it has come back in the Go and Java runtimes — and Linux chose 1:1 with NPTL.
And the revelation of the lesson: Linux does not implement threads. It implements clone(), and a thread is a task created with CLONE_VM|CLONE_THREAD|CLONE_FILES|... while a process is a task created without those flags. Hence the scheduler shares CPU between threads and not between processes, the boundary is a continuum that containers will exploit in module 6, and every thread has its directory in /proc/<pid>/task/, visible with ps -eLf.
In C, pthread_create/pthread_join and the pattern that matters: the ingestor's four threads reach 3.49× speedup without a single synchronization primitive, because each one works on a disjoint range of the Reading array and writes its result into its own structure. Partitioning the data beats protecting the data. In Python, the GIL makes 4 threads with a CPU load take 5.88 s against 1.42 s for a single one — worse, not equal — while with an I/O load they scale perfectly; for computing, multiprocessing gives 3.8×. Thread pools avoid the repeated creation cost and the thread explosion, and the sizing comes from cores × (1 + wait/cpu). Among the three server architectures, meteo-api uses threads and the ingestor uses epoll, each because of its load profile. And termination is done cooperatively, never with pthread_cancel, and never with exit() from a worker.
But throughout the lesson we have been cheating. The ingestor's four threads shared nothing that was written, and that is why they worked. As soon as real sharing is needed, and above all as soon as the flows are separate processes — like ingestor, aggregator and meteo-api, which do not share an address space — an explicit mechanism is required for them to pass data around. How does the ingestor send a batch of readings to the aggregator, if they are different processes with separate memories? What really is a pipe on the inside? And how do you tell a process "reload your configuration" without stopping it?
It is the turn of Inter-Process Communication (IPC).
Operating Systems Fundamentals
Module 1: Introduction to Operating Systems
- Basic Concepts of Operating Systems
- History and Evolution of Operating Systems
- Types of Operating Systems
- Main Functions of an Operating System
- Kernel Architecture: Monolithic, Microkernel and Hybrid
- User Mode, Kernel Mode and System Calls
Module 2: Resource Management
- Process Management
- CPU Scheduling
- Memory Management
- Virtual Memory and Paging
- Storage Management
- Device Management
- Drivers, Interrupts and I/O Operations
Module 3: Concurrency
- Concurrency Concepts
- Threads and Processes
- Inter-Process Communication (IPC)
- Synchronization and Mutual Exclusion
- Classic Concurrency Problems
- Deadlocks: Prevention, Detection and Recovery
Module 4: File Structures
- File Systems
- Directory Structures
- Partitions, Mounting and the Virtual File System
- File Management
- Space Allocation, Journaling and Integrity
- File Security and Permissions
Module 5: System Protection and Security
- Protection Principles and Access Control
- Users, Authentication and Privilege Escalation
- Common Threats and System Hardening
- Auditing, Logging and Incident Response
Module 6: Virtualization and Containers
- Virtualization: Hypervisors and Virtual Machines
- Containers: Namespaces and cgroups
- The Operating System in the Cloud
- Mobile and Real-Time Operating Systems
