The previous three lessons have raised a perimeter: a controlled network, SSH with keys and no root, a firewall with an allowlist policy and fail2ban blocking whoever persists. All of that is prevention, and prevention has an uncomfortable property: when it fails, it does not tell you. An attacker who gets in with a legitimate credential trips no rule; their connection shows up as established, exactly like yours. This lesson changes the question from "how do I stop them getting in?" to "how do I find out that they are already in?", and answers it with concrete tools: file integrity with AIDE, kernel auditing with auditd, systematic reading of the logs, configuration auditing with Lynis and an incident response procedure you can follow at three in the morning without improvising.
Legal warning up front. Everything here is defensive and is applied to srv-tramontana, which is your own lab machine. Scanning, probing or attempting to access systems that are not yours without written authorisation is a criminal offence, regardless of intent. And on the compliance side: detecting intrusions on a system that processes guests' personal data triggers specific legal obligations — breach notification within 72 hours under the GDPR — and monitoring the activity of working people has limits that are not the administrator's to decide. Both are dealt with at the end of the lesson.
Contents
- Why detection is a different control from prevention
- Taxonomy: NIDS, HIDS, signatures, anomalies, IDS and IPS
- File integrity with AIDE
- Automating AIDE with a systemd timer
- Rootkit detection: rkhunter and chkrootkit
- auditd: the kernel's audit subsystem
- Log analysis for detection
- NIDS: an honest overview of Suricata
- Configuration auditing with Lynis
- Incident response and legal obligations
Why detection is a different control from prevention
A preventive control tries to stop something happening. A detective control assumes it can happen and makes sure it does not go unnoticed. They are not alternatives: they are different layers, and an organisation with only the first would learn about a compromise from a third party's phone call, weeks later.
The four ways in which srv-tramontana's prevention can fail, today:
| Failure route | Why the firewall does not help |
|---|---|
| An unpatched vulnerability (a 0-day or the exposure window) | The traffic arrives on a legitimately open port |
| A leaked credential | The authentication is correct as far as the system is concerned |
| A configuration mistake of your own | The rule you opened "just for a moment" is still open |
| Abuse of legitimate access | Whoever is acting is already inside and is allowed to be |
The mindset to adopt is called assumed compromise: not "am I secure?", but "if I were compromised, how would I know?". And that question breaks down into three, which are what structure the whole lesson:
- What has changed? An attacker who wants to persist has to modify something: a binary, a configuration file, an authorised key, a systemd unit. → file integrity.
- Who got in? Every session leaves a trail in the authentication logs. → log analysis.
- What did they do? Which files were read or written, which processes were launched. → kernel auditing.
Taxonomy: NIDS, HIDS, signatures, anomalies, IDS and IPS
The vocabulary of this field is confusing because it mixes three independent axes. It is worth separating them:
| Axis | Options | What distinguishes them |
|---|---|---|
| Where it observes | NIDS (network) / HIDS (host) | The NIDS looks at packets in transit; the HIDS looks at the state and activity of one machine |
| How it decides | Signatures / Anomalies | Signatures recognise the known bad; anomalies detect deviations from the normal |
| What it does on detection | IDS (warns) / IPS (blocks) | The IPS acts, with the risk of blocking legitimate traffic |
The practical consequences of each choice:
- Signatures: high precision, few false positives, and total blindness to anything not in the catalogue. They need constant updating.
- Anomalies: they can detect the unknown, in exchange for false positives and for needing a learning period to work out what "normal" is. And "normal" changes: the baseline you established in 05-07 expires.
- IPS: when it gets it wrong, it causes a service outage.
fail2ban, which you already have running, is precisely a narrow-purpose IPS — and you already saw in 06-03 the risk of it blocking you.
For a single server like srv-tramontana, the investment with the best return is an integrity HIDS plus kernel auditing. The NIDS gains value when there is a network with several machines and a single point all the traffic passes through; we will come back to that.
File integrity with AIDE
File integrity is the most reliable signal a server has, for a structural reason: an attacker who wants to persist has to write to disk. They can delete log entries, they can falsify the output of ps with a rootkit, but the modified binary or the added key is still there, and its cryptographic fingerprint does not match the one you had.
AIDE (Advanced Intrusion Detection Environment) builds a database with the attributes and checksums of the files you tell it about, and then compares the current state against it.
The selection rules
The configuration lives in /etc/aide/aide.conf and in the fragments under /etc/aide/aide.conf.d/. The essential part is the attribute groups, which define what is checked about each file:
| Attribute | What it checks |
|---|---|
p |
Permissions |
i |
Inode number |
n |
Link count |
u / g |
Owner / group |
s |
Size |
m |
Modification time (mtime) |
c |
Inode change time (ctime) |
md5 / sha256 |
Checksum of the content |
They are combined with + into reusable definitions, and then applied to paths with path group to watch and !path to exclude:
# /etc/aide/aide.conf.d/99_tramontana
# Full group: everything that can be checked about a file
TramoAll = p+i+n+u+g+s+m+c+md5+sha256
# Loose group: for files whose content changes legitimately
# but whose permissions and ownership must NEVER change
TramoPerms = p+u+g
# Configuration and binaries: any change is suspicious
/etc/tramontana$ TramoAll
/opt/tramontana/releases$ TramoAll
/home/operator/scripts$ TramoAll
/home/operator/bin$ TramoAll
# Logs grow constantly: watch the container, not the content
/var/log/tramontana$ TramoPerms
!/var/log/tramontana/.*\.log$
!/var/log/tramontana/.*\.gz$
# The backups change every night; only permissions and ownership matter
/srv/tramontana/backups$ TramoPerms
!/srv/tramontana/backups/.*Notice the criterion: the level of watchfulness is matched to what is expected to change. Watching the content of access.log with sha256 would produce an alert every minute and within two days you would have stopped reading the reports — which is the most common way for a detection system to stop being of any use.
Initialising the database, and where to keep it
$ sudo aideinit
Running aide --init...
Start timestamp: 2026-08-18 11:04:22 +0200 (AIDE 0.18.6)
AIDE initialized database at /var/lib/aide/aide.db.new
Number of entries: 231847
$ sudo mv /var/lib/aide/aide.db.new /var/lib/aide/aide.db
$ sudo chmod 600 /var/lib/aide/aide.dbAnd now the point that decides whether all of this is worth anything or is pure theatre:
An integrity database that lives on the server being watched is worth nothing. An attacker with root privileges modifies the binary, regenerates the database, and
aide --checkwill tell you everything is in order. The database — and, if you can afford it, theaidebinary itself — has to be out of reach: on read-only media, or on another machine.
Like the recovery runbook you already keep off the server (05-08), the reference copy leaves the machine:
$ sudo sha256sum /var/lib/aide/aide.db | sudo tee /var/lib/aide/aide.db.sha256
c4f1...e88a /var/lib/aide/aide.db
$ scp -3 operator@srv-tramontana:/var/lib/aide/aide.db.sha256 \
student@laptop-student:~/tramontana-reference/By keeping at least the checksum of the database off the server you can detect that the database itself has been tampered with, which is the attack you have to protect against.
Interpreting a change report
$ sudo aide --check
Start timestamp: 2026-08-18 11:31:07 +0200 (AIDE 0.18.6)
AIDE found differences between database and filesystem!!
Summary:
Total number of entries: 231847
Added entries: 1
Removed entries: 0
Changed entries: 2
---------------------------------------------------
Added entries:
---------------------------------------------------
f++++++++++++++++: /home/operator/.ssh/authorized_keys2
---------------------------------------------------
Changed entries:
---------------------------------------------------
f ... . C... : /usr/bin/openssl
f ... . C... : /usr/lib/x86_64-linux-gnu/libssl.so.3The notation is dense but mechanical: the first letter is the type (f file, d directory, l link), and the following positions say which attribute changed (C content, p permissions, u owner, s size, m mtime). A + in the column means it was added.
And here is the real skill, which is not reading the report but classifying it:
-
The two changed files are
opensslandlibssl. You have a documented explanation: on 18/08 at 06:12unattended-upgradesapplied the fix for a TLS CVE. It is a legitimate change, and it is verified independently:$ grep -E 'openssl|libssl' /var/log/apt/history.log | tail -2 Upgrade: libssl3t64:amd64 (3.0.13-0ubuntu3.4, 3.0.13-0ubuntu3.5), openssl:amd64 (3.0.13-0ubuntu3.4, 3.0.13-0ubuntu3.5) $ sudo debsums -c openssl # no output: every file in the package matches Debian's manifestdebsumsis the second opinion: it verifies the installed files against the package's own checksums. They match, so the binary is the one Ubuntu published. -
The added file is another matter.
authorized_keys2is an obsolete name that OpenSSH no longer reads by default, but that older configurations do, and nobody had any reason to create it. This has no documented explanation, and it is exactly the shape of a persistence mechanism. Here you do not delete the file: you trigger the incident response procedure at the end of the lesson.
Once the legitimate changes have been validated, the database is updated so that the next report is clean again:
There are alternatives: Tripwire is AIDE's ancestor and has a somewhat more robust model for cryptographically signing its own database, at the cost of considerably more complexity. You have already seen debsums, and it covers only files from packages. For a single server, AIDE is the right balance.
Automating AIDE with a systemd timer
A check you have to remember to run does not get run. With what you learned in 05-05 the pattern is straightforward, and it respects the principle of silence if all is well:
$ sudo tee /usr/local/sbin/aide-check >/dev/null <<'EOF'
#!/usr/bin/env bash
set -euo pipefail
readonly LOG_TAG="aide-tramontana"
output="$(mktemp)"; trap 'rm -f "$output"' EXIT
if aide --check >"$output" 2>&1; then
logger -t "$LOG_TAG" -p local0.info "integrity correct, no changes"
exit 0
fi
# aide returns a non-zero code when there are differences
logger -t "$LOG_TAG" -p local0.warning "AIDE detected changes: review the report"
mail -s "[srv-tramontana] AIDE detected changes" [email protected] <"$output" \
|| logger -t "$LOG_TAG" -p local0.err "could not send the notification"
exit 1
EOF
$ sudo chmod 700 /usr/local/sbin/aide-check# /etc/systemd/system/tramontana-integrity.service
[Unit]
Description=File integrity check with AIDE
Documentation=man:aide(1)
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/aide-check
Nice=19
IOSchedulingClass=idle# /etc/systemd/system/tramontana-integrity.timer
[Unit]
Description=Daily integrity check
[Timer]
OnCalendar=*-*-* 05:40:00
Persistent=true
RandomizedDelaySec=600
[Install]
WantedBy=timers.target$ sudo systemctl daemon-reload
$ sudo systemctl enable --now tramontana-integrity.timer
$ systemctl list-timers tramontana-integrity.timer --no-pager
NEXT LEFT LAST PASSED UNIT ACTIVATES
Tue 2026-08-19 05:44:12 CEST 18h left - - tramontana-integrity.timer tramontana-integrity.service05:40 is not arbitrary: the backup finishes before that time and the purge runs on Mondays at 05:10, so the integrity check sees a system at rest and does not compete for I/O — the same reasoning you applied in 05-07 when you moved the backup to 02:30. Nice=19 and IOSchedulingClass=idle are the additional guarantee.
Rootkit detection: rkhunter and chkrootkit
A rootkit is software installed after the compromise to maintain access and hide its presence, typically by replacing system binaries (ps, ls, netstat) or loading a kernel module that lies to user space. It is the reason the golden rule at the end of this lesson exists: if the system lies about its own state, no tool running inside it can be trusted.
$ sudo apt install rkhunter chkrootkit
$ sudo rkhunter --propupd # baseline of the binaries' properties
$ sudo rkhunter --check --skip-keypress
[ Rootkit checks ]
Rootkits checked : 479
Possible rootkits: 0
[ Applications checks ]
Warning: The SSH configuration option 'PermitRootLogin' has not been set
to 'no'.
Warning: Package manager verification has failedAnd here comes the real lesson of these tools: both warnings are false positives, and knowing why they are is the job.
-
The first is a limitation of
rkhunter's parsing: you did configurePermitRootLogin noin 06-02, but you did it in a file under/etc/ssh/sshd_config.d/, which the tool does not read. It is verified with the authoritative source,sshd -T, which shows the effective configuration:$ sudo sshd -T | grep -i permitrootlogin permitrootlogin no -
The second is explained by the openssl update: the binaries changed after you ran
--propupd. Once validated withdebsums, the properties database is regenerated.
chkrootkit covers similar ground with different heuristics and it commonly flags interfaces in promiscuous mode or hidden processes that are in fact artefacts of virtualisation. The operational conclusion is that these tools are a complement, not the core: they are run periodically, their warnings are investigated one by one, and no action is ever automated on the basis of them.
auditd: the kernel's audit subsystem
AIDE answers "what has changed?", but not "who changed it, when and with which process?". That is auditd, the Linux kernel's audit subsystem: it records system calls and file accesses at the moment they happen, with the real user, the effective user, the PID and the command.
Persistent rules go in /etc/audit/rules.d/, and augenrules compiles them at boot:
# /etc/audit/rules.d/50-tramontana.rules
# -w path -p permissions -k key
# p: r read, w write, x execute, a attribute change
# k: label to search for later with ausearch -k
# The application's configuration: it contains credentials
-w /etc/tramontana/app.conf -p wa -k tramontana_conf
-w /etc/tramontana/ -p wa -k tramontana_conf
# Identities and privileges
-w /etc/passwd -p wa -k identities
-w /etc/shadow -p wa -k identities
-w /etc/group -p wa -k identities
-w /etc/sudoers -p wa -k privileges
-w /etc/sudoers.d/ -p wa -k privileges
# Authorised SSH keys: the most common persistence mechanism
-w /home/operator/.ssh/ -p wa -k ssh_keys
-w /root/.ssh/ -p wa -k ssh_keys
# Scripts that run with privileges
-w /home/operator/scripts/ -p wa -k operator_scripts
-w /usr/local/sbin/ -p wa -k local_binaries
# System calls: loading kernel modules (a classic rootkit signal)
-a always,exit -F arch=b64 -S init_module,finit_module,delete_module -k kernel_modules
# Ownership and permission changes made by ordinary users
-a always,exit -F arch=b64 -S chmod,fchmod,fchmodat -F auid>=1000 -F auid!=unset -k permission_changes$ sudo augenrules --load
$ sudo auditctl -l | head -4
-w /etc/tramontana/app.conf -p wa -k tramontana_conf
-w /etc/tramontana -p wa -k tramontana_conf
-w /etc/passwd -p wa -k identities
-w /etc/shadow -p wa -k identitiesQuerying what auditd has recorded
Remember that in 05-02 you set chattr +i on app.conf. Look at what turned up:
$ sudo ausearch -k tramontana_conf -ts today -i | tail -12
type=PROCTITLE msg=audit(18/08/26 09:47:31.882:1043) : proctitle=vim /etc/tramontana/app.conf
type=PATH msg=audit(18/08/26 09:47:31.882:1043) : item=0 name=/etc/tramontana/app.conf
inode=262149 dev=fd:00 mode=file,640 ouid=root ogid=tramontana
type=SYSCALL msg=audit(18/08/26 09:47:31.882:1043) : arch=x86_64 syscall=openat
success=no exit=EPERM(Operation not permitted) auid=luis uid=luis gid=luis
euid=luis comm=vim exe=/usr/bin/vim key=tramontana_confRead it slowly, because it contains a complete incident: the user luis tried to open app.conf for writing with vim, and the openat call failed with EPERM. The immutable attribute did its job. And — this is what auditd adds over any other control — the attempt was recorded with a name, a time, a PID and a command, even though it never produced any change AIDE could see.
It is not necessarily malicious: the most likely explanation is that Luis wanted to raise max_connections to resolve the db_timeout errors and did not know about the chattr. But it is a piece of information you did not have before, and the conversation it prompts — channelling configuration changes through deploy-safe instead of editing by hand in production — is exactly the value of a detective control.
Aggregate reports come out with aureport:
$ sudo aureport --file --summary -ts this-week | head -6
File Summary Report
===========================
total file
===========================
7 /etc/tramontana/app.conf
3 /home/operator/.ssh/authorized_keysAnd that second line points once more at the same thing AIDE detected. Two independent controls pointing at the same place is a strong signal.
The cost of auditing too much
auditd is not free. Every rule adds work on the system call path, and a broad rule over a busy path can degrade performance noticeably and fill the disk.
| Rule | Effect |
|---|---|
-w /etc/tramontana/ -p wa |
Negligible cost: little activity |
-w /var/log/ -p wa |
Very expensive: every write of every log generates an event |
-a always,exit -S all |
Unusable in production |
Watch what matters, measure the volume (du -sh /var/log/audit/) and control the retention in /etc/audit/auditd.conf with max_log_file, num_logs and max_log_file_action. And once the configuration is settled, -e 2 at the end of the rules makes them immutable until the next reboot: not even root can change them, which stops an attacker disabling the auditing before acting. The price is that a legitimate change requires a reboot.
Log analysis for detection
You have had a persistent journald since 05-06 and you know how to drive journalctl. What is missing is knowing what to look for. These are the signals an administrator reviews, with the query that gets them:
# 1. Failed authentications grouped by source
$ sudo lastb -F | awk '{print $3}' | sort | uniq -c | sort -rn | head -5
47 203.0.113.44
3 10.0.2.31
# 2. Accepted SSH sessions: who got in, from where and with which method
$ sudo journalctl -u ssh --since "7 days ago" | grep -E 'Accepted' \
| awk '{print $1, $2, $3, $9, $11, $7}' | tail -5
Aug 18 08:12:04 operator 10.0.2.31 publickey
Aug 18 09:41:57 luis 10.0.2.44 password
# 3. Logins outside normal hours (before 07:00 or after 21:00)
$ sudo journalctl -u ssh --since "30 days ago" -o short-iso | grep 'Accepted' \
| awk -F'T' '{split($2,h,":"); if (h[1] < 7 || h[1] > 21) print}'
# 4. Use of sudo: what was run, by whom and from where
$ sudo journalctl --since today | grep -E 'sudo:.*COMMAND' | tail -3
# 5. Changes to identities and privileges
$ sudo journalctl --since "7 days ago" | grep -E 'useradd|usermod|groupadd|passwd\['Two observations about these results, which are what turn a query into a decision:
- Line 2 shows that
luislogged in withpassword. But in 06-02 you configuredPasswordAuthentication no. The fact that this authentication succeeded means there is an exception in someMatchblock undersshd_config.d/, or that it was applied to an interface you did not review. It is the outstanding item 06-03 left open, and now it has evidence: it has to be resolved, not merely met with a request that Luis use his key. - The 47 attempts from
203.0.113.44are already blocked byfail2ban, but the count is still useful as a baseline: if tomorrow it is 4,700, or if they appear from a new range, the change in shape is the signal.
The technique you are applying is the same one from Module 3 — grep, awk, sort | uniq -c | sort -rn — over a different source. And the reason you insisted on a persistent journal in 05-06 is precisely this: without it, query number 3 would have no 30 days of history to look at.
One important warning: the machine's own logs are tamperable evidence. An attacker with root can delete journal entries. That is why the log centralisation mentioned in 05-06 is not an organisational luxury: sending the logs to another machine the moment they are generated is what stops them being deleted after the fact.
NIDS: an honest overview of Suricata
Suricata is the reference NIDS in free software: it inspects network traffic, compares it against a catalogue of signatures (the free ET Open set from Emerging Threats is the usual starting point) and logs or blocks the matches.
Where it is placed determines what it sees:
graph LR
I[Internet] --> R[Router / firewall]
R -->|copy of the traffic<br/>IDS mode| S[Suricata]
R --> SW[Internal network]
SW --> A[srv-tramontana]
SW --> B[laptop-luis]
S -.->|alerts| L[Central log store]
In IDS mode it receives a copy of the traffic and only warns. In IPS mode it sits in the traffic's path and can drop packets, with the associated risk: if it gets it wrong or falls over, it cuts the service.
And now the honest part, because setting up Suricata on srv-tramontana would be a lapse of judgement at this point:
| Situation | What a NIDS adds |
|---|---|
| A single server, HTTPS traffic encrypted end to end | Very little: it cannot inspect what it cannot decrypt |
| Several machines and a single point the traffic passes through | A great deal: it is the only control that sees the whole network |
| A need to detect lateral movement between machines | It is the right tool |
In a single-machine infrastructure, the effort invested in AIDE and auditd pays off considerably more. When Tramontana grows to several servers — and in 07-07 you will see how — the NIDS goes on the list.
Configuration auditing with Lynis
A different axis: instead of detecting activity, Lynis evaluates the system's configuration against a catalogue of good practices and returns a score with suggestions.
$ sudo apt install lynis
$ sudo lynis audit system --quiet
...
Hardening index : 68 [############# ]
Tests performed : 267
Suggestions : 31
Warnings : 2How to read it, which is less obvious than it looks:
- The hardening index is not an exam mark and there is no "pass" value. It is a tracking metric: what matters is its trend and that it does not drop without you knowing why. A 68 on a server with a defined purpose can be perfectly correct; a 95 achieved by applying suggestions without understanding them is worse.
- The suggestions are not applied blindly. Lynis evaluates against a generic profile and knows nothing about your machine's purpose. Some of its recommendations here would be actively counterproductive.
A real example from each category:
| Lynis suggestion | Reasoned decision |
|---|---|
| Install an audit daemon | Already done: auditd is active. The warning comes from a check that runs before installing it |
Configure noexec on /tmp |
Accepted, but verified first: some installers fail (this is done in 06-06) |
| Install an antivirus (ClamAV) | Rejected: on a server with no third-party files it adds little and consumes memory that MemoryMax=512M does not have to spare |
| Disable IP forwarding | Rejected: it will be needed when the VPN is built in 08-04 |
The outcome of the exercise is not a high score, but a report with documented decisions. That is what is handed to Marta and what stands up to an audit, whereas "we applied everything the tool said" stands up to none.
Lynis is run recurrently — monthly, with its own timer — and its index is recorded alongside the performance baseline from 05-07.
Incident response and legal obligations
You have a real detection: an authorized_keys2 nobody created. What you do in the next twenty minutes determines whether you keep the information needed to understand what happened. This is the procedure, and the order matters:
-
Do not power the machine off. Powering off destroys all the evidence that lives in memory: processes, open connections, files deleted but still open, keys in the clear. Isolate it instead: cut the traffic with the firewall or disconnect the virtual network interface from the hypervisor, keeping console access.
-
Write down the time and the finding, off the server. From here on, every action is recorded with its time. This log is what later lets you tell your own footprints from the attacker's.
-
Preserve the volatile evidence before touching anything, saving the output off the machine:
$ ss -tunap # connections and which process they belong to $ ps auxf # the full process tree $ sudo lsof -n # open files and sockets $ sudo ls -l /proc/*/exe 2>/dev/null | grep deleted # deleted binaries still running $ w; last -F | head -20 $ sudo cp -a /var/log /media/evidence/ # and the journal: journalctl -o exportThe last query is especially valuable: a process whose binary no longer exists on disk but is still running is a very strong signal.
-
Take a snapshot of the disk with the machine still powered on. On your VM that is a hypervisor operation; on LVM, a snapshot like the one in 05-04. That image is the copy the investigation works on, so as not to alter the original.
-
Determine the scope: what was accessed, since when, and whether personal data is involved.
auditdand the journal are the sources. The question you have to be able to answer is whetherbookings.csv— with guests' names in it — was read or exfiltrated. -
Communicate. Marta first, and as soon as there is any suspicion of access to personal data, the security officer and the data protection officer. This is not a technical decision.
-
Recover by reinstalling.
The chain of custody is the concept underpinning steps 2 to 4: a record of who had access to each piece of evidence, when, and what they did with it, together with the checksums that prove it has not been altered. Without it, the evidence serves to understand what happened but not to support a claim or a criminal complaint.
The rule nobody wants to hear
A compromised server is reinstalled, not cleaned.
That is not professional pessimism, it is arithmetic. To "clean" it you would have to prove that you have found every persistence mechanism: every replaced binary, every added systemd unit, every crontab line, every authorised key, every kernel module, every library preloaded through LD_PRELOAD. A competent attacker leaves several, some of them hard to find. And the tools you would search with run on top of the very system that may be lying to you.
That is why all the work of Module 5 has a value that was not obvious at the time:
- The
resticbackup (05-08) lets you restore the data, not the compromised system. - The runbook kept off the server contains the rebuilding procedure.
- The documented configuration — netplan,
sshd_config,ufwrules, systemd units,app.conf— lets you rebuild the machine. - The 8-hour RTO Marta approved is exactly the time budget for this operation.
Reinstalling is the fast, safe option, not the drastic one. And the rebuilding from code that you will see in 07-06 with Ansible turns those 8 hours into considerably less.
Legal and compliance obligations
- Breach notification (GDPR, art. 33). If there is a security breach affecting personal data, the organisation must notify the supervisory authority within a maximum of 72 hours from becoming aware of it, and in some cases the affected individuals as well.
srv-tramontanastores guests' names inbookings.csv: the scenario applies in full. The clock starts from awareness, not from when you finish the investigation, so the communication in step 6 cannot wait. - Limits on monitoring people. The
auditdrules that watch whatluisdoes record the activity of an identifiable working person. That is subject to employment and data protection law: it requires a legitimate purpose, proportionality and — decisively — prior information to the people affected. Auditing staff activity in secret is not a decision that belongs to the administrator. - In a real environment, both the design of the detection and any decision about a breach must be reviewed by the security officer and the data protection officer. This course gives you the technical tools; it does not replace that judgement or a formal audit.
The report for Marta
Following the course's convention — saying what each measure protects and what it does not — this lesson's deliverable:
To be fixed:
- The unexplained
authorized_keys2: the incident response procedure is triggered; until it is closed, compromise is assumed. - Password authentication for
luis, which works despite being disabled: there is an exception in the SSH configuration that has to be located and removed. - Configuration changes made by editing in production: they are channelled through
deploy-safe, andauditdverifies that this is complied with.
Accepted, with reasons:
- Without a NIDS, there is no visibility of network traffic. Accepted while there is a single server; to be reviewed when there are several.
- Without centralised logs, an attacker with root can delete the journal. Accepted at the current cost; it is the first investment when the budget allows.
What this protects and what it does not. AIDE and auditd detect changes and accesses, and therefore shorten the time to detect a compromise, which is the variable that determines the damage. They do not prevent it. And neither of them protects against an attacker who arrives with valid credentials and modifies nothing — for that, the secrets have to stop being where they are.
Common Mistakes and Tips
- Keeping the AIDE database only on the watched server. It is the mistake that voids the entire measure. The database, or at least its checksum, has to be somewhere else.
- Watching files that legitimately change with
sha256. It produces a report full of changes every day, and within a week nobody reads them. Alert fatigue is the leading cause of death of a detection system: if it always rings, it never rings. - Updating the AIDE database without investigating the change. Running
aide --updateafter every report turns the tool into a history log with no ability to alert. First each change is classified, then the database is updated. - Auditing too much with
auditd. A broad rule over/var/logor/procdegrades performance and fills the disk. Start narrow and widen with judgement, measuringdu -sh /var/log/audit/. - Trusting a single control. AIDE detected the added file and
aureportconfirmed it independently. Two sources that agree give a confidence neither gives on its own. - Powering off the machine on detecting a compromise. It is the natural reflex and it destroys the memory evidence. Isolate the network, keep the machine powered on.
- Treating false positives as noise to be silenced. Every
rkhunteror Lynis warning you dismiss must be documented with its reason. Silencing without recording is how the one warning that was real gets lost. - A tip on method. Keep the detection queries in a script (
security_check.sh) that useslib/common.shand runs from a timer. A check that depends on your memory is not a control.
Exercises
Exercise 1
Design the AIDE configuration to watch /home/operator/bin/, the directory that contains the deploy-safe wrapper with 0755 permissions owned by root. Justify the attribute group you choose and explain which specific attack your rule would detect that watching only the checksum of the content would not.
Exercise 2
Write the auditd rule that records any execution of the binary /home/operator/bin/deploy-safe, and the ausearch query that shows who has run it today. Explain what information this adds that the deploy.log written by deploy.sh does not already provide.
Exercise 3
aide --check reports a change in /etc/tramontana/app.conf: the content attribute and the mtime have changed, and the permissions have gone from 640 to 644. Describe the complete investigation procedure, saying which sources you would consult and in what order, and what decision you would take in each of the two possible outcomes.
Solutions
Solution 1
# /etc/aide/aide.conf.d/99_tramontana (addition)
TramoAll = p+i+n+u+g+s+m+c+md5+sha256
/home/operator/bin$ TramoAllThe full group, and not just sha256, for three concrete reasons, each associated with a different attack:
p(permissions).deploy-safeis in thesudoersrule as a command permitted tooperator. If an attacker manages to change its permissions to0777, any user on the system can rewrite its content and, through thesudorule, execute arbitrary code with privileges. The content would not have changed yet, so a rule that only checkssha256would see nothing. This is the attack the answer has to identify.uandg(owner and group). The file must beroot:root. If it becomes owned byluisor by thetramontanagroup, its owner can modify it with no need for privileges, and once again the content still matches.i(inode) andn(links). An inode change with the same content means the file was replaced, not edited — for example, swapped for a symbolic link to another binary. A link count going from 1 to 2 indicates somebody created a hard link to the file, which allows them to keep access to the original binary even if the one in/home/operator/bin/is replaced.
In short: the checksum detects the modification of the binary, but the attributes detect the preparation for modifying it, which happens earlier and is the moment when detection can still prevent the damage.
Solution 2
# /etc/audit/rules.d/50-tramontana.rules (addition)
-a always,exit -F arch=b64 -F path=/home/operator/bin/deploy-safe \
-F perm=x -k deploy_executionReading the rule: -a always,exit always records, on leaving the call; -F arch=b64 limits it to 64-bit binaries (necessary because filtering by path with perm operates on the architecture); -F path= is the specific file; -F perm=x restricts the recording to executions and not to reads or writes; -k labels the event.
What it adds over deploy.log, which is the substance of the exercise. deploy.sh writes into its log whatever it itself decides to write, and only if it gets as far as running:
| Situation | deploy.log |
auditd |
|---|---|---|
| A normal deployment | Records it | Records it |
| The script fails before opening its log | No trace | Records the execution |
| The script is replaced by another binary | Records whatever the attacker wants | Records the real execution, with uid, auid and PID |
Somebody deletes deploy.log |
It disappears | The record is in /var/log/audit/, with immutable -e 2 rules |
The essential difference is one of trust: deploy.log is the application's testimony about itself; auditd's record is generated by the kernel, underneath the process, and does not depend on the good faith or the correct working of what is being audited. On top of that, auditd keeps the auid — the user who started the original session — which survives identity changes through sudo and answers "who was it really?" when uid is already root.
Solution 3
The permission change from 640 to 644 is the relevant part: it means that any user on the system can now read the file, and that file contains db_password in the clear. It is treated exactly like an exposed credential, just like the incident with the backup with 644 permissions.
The procedure, in this order and for this reason:
-
auditdfirst, because it answers who and when, and it is the source an attacker would find hardest to tamper with:$ sudo ausearch -k tramontana_conf -ts today -iLook for the successful
chmod/fchmodatevent and note theauid,uid,comm,exeand the exact time. -
Correlate with the sessions, to find out where that user came from:
$ sudo journalctl -u ssh --since today | grep -E 'Accepted|Disconnected' $ sudo journalctl --since today | grep -E 'sudo:.*COMMAND' -
Check whether there is a legitimate explanation: was there a deployment in that window? Does it coincide with
apt?$ tail -20 /var/log/tramontana/deploy.log $ grep -E "$(date +%Y-%m-%d)" /var/log/apt/history.log -
Assess the exposure, which is the question that determines the severity: for how long the file was readable, and whether anybody read it. The changed
mtimealso indicates that the content was modified too, so you have to see what was changed by comparing against the last backup:$ sudo restic dump latest /etc/tramontana/app.conf | diff -u - /etc/tramontana/app.confAnd look for reads of the file in that interval:
$ sudo ausearch -k tramontana_conf -ts recent -i | grep -E 'syscall=openat.*success=yes'
Outcome A — a legitimate change: the auid is operator, the time coincides with a recorded deployment, and the diff shows only max_connections adjusted. Even so there are two compulsory actions, because the result is wrong even though the intention was good: restore chmod 640 and chown root:tramontana, and fix the root cause — the procedure or the script that leaves the file at 644 — so that it does not happen again. It is documented and the AIDE database is updated. The underlying lesson: an authorised change with an insecure result is still a finding.
Outcome B — no explanation, or an unexpected auid, or a diff with changes nobody recognises: it is treated as a confirmed compromise. The full incident response procedure is applied (isolate, do not power off, preserve volatile evidence, snapshot the disk, determine the scope, communicate). And regardless of how the investigation ends, the db_password is considered compromised and is rotated immediately, because it was readable by the whole system for a period you cannot bound with certainty. Since personal data is involved, the communication to the data protection officer falls within the 72-hour deadline.
Both outcomes reach the same structural conclusion: while the password sits in the clear in a configuration file, any permission failure is a leak. That problem is not solved by watching the file more closely.
Conclusion
You have gone from a server that defends itself to a server that also observes itself. AIDE watches the integrity of the configuration, the binaries and the scripts, with its database protected off the machine and a daily check from a timer that only speaks when there is something to say. auditd records in the kernel who touches app.conf, who modifies identities and privileges, and who tries to load a module — and it has already given you a real finding no other control could see: Luis's failed attempt against an immutable file. You know how to read the logs looking for specific signals instead of glancing over them, you know that rkhunter and Lynis are interpreted and not obeyed, and you have an incident response procedure that does not depend on improvisation, uncomfortable rule included: a compromised server is reinstalled, not cleaned.
And detection has done exactly what is asked of it: it has told you where the underlying problem is. The three findings of this lesson point at the same place. The unexplained authorized_keys2, the luis session authenticated by password, and above all the conclusion of the last exercise: while db_password lives in the clear inside app.conf, any permission failure — yours, a script's, an attacker's — is a credential leak, and no amount of watchfulness over that file changes it. To that is added a second piece you have been dragging along since Module 5: the booking traffic, with the guests' names you have taken such care to protect in the backups, is still travelling unencrypted over port 8080. Lesson 06-05: Secrets Management and TLS Certificates resolves both: you will move the password out of the configuration file into an encrypted secret that systemd delivers to the service and nobody else can read, you will learn to manage keys with GPG and pass and to encrypt at rest with LUKS, and you will build the cryptographic material for bookings.tramontana.example — key, CSR, Let's Encrypt certificate and expiry monitoring — so that the day the reverse proxy comes into play in Module 8, encryption in transit is already solved and verified.
Linux Course: From Beginner to System Administrator
Module 1: Introduction to Linux
- What Is Linux?
- History of Linux
- Linux Distributions
- Installing Linux
- First Contact with the System
- The Linux File System Structure
Module 2: Basic Linux Commands
- Introduction to the Command Line
- Getting Help and System Documentation
- Navigating the File System
- File and Directory Operations
- Viewing and Editing Files
- Hard and Symbolic Links
- File Permissions and Ownership
Module 3: Advanced Command-Line Skills
- The Shell Environment: Variables, Aliases and History
- Using Wildcards and Regular Expressions
- Searching Files and Content: find, locate and grep
- Pipes and Redirection
- Text Processing: cut, sort, uniq, sed and awk
- Process Management
- Scheduling Tasks with Cron
- Networking Commands
Module 4: Shell Scripting
- Introduction to Shell Scripting
- Variables and Data Types
- Script Input, Output and Arguments
- Control Structures
- Functions and Libraries
- Debugging and Error Handling
- Production Scripts: Best Practices
Module 5: System Administration
- User and Group Management
- sudo and Special Permissions
- Package Management
- Disk Management
- systemd and Service Management
- System Logs: journald and syslog
- System Monitoring and Performance Tuning
- Backup and Restore
Module 6: Networking and Security
- Network Configuration
- SSH and Remote Access
- Firewalls and Perimeter Security
- Intrusion Detection Systems
- Secrets Management and TLS Certificates
- Securing Linux Systems
Module 7: Advanced Topics
- The Boot Process and System Recovery
- Advanced Diagnostics: strace, perf and eBPF
- Linux Kernel Tuning
- Virtualization with Linux
- Linux Containers and Docker
- Automation with Ansible
- High Availability and Load Balancing
