DevOps Foundations · Module 1: Linux working model
Processes, Signals and Services
A service is a process with a supervisor, a user, and a contract about how it stops. This lesson connects PIDs, signals and units so restarts become deliberate instead of hopeful.
11 min reading
Objectives
- Trace a running service from its PID to its command line, owner and open ports
- Choose the right signal for stop, reload and diagnosis instead of reaching for -9
- Read a systemd unit and predict what happens on start, crash and reboot
- Explain why kill -9 is a last resort and what it destroys
Why this matters
"Just restart it" resolves most incidents and teaches nothing. The restart that matters is the second one: the service dies again at 3 AM because the first restart never asked why the process exited, whether it flushed its state, or whether the supervisor will even bring it back. Process literacy turns restarts from a reflex into a procedure with a known end state.
Concepts
A process is a running program with four facts worth knowing: its PID, its parent (PPID), the user it runs as, and its current state. ps aux shows all four plus the full command line; /proc/<pid>/ exposes the same facts as files, which is why lesson 1 matters here too. Orphaned processes get adopted by PID 1, and zombies are already dead processes whose parents have not collected their exit status: they consume a PID slot and nothing else, and killing them is impossible because there is nothing left to kill.
Signals are the vocabulary for talking to processes. The number is not the point; the contract is. SIGTERM (15) asks the process to stop and lets it clean up: flush buffers, close connections, deregister. SIGHUP (1) by convention asks a daemon to reload its configuration without dropping work. SIGKILL (9) is not delivered to the process at all; the kernel destroys it immediately, which means no cleanup, no flush, and a real chance of a half-written file. kill without arguments sends SIGTERM, which is the correct default. Reaching past it should require a written reason.
systemd turns a process into a service: a unit file declares what to run, as whom, with which environment, and what to do when it fails. Read a unit in three parts. [Unit] names dependencies (start me after the network). [Service] names the command, the user, the restart policy and the environment, including EnvironmentFile lines that are the usual hiding place of boot failures. [Install] decides whether the service starts on boot at all. Restart=on-failure with a RestartSec delay is the difference between a self-healing service and a crash loop that hammers a dependency five times a second.
Worked example
A service is up but not responding, and the dashboard suggests a restart. The deliberate sequence:
$ ps -o pid,ppid,user,stat,cmd -p $(pgrep -f svc-app)
PID PPID USER STAT CMD
4182 1 svc-app Ssl /usr/bin/svc-app --config /srv/app/config.yaml
$ kill 4182
$ sleep 3; ps -p 4182
PID TTY TIME CMD
$ systemctl is-active svc-app
activeExpected reading: the process belonged to the service user, slept normally (S), and its parent was PID 1 through the unit, so SIGTERM gave it three seconds to exit and the supervisor started a fresh one. If the PID survives SIGTERM, escalate with evidence: check the state column (a D state means uninterruptible I/O wait, where even SIGKILL waits), capture the stack or at least the last log lines, then SIGKILL once and record why TERM failed. The incident record should show TERM first, the wait, and the reason for KILL.
The common wrong move
kill -9 as a first resort destroys exactly the evidence the diagnosis needs: buffered logs never flush, temp files stay behind, and lock files claim a holder that no longer exists, so the replacement process refuses to start and the incident doubles. The same reflex applied to a database or a queue consumer risks torn writes that surface days later as corruption with no trail. Treat -9 like a fire axe behind glass: available, documented when used, and never the daily tool.
Lab and next step
Lab L02 replays a boot failure caused by a wrong environment variable and a wrong service definition, diagnosed through the unit before anything is restarted. Next, lesson 3 covers what the machine itself runs out of: disk, inodes and memory.
Quick check
An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.
Lesson feedback
No published feedback yet.
Log in and complete the lesson to leave feedback.
Exercise
Write a one-page restart procedure for any service you run: the exact ps check, the TERM command, how long you wait, what you read if it survives, and the single condition under which you escalate to KILL. Run it once against a harmless process (e.g. sleep 600 in the background).
Pass criteria
Procedure names the PID check, the TERM command, a concrete wait time, the evidence read before escalation, and one explicit KILL condition; the trial run shows the PID gone after TERM.