The most detailed free FDE + DevOps library: 140+ lessons, 70+ labs and 80 long-form articles, in English and Turkish. Start learning →

DevOps Foundations · Module 1: Linux working model

Disk, Inodes and Memory

Three resources run out with identical symptoms and different cures. This lesson teaches the three-way split so cleanup fixes the cause instead of deleting the evidence.

10 min reading

Objectives

  • Tell disk-full apart from inode-full apart from out-of-memory using three commands
  • Find what is actually consuming space without guessing or deleting blindly
  • Explain why deleting an open log file frees nothing
  • Apply a safe cleanup order that never touches data you cannot restore

Why this matters

Writes start failing, and every instinct says the disk is full. Half the time it is not: the filesystem has free blocks but zero free inodes, or plenty of both while the machine is out of memory and the OOM killer is choosing victims. Each wrong guess costs a restart or a deletion you cannot take back. The three-way check takes thirty seconds and has saved more data than any recovery tool.

Concepts

Disk space has two counters, and either one stops writes. Blocks hold file contents; inodes hold file metadata, one per file regardless of size. df -h reports blocks, df -i reports inodes, and a filesystem at 100 percent on either refuses new files. Millions of tiny session or cache files exhaust inodes while blocks look healthy; one runaway core dump does the reverse. Always read both before acting.

du finds the consumer: du -xhd1 /srv 2>/dev/null | sort -rh stays on one filesystem, summarizes one level deep, and sorts largest first. Deleted but open files are the classic trap: removing a name with rm only drops the directory entry while a process still holds the file open, so df shows no freed space until the holder exits or truncates. lsof | grep deleted names the holders; truncating with : > file or restarting the service actually releases the blocks. Compressing or moving to another filesystem has the same non-effect until the open handle closes.

Memory pressure looks different: the machine slows, then processes die with no disk error at all. free -h separates used RAM from reclaimable cache; high cache is healthy, high used with zero available is not. dmesg and the journal carry the OOM killer's verdict naming its victim. Swap delays the decision but never resolves it: a swapped service answers slowly enough to trip every timeout above it. The honest fix for memory exhaustion is less concurrent work or a bigger box, chosen with numbers, not a swappiness tweak copied from a forum.

Worked example

Writes to /srv fail. The split, in order:

$ df -h /srv | tail -1
/dev/sda1  20G  9.1G  9.9G  48% /srv
$ df -i /srv | tail -1
/dev/sda1  1.3M  1.3M   0   100% /srv
$ du -xhd1 /srv/app/tmp 2>/dev/null | sort -rh | head -3
4.0K  /srv/app/tmp/sess_9f2c
...
$ ls /srv/app/tmp | wc -l
1298433

Expected reading: blocks half free, inodes exhausted, 1.3 million session files in one directory. The cure is deletion of files the application regenerates, oldest first, in bounded batches (a single rm * with a million arguments fails on argument limits; find ... -delete in slices does not), followed by fixing the session cleanup job that caused it. Note what this example rules out: no memory evidence, no block pressure, so no reboot and no service restart beyond what the cleanup needs.

The common wrong move

Deleting by size ranking alone: the biggest file is often the database or the only backup on the machine, and it is exactly what cannot be restored. The safe order is: identify the consumer with du, confirm nothing important shares the directory, prefer truncating logs over deleting databases, delete regenerable data first, and verify with df after each step. Anything irreplaceable gets moved, never removed, until the incident is over. "The disk was full so I deleted the largest files" is how postmortems begin.

Lab and next step

Lab L03 makes you produce inode pressure inside a bounded training directory and clean it safely, with verification after every step; host disks stay untouched by construction. Next, lesson 4 turns these checks into a repeatable diagnosis method using logs.

Quick check

An optional 4-question self-check. Answers never leave your device, are not stored, and never count toward any assessment.

Lesson feedback

No published feedback yet.

Log in and complete the lesson to leave feedback.

Exercise

In /tmp only, create 50,000 small files with a loop, observe df -i change, then remove them in find batches and confirm recovery. Record the three commands and their before/after numbers.

Pass criteria

Inode use visibly rises then falls in df -i output; removal uses find (not rm *) ; before/after numbers recorded for at least one df -i read.

Sources

Log in to track progressFree account: stores only your lesson progress and quiz results.