Exit, Zombies, and Orphans
A process's exit status has to survive the process, because POSIX gives the parent the right to ask
for it — wait() and its relatives — and there is no guarantee the parent has asked yet, or ever will,
by the instant the child actually finishes running. The kernel's answer to "keep the status around for a
question that hasn't been asked" is to keep a corpse: no memory, no address space, no open files, no
usable resources of any kind — just enough of a task_struct left standing to answer one question,
correctly, whenever the parent gets around to asking it.
What exit() releases immediately
A task calling exit_group() (the glibc-level exit() you write ends up here for a multi-threaded
process; kernel/exit.c's do_exit() is the underlying path either way) releases almost everything it
holds well before it becomes a zombie, and the order matters — it is exactly why a process can already
be showing Z in ps while every byte of its memory has already been given back to the system. Read
directly out of do_exit() at v6.18, the release sequence is: signal-related bookkeeping first
(exit_signals(), which sets the PF_EXITING flag other code checks to know not to touch this task
anymore), then the address space (exit_mm()), then System V IPC attachments (exit_sem(),
exit_shm()), then the open-file-descriptor table (exit_files()), then namespaces
(exit_task_namespaces()), and only after all of that does exit_notify() run — the function that
actually marks the task a zombie. By the time a task is visible as Z, its address space is gone, its
file descriptors are closed, its POSIX timers are canceled — there is nothing left to reclaim from it
except the shell that remains.
What it cannot release
Three things persist past exit_notify(): the task_struct itself (reduced, but not freed), the PID
(kept alive via struct pid's own refcount specifically so a concurrent waitpid() can still name it),
and the exit status (the code or signal the process died from). Say the actual cost plainly, because the
word "zombie" invites exaggeration: a zombie occupies a few kilobytes of kernel memory and one slot in
the PID space. It does not hold memory in any quantity that matters, has no threads, and consumes no CPU
time.
The real problem thousands of zombies cause is not memory — it is PID exhaustion. The PID space is
finite (/proc/sys/kernel/pid_max, commonly 4194304 on a 64-bit system but frequently configured much
lower), and every zombie is holding one PID that cannot be reused until it is reaped. A parent that leaks
zombies fast enough — a process-pool server that forks constantly and never calls wait() — can run the
counter dry, at which point fork() itself starts failing for every process on the system, not just
the buggy one. A handful of zombies is cosmetic; a systemic failure to reap is an availability incident.
What actually happens
Here is a parent that forks a child, and the child exits immediately while the parent never calls
wait():
#include <stdio.h>
#include <unistd.h>
#include <sys/types.h>
int main(void) {
pid_t pid = fork();
if (pid == 0) {
_exit(0); /* child: exit immediately */
} else {
printf("parent %d, child %d, sleeping without wait()\n", getpid(), pid);
fflush(stdout);
sleep(20); /* parent: never reaps */
}
return 0;
}
Compiled and run in this environment (gcc -O0 -o zombie zombie.c && ./zombie &), ps shows the child
sitting in Z state for the full twenty seconds the parent sleeps:
$ ps aux | grep zombie
dev 14603 0.0 0.0 2776 1892 ? SN 01:47 0:00 ./zombie
dev 14605 0.0 0.0 0 0 ? ZN 01:47 0:00 [zombie] <defunct>
Real, captured output — not invented. The child (PID 14605) shows STAT ZN: Z for zombie,
N for a positive nice value inherited from the parent's shell. Its RSS and VSZ columns both read
0 — no memory is charged to it — and its command name is rewritten in brackets with the <defunct>
marker, the kernel's own way of telling ps "this is a name, not a running program." When the parent
(PID 14603) was killed, the reparented zombie was immediately reaped by its new parent, confirming there
was nothing further holding it beyond the parent's neglect:
$ kill -9 14603
$ ps -eo pid,ppid,stat,comm | grep -E "14603|14605"
(no output — both gone)
There are exactly two correct fixes, and one common wrong one. The first correct fix is to install a
SIGCHLD handler that calls wait() or waitpid() (typically in a loop with WNOHANG, to drain
however many children exited at once) — reaping happens as a side effect of the handler running. The
second, Linux- and POSIX-standardized but far less known, is to explicitly set SIGCHLD's disposition to
SIG_IGN:
signal(SIGCHLD, SIG_IGN);
This is not "ignore the notification and hope for the best" — it is a documented instruction to the
kernel meaning "auto-reap my children the instant they exit; I am not going to ask for their exit
status." exit_notify() checks exactly this at the moment a child would become a zombie and, when the
parent's disposition says so, sets exit_state straight to EXIT_DEAD and calls release_task()
inline — the zombie state is skipped entirely rather than created and then cleaned up.
The fix that does not work, and needs to be said plainly because the instinct is so natural: you cannot
kill a zombie. kill -9 delivers a signal to a running process so it can be acted on; a zombie has
already exited; there is no execution left to interrupt, no pending-signal check that will ever run, no
handler to invoke. kill -9 <zombie-pid> either fails (ESRCH, if the PID has already been fully
reclaimed) or silently sets a pending bit on a task_struct that will never again reach a return-to-user
checkpoint to act on it. The bug is never in the zombie. It is always in the parent, and the fix is
always in the parent.
Reparenting
When a parent dies before its children do, those children do not become unowned — every process must
have a parent for wait() bookkeeping to make sense. forget_original_parent()
(kernel/exit.c) walks the dying task's children and reparents each one via find_child_reaper(), which
picks a new parent by a specific rule: the nearest ancestor that has marked itself with
prctl(PR_SET_CHILD_SUBREAPER, 1), or, if none exists in the child's PID namespace, PID 1 of that
namespace.
Subreapers exist because PID 1 being the sole fallback was too coarse for how modern systems are
structured. A session manager, an init system's service supervisor, or a container runtime often wants
to inherit its own orphans rather than have them silently vanish into the system's real PID 1, where
diagnosing "who orphaned this" becomes much harder. PR_SET_CHILD_SUBREAPER lets any process opt in to
that role for its own descendant tree without being PID 1 itself.
Why PID 1 must reap
Every orphan a system produces ends up, eventually, parented to some task_struct that must call
wait() on it — and if no subreaper claims that duty, that task_struct is PID 1. This is a genuine
obligation, not a formality: a PID 1 that never calls wait() accumulates every orphan the system ever
produces, each holding its small task_struct and its PID slot, for as long as the system runs. This is
exactly the bug people hit when they run an ordinary application binary directly as a container's PID 1
— the application was never written to reap children it does not know it has, and every process the
container spawns and abandons becomes a permanent zombie under it. The duty a container's actual init
process (prose reference only — the design and responsibilities of that process belong to the container
and orchestration material later in this documentation) exists to take on is precisely this one.
Orphaned process groups and SIGHUP
One more piece completes the picture of what happens when a controlling process disappears. When the
process that led a foreground process group exits — a shell, when its controlling terminal is closed —
and that exit leaves the group "orphaned" (no member has a parent left in the same session), the kernel
sends SIGHUP to every stopped member of that group, and this is exactly why closing a terminal
window kills whatever job was running in it, and exactly why nohup (which sets SIGHUP's disposition
to SIG_IGN before running the command) and setsid (which detaches the command into its own new
session entirely, so the terminal closing never touches it) are the two standard ways to keep a
long-running job alive past its terminal's lifetime.
Misconceptions
- "Zombies leak memory." A zombie holds a PID slot and a small
task_structremnant — its memory was already released before it became a zombie at all. The real resource pressure is PID exhaustion, not memory. - "You can
kill -9a zombie." There is nothing left to signal — a zombie is not executing and never will again. The correct target for a signal is the parent, to make it callwait(). - "Orphans become zombies." The opposite: orphans are reparented (to a subreaper or PID 1) and, by convention, promptly reaped by whatever process takes them on. An orphan that lingers as a zombie means its new parent is also failing to reap — the same bug, one level up.
The two ways a task's remains are cleaned up: the parent asks, or the parent dies and someone else inherits the duty.
References
wait(2)— the reaping interface, including theSIGCHLD-to-SIG_IGNauto-reap behavior, which is Linux-specific and rarely known.prctl(2), thePR_SET_CHILD_SUBREAPERsection — the mechanism that made container init processes practical.kernel/exit.c:do_exit()— the release order, read in the order it actually happens; verified at v6.18 to runexit_signals()→exit_mm()→exit_sem()/exit_shm()→exit_files()→exit_task_namespaces()→exit_notify().kernel/exit.c:exit_notify()andkernel/exit.c:release_task()— the zombie-creation and reaping functions themselves;exit_notify()is also where theSIGCHLD == SIG_IGNautoreap fast path (skippingEXIT_ZOMBIEand going straight toEXIT_DEAD) actually lives.credentials(7)andsetsid(2)— for the process-group andSIGHUPsection above.