Demand Paging and Copy-on-Write
The Page Fault Handler named do_anonymous_page() and do_wp_page() as
two leaves of its classification tree without dwelling on them. This page is that dwelling — it's the
policy those two functions implement, stated plainly: never do work until someone proves they need it.
Address space is free — a VMA is bookkeeping, a handful of bytes in a maple-tree node — and physical pages
are not. Linux issues the first liberally, on the promise that most of it will never be touched. This is
also, directly, why memory accounting on Linux confuses almost everyone who first encounters it: top,
ps, and /proc/PID/status are each answering a slightly different question, and the confusion is a
straight consequence of a deliberate design choice, not an accident of bad tooling.
First touch
mmap() — directly, or via malloc()'s large-allocation path — creates a VMA and maps no pages. The
first read of a freshly mapped anonymous page doesn't allocate anything either: do_anonymous_page()
maps the shared zero page into the faulting address, read-only. Only the first write
allocates a real, private, zeroed physical page.
This read/write asymmetry is worth stating explicitly, because it has a consequence programmers routinely
get wrong: a program that mmap()s a large region and only reads it — scanning for a value, say, without
writing — uses almost no physical memory no matter how large the region is. The first write is the event
that actually costs something.
The zero page
empty_zero_page is one physical page, filled with zeros, shared read-only by every process on the
system for every untouched anonymous read. Defined in arch/x86/kernel/head_64.S:
__PAGE_ALIGNED_BSS
SYM_DATA_START_PAGE_ALIGNED(empty_zero_page)
.skip PAGE_SIZE
SYM_DATA_END(empty_zero_page)
EXPORT_SYMBOL(empty_zero_page)
do_anonymous_page() maps it in for a read fault with pfn_pte(my_zero_pfn(address), ...) — the same
physical frame, mapped read-only into however many processes have touched however many untouched pages.
It's cheap (one physical page, system-wide, no matter how many mappings point at it) and elegant (it
requires no special-casing anywhere else: the PTE that results looks like an ordinary present, read-only
entry to everything downstream), and it's the reason calloc() of a large region can be nearly free —
calloc promises zeroed memory, and mapping the zero page read-only satisfies that promise for every page
the caller never writes to, without allocating or zeroing anything real.
COW after fork
folder 06 described fork()'s
copy-on-write behavior from the process side; this is the same mechanism in mm terms. fork() doesn't
duplicate physical pages — it duplicates the parent's page tables, marking every writable anonymous PTE
read-only in both the parent's and the child's tables (both processes now share the same physical
pages, temporarily, neither with write access). The first write by either process — X86_PF_WRITE set,
The Page Fault Handler's "present, read-only, write
access" leaf — lands in do_wp_page().
do_wp_page() doesn't unconditionally copy. It first checks whether this process is provably the only
remaining owner of the page — PageAnonExclusive(), or a folio refcount/mapcount check
(wp_can_reuse_anon_folio()) confirming no other mapping references it — and if so, reuses the same
physical page in place, just flipping the PTE back to writable. Only when the page is genuinely shared
does it allocate a fresh page and copy (wp_page_copy()). This distinction is worth a paragraph on its
own, because it's exactly why a fork()-then-immediately-exit() child (the shape of every system()
call, every shell pipeline, every process-per-request server before it execs something else) costs almost
nothing: by the time the parent or child writes anything at all, the other side has often already
released its reference, and the write becomes a cheap in-place reuse rather than a copy.
One anonymous page's states, and the two events that move it: the first read, and the first write.
What actually happens
Take malloc(100 * 1024 * 1024 * 1024) — a 100 GiB allocation — on a machine that does not have 100 GiB
of RAM, and walk what actually happens, with real numbers from the machine this page was written on
rather than an invented transcript.
This machine: MemTotal 20481144 kB (≈19.5 GiB), SwapTotal 10485760 kB (10 GiB) — call it ≈29.5 GiB
of RAM+swap combined — with vm.overcommit_memory = 0 (the default heuristic mode; see Overcommit
modes below).
glibc's malloc() routes any allocation past M_MMAP_THRESHOLD (a few hundred KB) directly to mmap()
rather than extending the heap with brk(). A real 100 GiB request on this machine:
$ ./malloc_demo 100
Requesting malloc of 107374182400 bytes (100 GiB)
malloc FAILED (returned NULL)
It fails — and this is itself the honest, real result on a machine with ≈29.5 GiB of RAM+swap, not a
manufactured "succeeds on 8 GB" scenario. Under overcommit mode 0, the kernel's __vm_enough_memory() heuristic refuses an "obvious"
overcommit — roughly, a request too large relative to total RAM+swap — and 100 GiB against ≈29.5 GiB of
combined RAM+swap trips that refusal. Binary-searching the actual boundary on this machine:
$ ./malloc_demo 30 → malloc FAILED (returned NULL)
$ ./malloc_demo 28 → malloc succeeded, base address = 0x742169dff010
28 GiB — comfortably past the 19.5 GiB of physical RAM, safely under the ≈29.5 GiB RAM+swap ceiling —
succeeds. That number is this page's "100 GB malloc on an 8 GB machine": a successful allocation for far
more memory than physically exists, made possible because a successful mmap() is a promise about address
space, not a guarantee of physical pages. The maps entry and RSS immediately after:
maps entry: 742169dff000-742869e00000 rw-p 00000000 00:00 0
--- after malloc, before touching anything ---
VmSize: 29362908 kB (≈28 GiB — the full request, reserved)
VmRSS: 1892 kB (≈1.8 MB — this process's baseline; none of it is the new mapping)
VmData: 29360360 kB
VmSize already reflects the entire 28 GiB — the address space exists. VmRSS — the resident set,
physical pages actually backing this process — is unchanged from before the call: the mapping is real to
the page tables' absence of entries, and to nothing else yet.
Then, touching one byte per gigabyte (28 touches, one per GiB of the mapping) and reading VmRSS after
each:
touched byte 1 (offset 0 GiB) -> VmRSS = 1892 kB
touched byte 2 (offset 1 GiB) -> VmRSS = 1896 kB
touched byte 3 (offset 2 GiB) -> VmRSS = 1900 kB
touched byte 4 (offset 3 GiB) -> VmRSS = 1904 kB
...
touched byte 28 (offset 27 GiB) -> VmRSS = 2000 kB
--- after touching one byte per GiB, all touches done ---
VmRSS: 2000 kB
RSS rises by exactly 4 kB — one PAGE_SIZE (getconf PAGESIZE confirms 4096 on this machine) — per touch,
for 27 of the 28 touches (1892 → 2000 kB is a 108 kB rise, 27 × 4 kB). The first touch (offset 0 GiB)
shows no visible rise: glibc's mmap-backed allocator writes a small chunk header into the start of a
large mmap()-obtained region as part of malloc()'s own bookkeeping, so the very first page of this
particular allocation was already resident before the demo ever wrote to it — a real, if slightly
surprising, artifact of glibc's allocator rather than a discrepancy in the kernel's accounting. Every
subsequent touch lands on a genuinely untouched page and costs exactly one page of RSS, matching
First touch precisely: a successful malloc() reserved 28 GiB of address space instantly,
and each gigabyte only became real, physical memory the instant this code wrote to it.
Overcommit modes
vm.overcommit_memory (/proc/sys/vm/overcommit_memory) controls how strictly the kernel enforces the
"can I actually back everything I've promised" question at allocation time, not at fault time:
0— heuristic (the default, and this machine's setting). The kernel refuses allocations it judges an "obvious" overcommit — the__vm_enough_memory()check demonstrated above — but is otherwise permissive: normal-sized allocations, even ones that collectively exceed physical RAM, are allowed on the expectation that most memory promised is never fully touched at once.1— always overcommit. No accounting check at all; everymmap()/brk()extension is allowed regardless of how much has already been promised. Useful for workloads (some scientific/HPC code) that deliberately reserve far more address space than they'll ever use and would rather never seemallocfail.2— strict. Total committed address space across the system is capped atswap + physical_RAM × overcommit_ratio / 100(vm.overcommit_ratio, default 50 on this machine — orvm.overcommit_kbytesfor an absolute figure instead of a percentage). Once that cap is reached, allocation requests fail outright.
The practical guidance is worth stating honestly rather than diplomatically: mode 2 moves allocation
failure from "the OOM killer picks a victim process at some unpredictable later moment, possibly not even
the process that over-allocated" to "malloc() returns NULL right now, at the call site that asked for
too much." That's exactly what some workloads (databases, anything that must never be silently killed)
want — and it's also a mode most ordinary software is not written to handle, because most software treats
"malloc never really fails" as an unstated assumption baked into decades of C and C++ code that never
checks the return value.
MAP_POPULATE, mlock, and pre-faulting
Three ways to demand physical pages now, at mapping time, instead of paying for each one lazily at first touch:
mmap(..., MAP_POPULATE, ...)asks the kernel to pre-fault the entire mapping as part of themmap()call itself — the pages are resident by the timemmap()returns, at the cost ofmmap()itself now doing all the work every future touch would otherwise have spread out.mlock()/mlockall()goes further: it pre-faults and pins the pages, preventing reclaim from ever evicting them.MAP_POPULATEalone doesn't stop the pages from later being reclaimed under memory pressure and re-faulted.- Manual pre-faulting — simply looping over the region writing a byte per page yourself — achieves the
same resident-and-mapped effect as
MAP_POPULATEwithout a dedicated flag, useful when the mapping already exists and can't be re-created with the flag.
Worth it for real-time or latency-sensitive work specifically, where a page fault mid-critical-section is a latency spike the workload cannot tolerate — audio processing, some trading systems, anything with a hard deadline. The cost is exactly what demand paging was designed to avoid: paying for pages that might never be touched, up front, unconditionally.
Where the accounting goes wrong
/proc/meminfo's Committed_AS — the sum of everything the kernel has promised, across every process,
under the current overcommit accounting rule — and CommitLimit — the ceiling that promise is checked
against under mode 2 (informational under modes 0 and 1) — are the two fields that actually answer "is
this machine over-promised," and neither is what most people reach for first. VSZ (virtual size, ps's
VSZ column, VmSize above) is nearly meaningless as a "how much memory is this process using" answer
precisely because of everything on this page: it counts address space reserved, not physical memory
resident, and a process can have a VSZ in the hundreds of gigabytes while genuinely using a few
megabytes. The full treatment of what to trust instead — RSS's own caveats included — belongs to
What free and RSS Really Tell You, not here.
Misconceptions
- "
mallocallocates memory." It reserves address space. First touch is what actually allocates a physical page, one page at a time, on the first write —malloc()'s return isn't the moment memory becomes real, it's the moment a promise is made. - "If
mallocsucceeded, the memory is mine." Under the default heuristic overcommit policy (mode 0, this machine's setting), a successfulmalloc()is a promise the kernel may not be able to keep in full. If the system as a whole is over-committed when the pages are actually touched, the failure doesn't arrive as aNULLfrommalloc— it arrives later, as a page fault the kernel cannot satisfy, and then as the OOM killer choosing a process to kill, which is not necessarily the process that asked for too much. - "COW means
fork()is free." It defers the cost, it doesn't eliminate it. COW after fork marks pages read-only and shares them — cheap — but the first write to each shared page still costs exactly what an ordinary first-touch write costs, deferred to whenever that write happens rather than paid all at once atfork()time.
References
- Overcommit accounting — the kernel's own definition of the three modes, including exactly what the mode 0 heuristic checks.
mm/memory.c:do_anonymous_page()— the first-touch path, including the zero-page shortcut for reads verified against v6.18 source above.man 2 mmap, theMAP_NORESERVEandMAP_POPULATEsections — the interface-level controls over this behaviour.man 5 proc, the/proc/meminfosection — forCommitted_ASandCommitLimit.