Skip to main content

Updated Sep 14, 2026

Copying Data Across the Boundary

There is a rule that sounds like bureaucracy and is not: the kernel may never dereference a user pointer directly, even though on ordinary hardware it is perfectly capable of doing so — a plain *ptr on a user address would, mechanically, just work, right up until the moment it doesn't. Three separate things can go wrong with a pointer a syscall was handed: it may not point at valid memory at all, it may point at kernel memory the caller has no business reading or writing, or its contents may change out from under the kernel between two reads of it. The copy routines this page covers exist to turn each of those three into a handled, well-defined case, rather than a crash or a hole.

copy_from_user and copy_to_user​

include/linux/uaccess.h:copy_from_user() and include/linux/uaccess.h:copy_to_user() are the interface almost every syscall that touches a buffer goes through. Both take a kernel pointer, a __user pointer, and a length; both return the number of bytes that could not be copied — not a boolean, not an error code, a count. Zero means complete success. This trips up everyone once, because it means the natural-looking if (copy_from_user(...)) reads correctly (any nonzero remainder is truthy, i.e. a failure) but copy_from_user(...) == -1 or treating the return as a byte-count success value does not:

if (copy_from_user(&kbuf, ubuf, len))
return -EFAULT;

copy_from_user additionally zero-pads the destination on a short copy, so a partially-failed copy does not leave uninitialised kernel stack or heap bytes sitting behind a buffer the caller believes is fully populated.

access_ok, and what it does not check​

arch/x86/include/asm/uaccess.h:access_ok() checks exactly one thing: that the address range given falls entirely within the user half of the address space, as opposed to reaching into kernel addresses. That is all it checks. It does not check that the range is actually mapped to anything — whether a page exists at that address, whether it is present in physical memory right now, or whether the calling process has permission to touch it — because that is the page-fault handler's job, decided at the moment of actual access. This is precisely why access_ok alone is never sufficient by itself for a raw copy: passing access_ok only means the pointer is in range, not that dereferencing it is safe, which is why the copy routines still need the mechanism below even after the range check has passed.

The exception table​

This is the single most elegant thing in this folder, and it earns the space. A plain, unprotected mov instruction that faults on a bad address would, in ordinary kernel code, be treated as a kernel bug — an oops. But the instructions inside copy_from_user/copy_to_user that actually touch user memory are deliberately not ordinary: each one is registered in a table, arch/x86/include/asm/extable.h:exception_table_entry(), alongside a fixup address to jump to if that specific instruction faults:

struct exception_table_entry {
int insn, fixup, data;
};

(The three fields are relative offsets, not raw pointers — kept compact and position-independent so the table can be built into read-only, relocation-free memory.) When a page fault happens anywhere in the kernel, the fault handler calls arch/x86/mm/extable.c:fixup_exception(), which looks up the faulting instruction's address in this table. If it finds an entry, the fault is not a kernel bug at all — it is a user-copy routine that was handed a bad pointer, doing exactly what it is supposed to do when that happens. fixup_exception redirects execution to the registered fixup address instead of the instruction that faulted; that fixup code sets the appropriate registers so the copy routine returns the remaining byte count as if the copy stopped right there, and the caller's if (copy_from_user(...)) check turns that into -EFAULT. If the faulting address is not found in the table, fixup_exception returns failure and the fault handler falls through to treating it as the kernel bug it actually is.

The effect: a single mechanism turns "the kernel touched a bad user pointer" from an oops into an ordinary, checkable return value, without the copy routine needing a conditional branch around every single memory access.

SMAP and SMEP​

The exception table is a software safety net; SMAP and SMEP are hardware backing the same rule from underneath it. SMEP (Supervisor Mode Execution Prevention) stops the kernel from ever executing instructions that live on a user page — no amount of kernel-mode confusion about RIP can jump into attacker-controlled code. SMAP (Supervisor Mode Access Prevention) goes further and stops the kernel from even accessing user pages at all, except inside a narrow window explicitly opened with the stac instruction and closed again with clac — exactly the window the copy routines execute inside.

The consequence worth stating plainly: on a machine with SMAP enabled, an accidental, unprotected user dereference in kernel code — a bug that on older or SMAP-less hardware might have silently worked and gone unnoticed for years — faults immediately, every time, rather than quietly succeeding. SMAP turns a latent, hard-to-find class of bug into one that reproduces on the first bad access.

__user and sparse​

The __user annotation seen throughout this page (const void __user *from, and so on) is what lets a static checker catch, at build time, the same class of bug SMAP catches at run time: sparse treats __user as marking a distinct address space, and flags any attempt to dereference a __user pointer directly or assign it to a plain kernel pointer without going through a conversion function. The annotation itself, and the rest of the kernel's address-space and type-checking vocabulary sparse understands, belongs to the kernel C dialect — this page only uses it, it does not re-derive it.

Double-fetch bugs​

The copy interface also creates its own bug class. The shape: a syscall reads a length (or some other control value) out of user memory, validates it against some limit, and then — instead of using the already-validated copy — reads the same user memory a second time and acts on whatever it finds there. Between the first read and the second, another thread in the same process can change the value, and the kernel ends up validating one value while acting on a different one it never checked. This is a genuine, recurring, real-world CVE class — a time-of-check-to-time-of-use (TOCTOU) bug specific to the user/kernel boundary — not a theoretical concern. The rule that avoids it: copy the data into kernel memory exactly once, and validate the kernel-side copy from then on. Never re-read the user pointer after validation.

Structures that grow​

include/linux/uaccess.h:copy_struct_from_user() generalises the same idea to whole structs, and is the mechanism behind the size-argument convention mentioned when arguments beyond the register limit were introduced. It takes a kernel destination and its size, a user source, and the caller-supplied size of the user struct, and handles the two cases that let a struct grow new fields across kernel versions without a new syscall number:

  • An old binary built against a smaller struct passes a smaller usize. The kernel copies that many bytes and zero-fills the rest of the (newer, larger) kernel struct — any field the old binary never knew about reads as zero, which by convention means "not set."
  • A binary passes a larger usize than the kernel struct it knows about — for instance, a newer binary talking to an older kernel. copy_struct_from_user checks that every trailing byte beyond what the kernel understands is zero; if any of them is nonzero, it fails with -E2BIG rather than silently discarding data the caller thought it was setting.

openat2's struct open_how and clone3's struct clone_args are both built this way.

A user pointer that was not valid, turned into an error return instead of an oops.

References​