The Entry Path
Some of what happens between the SYSCALL instruction and do_syscall_64 is done by the CPU, because
software cannot be trusted to do it — there is no valid kernel state yet for software to run in. The
rest is done by software, because the CPU does surprisingly little on its own. Knowing which is which
is the difference between reading entry_64.S and guessing at it.
What the CPU does, exactly
On x86-64, executing SYSCALL makes the CPU do exactly five things, and nothing else:
- Load
RIPfromarch/x86/include/asm/msr-index.h:MSR_LSTAR— the address the kernel registered at boot as the syscall entry point. - Save the instruction to return to in
rcx(the oldRIP). - Save the flags register in
r11(the oldRFLAGS), then maskRFLAGSagainstarch/x86/include/asm/msr-index.h:MSR_SYSCALL_MASK— clearing the bits the kernel does not want carried into ring 0. See the register diagram below. - Load
CSandSSfromMSR_STAR, which is how the privilege level actually changes. - Jump to the loaded
RIP.
That is the entire hardware contract. Two things it conspicuously does not do: it does not switch
stacks, and it does not save any register other than the old RIP (in rcx) and the old RFLAGS (in
r11). Every general-purpose register the caller was using is still exactly where the caller left it
when software takes over. Everything else — finding a stack, saving those registers, deciding what to
run — is software's problem from this point on.
swapgs and finding the kernel's own state
The very first instruction in arch/x86/entry/entry_64.S:entry_SYSCALL_64() is
swapgs. Before the kernel can do anything else — before it can even find its own per-CPU data — it
needs a register it can trust, and GS is that register: swapgs exchanges the value in GS.base
with a value saved in MSR_KERNEL_GS_BASE, so that GS-relative addressing now reaches the kernel's
per-CPU structures instead of whatever userspace was using GS for. Nothing before this instruction
can be written as ordinary C, because ordinary C on this kernel assumes per-CPU data is reachable, and
until swapgs runs it is not.
This is also a genuine hazard, not a formality. If an interrupt or an NMI could land in the narrow
window where GS has already been swapped once but the kernel state built on top of it does not exist
yet — or is only half-built — the handler would run with the wrong idea of whose state it is looking
at. The entry code handles this explicitly, with PARANOID-flavoured paths for NMI and machine-check
entry that re-check whether a swapgs is already in effect before deciding whether to issue another
one, rather than assuming the normal one-swap-per-transition invariant holds.
The stack switch
SYSCALL did not touch rsp. The very next instructions after swapgs do: the entry stub stashes the
old rsp in a per-CPU scratch slot, then loads the new stack pointer from the per-CPU
cpu_current_top_of_stack — the top of this task's kernel stack, not a stack shared across tasks or
CPUs.
Running kernel code on a stack pointer userspace controls would be an immediate privilege escalation:
userspace could point rsp at a location it can read and write, and every subsequent push in the
entry path would be writing kernel-chosen values to an address userspace chose, which userspace could
then read directly — or worse, point rsp somewhere unmapped or dangerous and let the very first
push fault or corrupt something the kernel trusted. The switch has to happen in this exact position:
after swapgs gives the kernel a trustworthy per-CPU pointer to find the stack, and before anything
in the entry path uses the stack for a single byte.
Building pt_regs
With a trusted stack under it, the entry stub pushes the saved registers — ss, the old rsp,
rflags (from r11), cs, the old rip (from rcx), and the syscall number (from rax), followed
by the rest of the general-purpose registers — onto that stack, in the fixed layout of
arch/x86/include/asm/ptrace.h:pt_regs(). This is why every syscall handler, every
tracer, and every oops dump can find the caller's registers by shape rather than by convention: a
struct pt_regs * is not a hint about where the registers probably are, it is the actual layout the
entry stub built, byte for byte.
Then C takes over
Once pt_regs exists, arch/x86/entry/syscall_64.c:do_syscall_64() is called with
a pointer to it. It, in turn, calls
include/linux/entry-common.h:syscall_enter_from_user_mode() before dispatch and
include/linux/entry-common.h:syscall_exit_to_user_mode() after it — the pair that
handles the bookkeeping the raw entry stub does not: syscall tracing (ptrace, audit), seccomp
filtering, and, on the way out, checking for a pending signal or a rescheduling request. The exit path
is where a pending signal or a need_resched flag is actually acted on — the kernel does not interrupt
itself mid-syscall to deliver a signal or preempt a task; it waits for a safe, well-defined point, and
syscall_exit_to_user_mode is that point. This is the fact that folder 06's signals page and folder
07's preemption page both build on.
KPTI, in one paragraph
With kernel page-table isolation active, CR3 — the register that points at the current page-table
root — gets flipped on both sides of a syscall, because the user-mode and kernel-mode page tables are
kept separate to prevent user code from using speculative execution to read kernel memory it cannot
legitimately access. The two directions aren't symmetric. On entry, SWITCH_TO_KERNEL_CR3 (right after
swapgs, before the kernel stack switch above) flips CR3 inline using a scratch register — no separate
stack is needed yet, because the handful of instructions doing the flip are still running on the
old, still-mapped user stack. It's on the way out that a trampoline stack earns its name: just before
sysretq/iret, the CPU is about to switch back to the user page tables, so the kernel stack would no
longer be mapped once that happens. The exit path saves the real stack pointer, switches to a minimal,
always-mapped trampoline stack, flips CR3 back to user tables from there, and only then restores the
real rsp and executes sysretq.
Whether KPTI is active, and what else is active alongside it, depends on the CPU model and on boot
parameters (pti=on/off/auto, and others). /sys/devices/system/cpu/vulnerabilities/ is the
authoritative answer on any given machine — read it rather than assuming a mitigation is or is not
active based on kernel version alone.
This path is simplified
Deliberately. Error paths (a bad rcx/rip that would fault on SYSRET, for instance), CONFIG_*
variants, IST-based stacks for faults that can occur at inconvenient times, and the 32-bit compat entry
points are all elided here. arch/x86/entry/entry_64.S:entry_SYSCALL_64() is the
real thing, comments and all — the comments in it are documentation in their own right.
arm64 does this differently
arm64 has no direct equivalent of SYSCALL/MSR_LSTAR. A syscall is requested with SVC, which traps
to a vector table entry selected by exception category (synchronous exception from a lower exception
level, using AArch64) rather than by a single fixed vector number the way x86-64's SYSCALL always
targets MSR_LSTAR. The return address and saved processor state live in ELR_EL1 and SPSR_EL1
rather than rcx/r11, and the kernel stack pointer for EL1 comes from SP_EL1, which the exception
entry already made current — there is no software swapgs-style exchange to reach it. The syscall
number is passed in x8, not rax.
The division of labour on syscall entry: five things the hardware does, everything else in
entry_64.S.
syscall_init() programs MSR_SYSCALL_MASK to clear every flag shown above (plus RF and ID, off
this strip) on entry — the comment above the write says it plainly: "clear as much as possible to
minimize user space-kernel interference." Of the bits shown, IF and DF are the two that matter most
in practice: IF because the kernel needs interrupts to start out disabled on entry rather than
inheriting whatever state userspace happened to be in, and DF because a huge amount of kernel code
uses rep-prefixed string instructions and the C calling convention already assumes DF is clear —
inheriting a set DF from userspace would silently run those instructions backwards.
References
arch/x86/entry/entry_64.S:entry_SYSCALL_64()— the path itself, and the comments in it are documentation.- Intel SDM Vol. 2B, the
SYSCALL/SYSRETinstruction reference — the exact list of what the hardware saves, loads, and masks. - 7. Kernel Entries — the in-tree x86-64 entry documentation, current at v6.18.
- LWN, "Meltdown and Spectre: kernel page-table isolation" — why
CR3moves on entry and what the trampoline stack is for.