Skip to main content

Updated Sep 14, 2026

Tracing and Intercepting Syscalls

Watching a syscall and changing what it does are completely different mechanisms with completely different costs, and the tool people reach for first — strace — is frequently the wrong one for the job they actually have. strace is precise and expensive; the tracepoint path is cheap and cannot change anything; and the only mechanism built to intervene correctly is neither of those.

What actually happens when you run strace ls​

strace attaches with PTRACE_TRACEME (from a forked child, just before its execve) or PTRACE_ATTACH (to an already-running process), then asks the kernel to stop the tracee at every syscall boundary. From there, one syscall means two stops: the tracee stops on entry, before the syscall body runs, and stops again on exit, after it has returned. Each stop is a full context switch — tracee to tracer, tracer decides what to do (usually just "let it continue"), then tracer to tracee again — not a lightweight callback. strace's job at each stop is mundane by comparison: read the frozen pt_regs out of the tracee, decode the syscall number and arguments into the readable form you see on your terminal, and let the tracee resume.

Give the cost honestly, because strace's reputation as "the tool to reach for" undersells it: a syscall-heavy workload under strace commonly runs one to two orders of magnitude slower than untraced, because the double context switch is paid on every single syscall, and a program that makes a lot of small read/write calls pays it a lot of times. The consequence is not a footnote — it changes what you can conclude. strace does not observe a program's timing, it replaces it with the tracer's timing. It is a correctness tool — "what did this program actually call, with what arguments, in what order" — and a poor performance tool, because the act of measuring is itself the dominant cost.

perf trace, and why it is cheaper​

perf trace gets most of the same information from a different mechanism entirely: the raw_syscalls:sys_enter and raw_syscalls:sys_exit tracepoints, the same tracepoint pair kernel/entry/syscall-common.c's syscall_trace_enter() fires via trace_sys_enter() when SYSCALL_WORK_SYSCALL_TRACEPOINT is set (see the entry path for where that sits in the syscall-exit work). A tracepoint firing writes a record into a per-CPU ring buffer and returns — no stop, no scheduling decision, no context switch to a separate tracer process. The tracee never leaves the CPU it was already running on.

The trade-off is real, not just a footnote to the speed advantage:

  • Less detail per call. A tracepoint records the raw register values at the moment it fires. It does not, on its own, walk a struct sockaddr * or decode a flags bitmask into names the way strace's argument-printing tables do — perf trace layers some of that decoding back on top, but it is working from a snapshot, not a live, stoppable tracee.
  • No ability to modify anything. A tracepoint is a read: it cannot change the syscall number, the arguments, or the return value, because nothing is stopped for it to change. Interception is simply not this mechanism's job.

The number to keep in your head: ptrace-based tracing costs two context switches per syscall; tracepoint-based tracing costs one ring-buffer write, on the CPU that was already running.

The same question, four ways​

"Which files did this process open?" — one question, answered by four different mechanisms with four different costs.

$ strace -e trace=openat ./app
openat(AT_FDCWD, "/etc/resolv.conf", O_RDONLY|O_CLOEXEC) = 3
openat(AT_FDCWD, "/var/lib/app/config.json", O_RDONLY) = 4

Full argument decoding (path, flags spelled out by name, the returned fd), for free. Costs the two ptrace stops above on every syscall the process makes, not just the openat calls being filtered — the filter only decides what gets printed, not what gets stopped for.

Four tools, two mechanisms: strace pays for a stoppable tracee it does not need for a read-only question, and the other three all ride the same cheap tracepoint underneath a different amount of convenience.

Interception, properly​

None of the four tools above can change a syscall's outcome — they can only watch. The mechanism built to intervene correctly is seccomp user-space notification (SECCOMP_RET_USER_NOTIF): a seccomp filter, instead of allowing, denying, or killing on a matched syscall, suspends the calling task and hands a file descriptor for that suspended call to a separate supervisor process. The supervisor reads a struct seccomp_notif (include/uapi/linux/seccomp.h:seccomp_notif()) off that descriptor — the syscall number, the architecture, and the full argument snapshot — inspects it, and decides: let it proceed unmodified, fail it with a chosen errno, or (carefully — see the kernel header's own caution about the flag) let it continue.

What makes this correct where ptrace-based interception is not: the supervisor sees the arguments in a race-free way. A ptrace-based interceptor stops the tracee, reads its memory to resolve any pointer arguments (a path string, a struct), decides, and resumes — but between the read and the resume, a second thread in the same traced process can rewrite that memory, and the interceptor's decision was made against data that no longer matches what the syscall will actually see when it runs. This is a real, documented class of TOCTOU bug in ptrace-based sandboxes. Seccomp user notification does not by itself close every such window (the kernel header for SECCOMP_USER_NOTIF_FLAG_CONTINUE warns about exactly this if the flag is used to resume the original syscall unmodified), but the redesigned model — one process supervising, using an explicit fd-based protocol built for this purpose, rather than the general-purpose debugging interface ptrace also is — is what container runtimes such as runc and gVisor's runsc are built on for the syscalls they need to intercept rather than merely filter.

Why LD_PRELOAD is not syscall interception​

LD_PRELOAD replaces library functions — it works by injecting a shared object earlier in the dynamic linker's symbol resolution order, so a call to open() resolves to your replacement instead of glibc's. It never touches the syscall boundary itself. A statically linked binary, a Go program (whose runtime issues syscalls directly and does not go through libc at all — see libc is not the kernel), or code that calls syscall(2) directly all sail straight past an LD_PRELOAD shim, because there is no dynamic symbol resolution step for it to intercept in the first place.

What each mechanism can see and do​

MechanismObservesCan modifyCostSurvives a static binary
ptrace (strace)Full arguments, return value, every entry and exitYes — registers and memory of a stopped traceeHigh: two context switches per syscallYes
Tracepoints (perf trace)Arguments and return value, from a ring-buffer snapshotNoLow: one ring-buffer write, no stopYes
kprobesAnywhere a probe is attached, including inside syscall handlersNot safely, for the syscall's own outcomeLow to moderate, depending on probe placementYes
seccomp-notifyFull argument snapshot, race-free, via a supervisorYes — the syscall's outcome, via the supervisor's responseModerate: one suspend/resume round trip per intercepted callYes
LD_PRELOADOnly calls that go through the dynamic linker to libcYes, but only at the library layerEffectively free (a redirected function call)No — dynamically-linked, libc-routed calls only

Why the ptrace stop is where it is​

The ptrace check in syscall_trace_enter() is not bolted on beside the syscall dispatch — it runs inside the same syscall entry work that also runs seccomp and decides what happens with a pending signal, and in that order deliberately: kernel/entry/syscall-common.c runs ptrace's report first and seccomp second, specifically "to catch any tracer changes" a debugger made to the registers before the filter evaluates them. The entry path already established that syscall_enter_from_user_mode/syscall_exit_to_user_mode is where tracing, seccomp filtering, and signal delivery all live — one code path, three features, and the ordering between the first two is not an accident.

One syscall under strace: two stops, two context switches, and the reason tracing is expensive.

References​

  • ptrace(2), the syscall-stop section — the definitive statement of when PTRACE_SYSCALL stops happen and what the tracer sees at each.
  • seccomp_unotify(2) — the modern interception interface, with a complete worked example in the man page itself.
  • perf-trace(1) — the low-overhead alternative and its option surface.
  • LWN, "Deferring seccomp decisions to user space" — why ptrace-based interception was inadequate and what replaced it.