Skip to main content

Real-Time Operating Systems

When a kernel earns its RAM, how tasks are scheduled and switched, and the synchronisation that comes back with pre-emption.

📄️Context Switching

A context switch is not a function call. A function call returns to its caller; a context switch enters on one task's stack and leaves on another's, and the code that resumes has no idea it was ever stopped. There is no C construct for that. What makes it possible on a Cortex-M is that the processor already performs most of it, for free, every time an exception is taken — and that an exception return reads its destination out of memory rather than from a register, so pointing it at a different stack points it at a different task.

📄️Semaphores and Mutexes

Adding a scheduler adds a second way for code to be interrupted. On bare metal the only thing that could run between your load and your store was an exception handler, and Critical Sections and Atomicity is the complete answer to that. Under a kernel, a task can also be pre-empted by another task, at any instruction, for reasons that have nothing to do with the NVIC. Masking interrupts no longer covers the hazard, because the thing that pre-empted you was the scheduler doing its job.

📄️Notifications and Event Groups

Queues, semaphores, mutexes and event groups are all objects. Each has its own allocation, its own handle, and its own place in the system's RAM budget, and a sender has to be given that handle before it can say anything. The kernel's own header states the distinction plainly: those four are "intermediary objects" for sending an event, and a task notification "is a method of sending an event directly to a task without the need for such an intermediary object."

📄️Software Timers and Delays

There are three clocks in a system running an RTOS and it is worth separating them before writing a single delay. There is the hardware timer, counting a crystal-derived clock in silicon, which does not know or care what software is doing. There is the tick, a periodic interrupt that increments a counter and is the only time the kernel can see. And there is the task's own experience of time, which is the tick as filtered through whether that task was actually running when something expired.

📄️Priority Inversion and Deadlock

A priority is a promise about who gets the core when two tasks want it. A lock is a mechanism that can break that promise, and the reason is structural rather than accidental: a task holding a lock has implicitly borrowed the urgency of every task that will ever wait for it, and the scheduler has no way of knowing that until someone actually waits. Until that moment the holder is running at its own priority, and the system is scheduling a task that is, in effect, on the critical path of something far more urgent.

📄️ISR-Safe APIs

Every kernel object in the previous six pages — the ready lists, the delayed list, each queue's two event lists, each mutex's holder field — is an ordinary linked list in RAM, and the kernel keeps it consistent by the only means a single-core Cortex-M offers: it masks interrupts around the update. That is the entire content of taskENTER_CRITICAL(). So a kernel API is not "thread-safe code" in the hosted sense; it is code that assumes nothing else on this core is running while it edits the list.

📄️Debugging and Tracing

A breakpoint stops time. On bare metal that is a fair trade, because the interesting question is usually "what is in this variable", and a halted core answers it perfectly. Under an RTOS the interesting questions have changed shape: why did the control loop run 4 ms late, which task was holding that mutex, what ran between the interrupt and the response. Every one of those is a statement about a sequence in time, and the instant you halt, the sequence you wanted to observe is gone.