Skip to main content

108 docs tagged with "embedded"

View all tags

A GPIO Driver from Scratch

The blink program drove one pin with two #defines and it was the right amount of code for one pin. The second pin costs another two, the first alternate-function pin costs four, and by the tenth you have shift arithmetic scattered across the codebase with the pin number written out by hand in each place. The fix is not a HAL. It is about eighty lines that name what the hardware already does.

ADC and DAC Drivers

A successive-approximation ADC is, physically, a capacitor with a switch in front of it. Converting a voltage happens in two completely different phases: first the switch closes and the capacitor is allowed to charge towards your signal through the source impedance, and then the switch opens and a comparator plays twelve rounds of twenty questions against the trapped charge. The second phase is fixed by the hardware and takes exactly twelve clocks. The first phase is the one you configure, the one every tutorial leaves at its reset value, and the one that decides whether your reading means anything at all.

Analog Basics: ADC and DAC

An analog-to-digital converter looks, from firmware, like a register you read. That framing hides the two things that actually determine whether the number is right. First, the conversion is a comparison against a reference, so the answer is a ratio, not a voltage — and a wrong or noisy reference is invisible in the result. Second, before any comparing happens the converter must charge a small capacitor through your circuit, and if you did not give it long enough, it will confidently report the voltage it managed to reach rather than the voltage that was there.

Bare-Metal, RTOS, or Linux

An engineer coming from application software tends to reach for the environment closest to what they already know — an RTOS, or better yet Linux, because it has threads and a filesystem and feels familiar. That instinct is worth resisting. Every layer of software you add between your code and the hardware costs something real: flash and RAM you don't get back, boot time, a scheduler whose behavior you now have to understand rather than one you wrote yourself, and — for Linux specifically — an MMU-capable microprocessor in the bill of materials at all. The right choice is the cheapest one that actually meets the product's requirements, and picking a heavier environment "to be safe" is itself a common and expensive mistake. This page exists to make that trade-off concrete instead of a matter of taste.

Brownout and Power-Loss Safety

A supply does not fail like a light switch. Pull the plug, drop a battery connector, or let a coin cell finally reach the end of its discharge curve, and VDD does not fall to zero — it decays, over a span set by whatever capacitance is still holding charge on the rail and whatever the circuit is still drawing from it. That decay is the entire subject of this page, because it is the only thing standing between "the supply is gone" and "the core has actually stopped running," and the honest answer to "how long do I have" is: a lot less than most designs assume, and the amount is calculable rather than a matter of taste.

Build Systems and Vendor Tooling

Choosing a build system for firmware feels like a taste question and is not. The real question underneath it is who owns the generated code, and it has consequences that outlive the project: whether a colleague can build your firmware without installing an IDE, whether CI can build it at all, and whether the day you need to change a pin assignment costs ten minutes or a merge conflict across forty files you did not write.

C Libraries: newlib, newlib-nano, picolibc

The C standard library is written as though a process exists. printf writes to a file descriptor. malloc asks the kernel for more address space. exit tells a parent that a child finished. fopen needs a filesystem, time needs a clock somebody set, errno needs somewhere thread-local to live. On a Cortex-M4 with no OS, not one of those things is true — and yet #include compiles fine and printf links, because the library was built with the bottom of every one of those paths left as a hole for you to fill.

Choosing a Toolchain

For a hobby project the toolchain question is nearly free: install Arm GNU, move on. For a product it is one of the longest-lived decisions in the codebase. A toolchain choice outlives the engineers who made it, because the compiler is baked into every build artefact you have ever released and into every certification argument you have ever made. Changing it later is not a flag change; it is a re-qualification.

Clock and Peripheral Gating

Two habits do most of the work in a low-power design that never touches Sleep, Stop, or Standby at all the energy cost of a fixed piece of work is not simply proportional to clock speed.

Clocks and Oscillators

On a desktop machine the clock is somebody else's problem — it was configured by firmware you never see, and by the time your program runs it is a constant. On a microcontroller you are that firmware. The chip comes out of reset running on a cheap internal RC oscillator at a fraction of its rated speed, with almost every peripheral's clock switched off, and the first job your code has is to build the clock tree the rest of the system will run on. Nothing you write behaves as intended until that is done.

CMake for Embedded

CMake's defaults encode one assumption so deeply that it is easy to miss: the machine running the build can also run what the build produces. That is what lets CMake test a compiler by compiling and linking a tiny program, what lets findpackage look in /usr/lib, and what lets checkcsourceruns exist at all. Every one of those assumptions is false for a Cortex-M4 with 128 KB of RAM and no operating system.

CMSIS and Vendor HALs

"Bare metal" does not have to mean "type every address yourself". Between raw pointer casts and a full vendor framework there are three or four distinct layers, each with a different bargain, and the useful skill is knowing which one you are standing on and why — not picking a side.

Configuring the Clock Tree

Out of reset the STM32F411RE runs at 16 MHz on an internal RC oscillator, which is a deliberately conservative choice: it works with no crystal, no configuration, and no risk. It is also one sixth of what the part can do, it is accurate to about ±1 % over temperature rather than the ±20 ppm a crystal gives, and it cannot produce the 48 MHz that USB requires. Somewhere in the first week of a real project you will need to change it.

Context Switching

A context switch is not a function call. A function call returns to its caller; a context switch enters on one task's stack and leaves on another's, and the code that resumes has no idea it was ever stopped. There is no C construct for that. What makes it possible on a Cortex-M is that the processor already performs most of it, for free, every time an exception is taken — and that an exception return reads its destination out of memory rather than from a register, so pointing it at a different stack points it at a different task.

Critical Sections and Atomicity

A bare-metal program with interrupts enabled is a concurrent program. There is one core and no scheduler, but there are still two threads of control — main and whatever handler just fired — and they share memory. Everything that makes concurrency hard is already present:atomic you can assume is lock-free, no kernel to block against.

Cross-Compilation

The compiler on your laptop is not a general-purpose translator that happens to be pointed at x86. It is a program that was built to emit x86-64 instructions, linked against a C library that was built assuming a Linux kernel is underneath it, and wired to a startup object that assumes something already created a process, set up a stack, and handed it argc and argv. Every one of those assumptions is false on a microcontroller. That is the whole reason a separate toolchain exists — not because the MCU is "different hardware", but because three independent layers of the build all encode a machine and an environment, and all three have to change together.

Cross-Compilation

Cross-compilation builds executables for a different platform (target) than the one running the compiler (host). Essential for embedded systems, mobile development, and deploying to different architectures.

Debugging and Tracing an RTOS

A breakpoint stops time. On bare metal that is a fair trade, because the interesting question is usually "what is in this variable", and a halted core answers it perfectly. Under an RTOS the interesting questions have changed shape: why did the control loop run 4 ms late, which task was holding that mutex, what ran between the interrupt and the response. Every one of those is a statement about a sequence in time, and the instant you halt, the sequence you wanted to observe is gone.

Deferred Work

"Keep your ISRs short" is the most repeated piece of firmware advice and the least actionable, because it never says short compared to what. The useful form is a consequence of how the NVIC works rather than a style preference: every microsecond a handler runs is a microsecond added to the worst-case latency of every interrupt at the same or a less urgent priority. A 200 µs handler is not "a bit slow". It is a 200 µs latency tax levied on most of the system, paid every time that handler runs, and it does not appear in any register you can read.

Determinism Killers

Almost every mechanism that makes a computer fast makes it less predictable, and the trade is usually invisible in source code. Caches, prefetchers, branch speculation, bus arbitration and dynamic memory allocation are all bets on history repeating: they are fast when the recent past resembles the present and slow when it does not. In a throughput-oriented system that is an unambiguously good bargain. In a system with deadlines it means the number you have to defend — the worst case — is set by the unlucky path through every one of them, while every measurement you take is dominated by the lucky one.

DMA

A DMA controller is not an accelerator bolted onto a peripheral. It is a second bus master: a small, dumb machine that sits on the same bus matrix as the Cortex-M4 and, when a peripheral raises a request line, performs the load and the store that your interrupt handler would otherwise have performed. It has no idea what the data means. It knows a source address, a destination address, a count, and whether to increment each pointer.

Embedded C Idioms

Embedded C is the same language as any other C. What differs is which of its underspecified corners you are standing on. On a desktop, int is 32 bits, structs are laid out the way you expect, unaligned access works, and the byte order matches whatever produced the file. In firmware you are parsing a protocol written by someone else's compiler, laying a struct over a hardware register, and running on a part where int might be 16 bits and an unaligned load might be a fault.

Embedded Systems

Embedded software runs on hardware that was never meant to run much of anything: a few hundred

Energy Budgets

Every battery-powered product answers one question before any other design decision matters a firmware team that only measures active current and ignores what the part draws asleep will overestimate lifetime by orders of magnitude, and a team that only reads the nameplate mAh off a battery datasheet without reading the discharge curves will do the same in the opposite direction.

Exceptions and the Vector Table

On most processors, getting from "an interrupt line went high" to "my C function is running" involves software: a dispatcher reads a status register, works out which source fired, and calls the right handler. Cortex-M does none of that. The hardware reads a table of function pointers at a known address, indexes it by exception number, and branches — having already pushed the registers a C function is allowed to clobber. Your handler is an ordinary function with an ordinary prologue, and it starts running a fixed and small number of cycles after the event.

External Memory and QSPI

The moment your data stops fitting on-chip, the interesting question is not which memory — it is whether the processor has to execute from it, or merely read it. Those two requirements lead to completely different hardware. Data you read into a buffer can live behind four wires and a software driver, and a plain SPI port is enough. Code the CPU fetches instructions from must appear in the address map, which means a controller that turns a bus read into a flash transaction with no software involved at all.

Flashing and Programming

On a hosted system, "running the program" means handing a file to a loader. Here it means writing your image into non-volatile memory inside the chip and then resetting it, using a second piece of hardware that talks to the silicon over a two-wire debug port. That second piece of hardware is doing considerably more than copying bytes: it halts the core, drives the flash controller through a sequence the reference manual specifies, verifies, and releases reset.

Floating Point and DSP Extensions

Writing float in C on a microcontroller does not tell you what the hardware will do. The same line of source can compile to one instruction, to a forty-cycle library call, or to a two-hundred-cycle double-precision emulation — and the compiler chooses silently, based on flags you may not have set deliberately. Nothing in the source distinguishes the three cases, which is why "why is my control loop suddenly missing deadlines" is so often a floating-point question.

Glossary

Embedded engineering has its own vocabulary, and a lot of it is acronyms that mean something quite specific in this field even when the letters look familiar from elsewhere. The problem isn't that the terms are hard — it's that skimming past one you half-recognize (assuming "MPU" means the same thing every time, or that "RTOS" is just "a small OS") is exactly how a plausible-but-wrong mental model gets built, and those are expensive to unlearn once you've written code around them. This page defines the terms every later folder in this section assumes you already know, once, in one place, so you can look one up instead of re-deriving it from context. Where a term's proper home is a folder that doesn't exist in this build yet, it's still defined here — it just isn't linked anywhere yet.

HardFault Forensics

A HardFault is not an error message. It is the processor announcing that it has already stopped being able to run your program, and that everything it knows about why is sitting in four registers and one stack frame. Nothing is printed, nothing is logged, and the default handler in every startup file ever shipped is b . — an infinite loop that discards all of it.

How a GPIO Pin Really Behaves

GPIOA->ODR |= (1 << 5); looks exactly like every other memory write you have ever done, and that resemblance is the problem. On the far side of that register is not a bit of storage but a pair of transistors, wired to a physical pin, with a maximum current, a maximum switching rate, and a real-world net on the other end that may already be being driven by something else. The register model hides all of it, right up until the moment it matters.

I2C in Depth

I²C is the only one of the three common serial buses whose electrical design is part of its protocol. SPI and UART drive their lines push-pull multiple controllers on the same two wires, targets that can pause the controller, collision detection that costs no extra hardware, and the ability to hang three sensors off two pins.

Input Capture and Encoders

PWM points the timer outward: the counter drives a pin. Input capture points it inward. The counter free-runs, an edge on a pin tells the hardware "now", and the value of CNT at that instant is copied into a capture register before software has had a chance to be late. That last clause is the entire value of the peripheral. A GPIO interrupt can also tell you an edge happened, but by the time your handler reads a counter it has been anywhere from 12 to several hundred cycles — jittering with whatever else the NVIC was doing — and the measurement carries that jitter. The capture unit's latch has no jitter at all.

Instruction and Event Tracing

There is a category of bug that a breakpoint destroys by looking at it, and a log line destroys almost as thoroughly. A race between two interrupts, an occasional priority inversion, a state machine that takes a wrong branch once in ten thousand iterations — halting the core to inspect any of these changes exactly the timing that produced them, and a printf line inserted to catch them costs enough cycles to close the window it was supposed to observe. The Debug Toolbox's perturbation table already ranks every instrument in this folder by how much it disturbs the thing it is measuring; trace is the answer at the bottom of that table, the one built specifically to be close to zero.

Internal Flash and EEPROM Emulation

Flash is not memory that happens to be non-volatile. It is a device with an asymmetric write model, and every design decision about storing settings on an MCU comes out of that asymmetry: you can clear a bit at any time, cheaply, one word at a time — but you cannot set a bit back to 1 without erasing an entire sector, which takes up to two seconds, during which the processor cannot fetch instructions from the same flash it is executing from.

Interrupt Latency

Ask what the interrupt latency of a Cortex-M4 is and you get "12 cycles", which is true and almost never the number you need. Twelve cycles is what the processor contributes when the memory is instant, nothing else is running, and the exception arrives at a convenient moment. The number that decides whether your product works is the one measured from the event at the pin to the first instruction of your handler that does something about it — and on a real board, at a real clock, with the rest of your firmware present, that number is dominated by things the core designer had no say in.

ISR-Safe APIs

Every kernel object in the previous six pages — the ready lists, the delayed list, each queue's two event lists, each mutex's holder field — is an ordinary linked list in RAM, and the kernel keeps it consistent by the only means a single-core Cortex-M offers: it masks interrupts around the update. That is the entire content of taskENTER_CRITICAL(). So a kernel API is not "thread-safe code" in the hosted sense; it is code that assumes nothing else on this core is running while it edits the list.

Lab Equipment and What It Answers

The debugger on your Nucleo can single-step your code, read every register, and show you the contents of memory — and it is blind to everything that happens outside the package. It will tell you, truthfully, that you wrote 0xA5 to the SPI data register. It cannot tell you whether 0xA5 left the pin, whether the clock that carried it was clean, whether the device on the other end was even powered.

Logging Without Breaking Timing

Here is the observation this page exists to explain. You have a fault that happens once every few minutes — a corrupted buffer, a missed deadline, a state machine that ends up somewhere it cannot reach. You add a printf to narrow it down. The fault stops happening. You remove the printf and it comes back.

Logic Analyzer Workflows

"The bus doesn't work" is not a bug report a logic analyzer can answer by itself, and the gap between clipping on four wires and actually learning something is where most of a session goes. Lab Equipment and What It Answers makes the case for reaching for the analyzer first and works one I²C bug through it end to end; this page assumes you have already made that choice and is about running the capture well enough that the answer it gives you is true.

Measuring Power

Every figure in this folder so far has come from a datasheet table or from arithmetic built on one. Neither tells you what your board actually draws. A datasheet's Stop-mode current is measured on ST's own characterisation board, with nothing attached to the GPIOs, a specific silicon revision, and every peripheral in the state ST chose to test it in; your board has a pull-up ST did not model, a status LED that never got gated, and firmware that may or may not be entering the mode it claims to. Energy Budgets's entire arithmetic is only as good as the I_avg fed into it, and the only way to know that number is real is to measure it on the actual hardware running the actual release firmware — the warning on that page about a demo build with logging left on exists precisely because nobody measured until it was too late.

Memory Sections and VMA vs LMA

Every variable and every function in your firmware ends up in one of about six buckets, and which bucket it lands in is decided by two things: whether it is code or data, and whether its initial value is zero. That is nearly the whole rule. int counter; goes in .bss because its initial value is zero. int counter = 5; goes in .data because it is not. const int limit = 5; goes in .rodata because it never changes and can therefore stay in flash. Nobody chose those placements for your variable; the compiler applied that rule and emitted a section name.

Microcontroller, Microprocessor, SoC

"It's an ARM chip" tells you almost nothing useful. The question that actually matters — can this thing run Linux, does it need a bootloader partition scheme, will your firmware fit without an external memory chip, is there hardware memory protection between tasks — all comes down to one boundary: where the code and data live relative to the CPU core, and whether there's a hardware unit that translates and protects memory addresses on the way there. That boundary is what separates a microcontroller from a microprocessor, and it's a hardware property you can check on a datasheet, not a marketing category.

Mocking Hardware

Unit Testing Firmware drew a clean line around the logic that needs nothing from the hardware at all — a parser, a state machine, a checksum — and pointed at this page for the harder case: logic whose whole job is to talk to a peripheral. A driver cannot be tested by giving it inputs and checking outputs the way a parser can, because its inputs and outputs are register writes and register reads, and there is no register to read from on a laptop. Testing it on the host requires something on the other end of the bus that behaves like the real device without being it.

Optimization for Size and Speed

Raising the optimization level is the one build change that routinely alters what a firmware does. Not what it does more quickly — what it does. A delay loop disappears. A register write that was there at -O0 is gone at -O2. Code that worked for two years starts failing, and nothing in the source changed.

Polling, Interrupt, or DMA

There are exactly three ways to get a byte out of a peripheral register and into your program's memory. The CPU can ask repeatedly until the answer is yes. The peripheral can raise a line that makes the CPU stop what it was doing. Or a second bus master can do the load and the store on the CPU's behalf and tell it afterwards. Every driver you will ever write picks one of these, and the choice is made badly far more often than it is made wrong — badly, meaning by reflex rather than from a budget.

Postmortem Debugging

Every technique in HardFault Forensics assumes something that is only true on your bench: a debugger attached at the moment of the fault, watching CFSR before anything clears it, holding the stack frame before the stack is reused. A device in the field has none of that. It faults, and unless you decided in advance what to do about it, the only evidence is a b . loop nobody is watching, or a reset that erases everything and starts the firmware running again as if nothing happened.

Power Supplies and Regulators

Firmware is written as though the supply rail were a constant — a number in the datasheet, 3.3 V, always there. The rail is not a constant. It is the output of a control loop with finite bandwidth, fed through traces with real resistance and inductance, feeding a load whose current draw your own code is modulating thousands of times a second. Every time the CPU switches from an idle loop to a burst of floating-point work, every time a GPIO drives an LED, every time the chip wakes from Stop mode, the load steps and the rail moves.

Priorities and Nesting

Every interrupt on a Cortex-M comes out of reset at priority 0 — the most urgent level there is. A program that enables six interrupts and never calls NVIC_SetPriority has six handlers that all sit at the top, none of which can pre-empt any other, serviced in exception-number order when several arrive at once. That configuration is not "no priority scheme". It is a specific, and usually wrong, priority scheme: it says every interrupt in the system is equally urgent and none may interrupt another, which means the worst-case latency of your fastest deadline is the sum of every other handler's execution time.

Priority Inversion and Deadlock

A priority is a promise about who gets the core when two tasks want it. A lock is a mechanism that can break that promise, and the reason is structural rather than accidental: a task holding a lock has implicitly borrowed the urgency of every task that will ever wait for it, and the scheduler has no way of knowing that until someone actually waits. Until that moment the holder is running at its own priority, and the system is scheduling a task that is, in effect, on the critical path of something far more urgent.

Privilege Modes and the Two Stacks

A Cortex-M has two independent switches that most bare-metal firmware never touches, and that an RTOS depends on completely. One decides which stack the processor is using; the other decides whether the code running is allowed to change anything important. They are separate — you can be unprivileged on the main stack or privileged on the process stack — and confusing them is the source of a lot of half-right explanations.

PWM

A microcontroller pin has two output voltages and nothing in between. Pulse-width modulation is the trick that gets the third: switch fast enough and whatever is downstream — an LED and your eye, a motor and its inductance, an RC filter and a slow ADC — averages the square wave into a level. The pin is still only ever fully on or fully off, which is why it dissipates almost no power doing it. That is the whole reason PWM won over analogue drive for everything from a status LED to a 10 kW inverter.

Queues and Message Passing

A queue looks like a data structure and is really a design decision. The structure part is unremarkable — a circular buffer with a head, a tail and two waiting lists. The decision is what makes it worth a page: a FreeRTOS queue copies the item in and copies it out again, and every property people value in queues follows from that one fact.

Reading a Datasheet

Coming from software, the instinct when you meet a new chip is to look for "the docs" — one document, searchable, that tells you everything. That document does not exist, and looking for it is the reason people bounce off hardware. Silicon vendors ship a set of documents, deliberately separated, because they answer questions that different people ask at different times: the person choosing a part, the person laying out the board, the person writing the firmware, and the person whose product works on the bench but fails one unit in fifty. Each document is written for one of those people and is close to useless for the others.

Reading a Schematic

A schematic is not a picture of a board. It is a graph: components are nodes, and the wires between them — nets — are edges. Two points drawn at opposite corners of the page with the same net label are the same electrical point, as surely as two references to the same object in memory. Once you read it as a graph rather than as a drawing, the intimidating density stops mattering, because you are never reading the whole thing. You are tracing one path.

Reading the Map File

Every firmware project reaches the same afternoon. The build that fitted last week does not fit this week, or it fits but the RAM figure has doubled, and nobody changed anything that should have cost 12 KB. The instinct is to start deleting features. The correct move is to ask the toolchain, which has known the answer the whole time and wrote it down.

Register-Level Programming

A peripheral is a piece of digital logic sitting on the same bus as your RAM. It has no API, no calling convention and no way to be invoked. The only interface it exposes is a small block of addresses: write a word to one of them and some flip-flops change state; read from another and you get the current state of some wires. That is the whole model.

Reset and Boot Configuration

There is a gap between the moment power reaches the chip and the moment your first instruction executes, and firmware engineers habitually treat it as empty. It is not. In that gap the supply supervisor decides whether the rail is trustworthy, a pulse generator stretches whatever event caused the reset into a signal long enough for every block on the die to see it, an option-byte loader runs, boot-mode pins are sampled and latched, an address decoder is reconfigured so that a completely different memory appears at address zero, and only then does the CPU fetch two words and start running.

RISC-V for Arm Developers

Everything in this folder so far has described one architecture. The reason that is a reasonable way to spend eleven pages is that Cortex-M is what the overwhelming majority of microcontroller work is written against. The reason it is not the only thing worth knowing is that RISC-V parts are now genuinely shipping in volume — the ESP32-C3 in a hobbyist's hands, the CH32V003 at ten cents, the RISC-V management cores inside SoCs whose application processors are Arm — and the transition is much easier than a new instruction set sounds, provided you know which of your Cortex-M assumptions are architectural and which are Arm's.

RTC and Timekeeping

Every other peripheral in this folder lives in your power domain, stops when you stop, and forgets everything when the supply goes away. The RTC does not. It is a small independent machine in a separate power domain with its own oscillator, its own supply pin, and its own reset — and the only things it shares with the rest of the chip are a bus interface and a couple of locks that exist specifically to stop your code from disturbing it by accident.

Scheduling Theory for Firmware

Most firmware priority assignments are opinions. Somebody decided the safety task was the most important thing in the system and gave it the highest priority, somebody else discovered the display flickered and bumped its task up, and three years later nobody can say whether the 1 kHz control loop meets its deadline — only that it seems to. Real-time scheduling theory replaces that with arithmetic. Given each task's execution time, its period and its deadline, it answers "does every task always meet its deadline" with a proof rather than a measurement campaign.

Semaphores and Mutexes

Adding a scheduler adds a second way for code to be interrupted. On bare metal the only thing that could run between your load and your store was an exception handler, and Critical Sections and Atomicity is the complete answer to that. Under a kernel, a task can also be pre-empted by another task, at any instruction, for reasons that have nothing to do with the NVIC. Masking interrupts no longer covers the hazard, because the thing that pre-empted you was the scheduler doing its job.

Shared Data and Race Conditions

A single-core MCU with interrupts enabled runs a concurrent program, and the reason people are surprised by that is that the concurrency is invisible in the source. There is one main, one call stack in view, no threads to create, and nothing that looks like it could run at the same time as anything else. But between any two machine instructions the hardware may insert an entire function of someone else's code, and if that function touches a variable you were half-way through updating, the update is lost — or worse, the value it reads is one that never existed.

Signal Integrity and Noise

A schematic draws a wire as a line with no properties. That abstraction holds beautifully for DC and falls apart on the edges — the few nanoseconds after a driver switches, when the wire is not a connection but a component, with inductance, capacitance, a characteristic impedance and a finite speed. For most of the time your signal is idle and the abstraction is fine. For the small fraction of time when it is changing, the wire is the circuit.

Simulation and Emulation

Every technique in the rest of this folder needs a board on the desk, or at minimum a physical binary running somewhere real. Unit Testing Firmware and Mocking Hardware sidestep that by pulling logic out from behind a seam and running it on the host — a deliberate, narrow substitute for one driver's register interface. Simulation is a different move entirely: instead of extracting the logic, it models the chip, and runs the actual, unmodified cross-compiled firmware.elf against that model — the same instruction stream, the same startup code, the same register pokes, executing against a virtual STM32 instead of a real one.

Sleep Modes

A low-power mode is not one switch, it is a set of trade-offs arranged along a slider first the CPU clock, then every clock and the main regulator, then the regulator's own supply to everything but a small backup island.

Software Timers and Delays

There are three clocks in a system running an RTOS and it is worth separating them before writing a single delay. There is the hardware timer, counting a crystal-derived clock in silicon, which does not know or care what software is doing. There is the tick, a periodic interrupt that increments a counter and is the only time the kernel can see. And there is the task's own experience of time, which is the tick as filtered through whether that task was actually running when something expired.

SPI in Depth

SPI is a shift register with a wire between two halves of it. That is the whole protocol, and holding it in mind explains everything the peripheral does. The controller has eight bits, the target has eight bits, and the clock the controller generates walks them past each other in a ring: the controller's MSB goes out on MOSI and into the target's LSB position, the target's MSB goes out on MISO and into the controller's. After eight clocks the two registers have swapped contents. There is no addressing, no acknowledgement, no error detection and no notion of a transaction — every one of those has to be built on top by whatever protocol the target's datasheet defines.

Stack Usage and Overflow

Stack overflow on a hosted operating system is an event. A guard page is hit, the process receives a signal, the debugger stops on the offending frame, and you have a stack trace pointing at the recursion you forgot to bound. Stack overflow on a bare-metal Cortex-M is not an event. It is a silent write to a variable that belongs to something else, discovered later, in code that is innocent.

Stacks and Heaps in an RTOS

On bare metal there is one stack, the linker script decides where it starts, and the map file tells you how much room it has. Adopting a kernel deletes all three of those facts at once. There are now N+1 stacks; the linker knows about exactly one of them; and the other N are ordinary arrays — either carved out of a heap at runtime or declared as static — with nothing above or below them that the toolchain considers special.

Startup Code: Reset to main

C has a set of guarantees that programmers stop noticing after the first month. Initialised globals hold their initialisers. Uninitialised globals are zero. The stack works. printf has somewhere to write. Static C++ objects have had their constructors run before main starts. On a hosted system a program loader and a crt0 object you have never opened deliver all of that before your first line executes.

Static Analysis and Sanitizers

There is a tool that finds real defects, requires no new install, no configuration file, and no CI integration beyond a flag already accepted by the compiler you are already invoking on every build — and most projects run it in a mode that throws its findings away. That tool is the compiler itself. arm-none-eabi-gcc already builds a fully type-checked internal representation of your program to generate code from; a warning is analysis it already did, offered back to you, and the default build configuration in a great many embedded projects is -Wall at best, or nothing at all, which means the cheapest static analysis available is routinely left switched off.

Static Memory and Why malloc Is Banned

Most embedded coding standards ban dynamic allocation after startup, and most engineers meet the rule before they meet the reason. Stated as a rule it sounds like superstition — malloc works, it is in the standard library, the vendor examples call it. Stated as a consequence it is obvious: a device that runs for three years without restarting cannot use a memory strategy whose correctness depends on restarting.

SWD, JTAG, and GDB

The thing that surprises people about on-chip debugging is how little of it is software. GDB does not single-step your firmware; the processor single-steps it, because Armv7-M defines a halting-debug mode with its own registers, comparator hardware for breakpoints, and comparator hardware for watchpoints. GDB is a client. OpenOCD is a translator. The debug capability is silicon that was on the die before you wrote anything, and it is finite in a way that software breakpoints on a hosted system are not.

SysTick and the Core Peripherals

Almost everything on a microcontroller is the vendor's. The timers, the UARTs, the clock tree, the GPIO blocks — all of it is ST's design, at ST's addresses, described in ST's reference manual, and none of it transfers to a part from a different vendor. A handful of peripherals are not like that. They are Arm's, they are built into the processor itself, they sit at the same addresses on every Cortex-M ever made, and code that uses them ports between vendors without changes.

Task Notifications and Event Groups

Queues, semaphores, mutexes and event groups are all objects. Each has its own allocation, its own handle, and its own place in the system's RAM budget, and a sender has to be given that handle before it can say anything. The kernel's own header states the distinction plainly: those four are "intermediary objects" for sending an event, and a task notification "is a method of sending an event directly to a task without the need for such an intermediary object."

Tasks and Scheduling

A task looks like an infinite loop written as if it owned the processor. That illusion is the whole product, and it is manufactured from three things: a block of RAM used as a stack, a small structure recording where that stack's pointer got to, and membership of exactly one linked list. Everything the scheduler does is move that structure between lists.

The Anatomy of a Peripheral

An STM32F411RE has a UART, three SPIs, three I²C controllers, eight timers, an ADC, a USB controller, an RTC, two watchdogs and two DMA engines. The reference manual gives each of them thirty to eighty pages, and read front to back they look like thirteen unrelated pieces of hardware. They are not. They are thirteen instances of one design, drawn by the same team, wired onto the same two buses, and configured by the same six steps in the same order every time.

The Cortex-M Family

Arm does not sell chips. It sells processor designs, and a silicon vendor — ST, NXP, Nordic, Raspberry Pi — licenses one, wraps it in memory and peripherals, and sells you the result. That arrangement is why "it's an Arm chip" tells you almost nothing on its own, and why the useful question is always which Arm core, in which configuration, from which vendor.

The Cortex-M Memory Map

A microcontroller has one 4 GB address space and everything lives in it: flash, RAM, every peripheral register, the interrupt controller, the debug hardware. There is no MMU, so what you write in a pointer is the physical address the bus sees. That is the simplification that makes bare-metal firmware tractable — and it means the layout of that space is not a vendor's private business but part of the architecture you program against.

The Debug Toolbox

Embedded debugging goes wrong in a particular way. You have a symptom — the board hangs after eleven minutes, the sensor returns zeros every hundredth read, the current draw is four times what the datasheet promises — and you reach for the instrument you are most comfortable with rather than the one that can answer the question. Then you measure the wrong thing very carefully for a day.

The Embedded Landscape

Picking a chip for a project isn't like picking a library — you can't easily swap it out six months in. The instruction set, the vendor's toolchain, the peripheral register layout, and the ecosystem of drivers and examples around a part are all things you commit to for the life of the product, and products in this field often live for years. So "which family of hardware" is really a question about which toolchain, which debugging workflow, and which vendor's support model you're signing up for — the raw specs are almost a secondary concern. This page maps the major families so that when a later folder says "on Cortex-M" or "targeting RISC-V," you have a sense of where that sits in the wider landscape and what it commits you to.

The Linker Script

On a hosted system nobody writes a linker script, because the answer to "where does this code go" is "wherever the loader decides", and the loader is part of the OS. On a microcontroller there is no loader. The addresses in the binary are the addresses the CPU will use, forever, and something has to choose them. That something is a text file, usually about a hundred lines long, that most projects copy once and never read.

The Memory Protection Unit

A null-pointer write on a Cortex-M does not crash. Address 0x00000000 is real memory — it is the start of flash, or the boot alias — so *(uint32t )0 = 42 on a fresh chip silently does nothing at all, and the program carries on. A stack that overflows its intended region does not crash either; it grows down into .bss and corrupts variables that belong to something else, and the failure surfaces minutes later in code that is entirely innocent. Both are the same problem: on a bare Cortex-M, memory has no permissions, so a wrong access is indistinguishable from a right one.*

The NVIC

The vector table answers where. The Nested Vectored Interrupt Controller answers whether, when, and in what order — and it is the only part of the exception model you actively program. Every interrupt line on the chip arrives at the NVIC; the NVIC decides whether that line is enabled, records it as pending if it is not yet serviceable, compares its priority against whatever the processor is currently doing, and hands the winner to the core.

The Oscilloscope for Firmware Engineers

A logic analyzer and an oscilloscope look like they answer the same question — "what happened on this wire" — and firmware engineers who own one tend to reach for it for everything, because a protocol decode is easier to read than a wobbly trace. They are not the same instrument. A logic analyzer's front end does one thing: at every sample instant, compare the voltage to a threshold and emit a 0 or a 1. Everything about how the signal got to that voltage — how fast, how far past it, whether it wobbled on the way — is discarded before the decoder ever sees a bit. An oscilloscope keeps that information. It is not a better logic analyzer; it answers a different class of question, and it is the only instrument in this folder that can.

The Register Model

A Cortex-M shows you sixteen 32-bit registers at any moment, and thirteen of them are genuinely interchangeable scratch space. The other three are the processor's control surface — a stack pointer that is secretly one of two, a link register that sometimes holds a return address and sometimes holds a magic number, and a program counter that is one bit weirder than it looks. Alongside those sit half a dozen special registers that are not in the main file at all: you cannot load or store them, and the only way to touch them is a pair of dedicated instructions.

The RTOS Landscape

Every kernel on this page implements fixed-priority pre-emptive scheduling with a ready list per priority, blocking primitives, and a tick. If you compare them on scheduling behaviour you will find almost nothing to choose between them, and you will have compared the wrong thing. The scheduler is the commodity part.

The Superloop and Cooperative Scheduling

Every embedded program is an infinite loop. The interesting question is what is inside it. A while(1) that calls three functions in order is the simplest architecture that can run a device, it ships in an enormous number of products, and it is entirely capable of being the right answer for the lifetime of a project. It is also the architecture that fails most quietly when it stops being the right answer, because nothing breaks — the loop just gets slower, and one day a button press is dropped.

Thumb-2 and Code Density

On a desktop CPU, instruction encoding is an implementation detail you can go a whole career without thinking about. On a microcontroller it is a budget. The STM32F411RE has 512 KB of flash and 128 KB of RAM, and that flash figure is one of the two or three numbers that set the price of the chip. Every byte an instruction occupies is a byte of product cost, multiplied by the production run.

Tickless Idle

Tasks and Scheduling established that the tick is a periodic interrupt — SysTick, by default 1000 Hz — that the kernel uses as its only window onto elapsed time. That periodicity is exactly what makes an RTOS system a bad candidate for the deep sleep modes earlier in this folder, for a reason that has nothing to do with how much idle time there is: a system that wakes once a millisecond to increment a counter never gets to stay in a low-power mode for longer than a millisecond, no matter how little else it has to do. Sleep Modes showed that Stop mode's wake latency comes from a regulator switch-back and, at the deepest setting, a flash re-power — real time that has to be paid on every entry and exit. Pay it a thousand times a second and the average current never gets anywhere near the deep-sleep floor in the table on that page; it is dominated by the wake-and-sleep-again overhead, not by the sleep current itself.

Timers and Counters

A timer is the least interesting peripheral to describe and the one you will use most. It is a counter that increments on a clock edge, and a comparator that notices when the count reaches a number you chose. Everything else in the chapter — PWM, input capture, encoder decoding, one-pulse output, triggering the ADC — is that counter with different plumbing bolted onto the comparator.

UART in Depth

A UART has no clock wire. That single fact generates every interesting property of the peripheral and every way it fails. SPI and I²C both ship a clock alongside the data, so the receiver is told exactly when to look; a UART receiver is told nothing. It sees a falling edge, starts its own counter, and from that moment guesses where the bit centres are using an oscillator the transmitter has never met. Everything below — the divider arithmetic, the oversampling modes, the tolerance budget, the overrun flag — is machinery built around that one guess.

Unit Testing Firmware

The ordinary embedded development loop is: change a line, build, flash, power-cycle or reset, and watch a UART or an LED to see whether the change did what you meant. On a fast board with a fast probe that loop is a few seconds; on a slow one, or one shared with a bench full of other work, it is tens of seconds, and every one of those seconds is spent re-verifying code paths that had nothing to do with the line you just changed. A state machine with a dozen transitions and a handful of edge cases does not get thoroughly exercised at that cadence — it gets exercised until the one case someone thought to try passes, and the rest ride along untested until a customer finds them.

Voltage Levels and Logic

There is no 1 on a wire. There is a voltage, and there is a receiver that has decided in advance which range of voltages it will call one and which it will call zero. Digital logic is an agreement layered on top of an analogue quantity, and the reason firmware normally gets to ignore that is that the agreement usually holds — the hardware on both ends was designed to the same convention, so the bits you read are the bits that were sent.

Wake Sources and Event-Driven Design

Every page so far in this folder has been about what happens once the core decides to sleep. This one is about the decision that matters more than any of them 90 ms of idle Run mode cost roughly two orders of magnitude more than the same 90 ms in Stop.

Watchdogs

A watchdog does not detect that your software is wrong. It detects that one specific piece of code stopped executing, and it reboots the system when that happens. Everything about designing a watchdog into a product follows from taking that sentence literally: the watchdog proves exactly what you make it prove, and not one thing more.

What "Embedded" Actually Means

Ask most software engineers what "embedded" means and they'll say something about small chips, or soldering, or blinking an LED. That's the wrong mental model, and it's why so many engineers who are perfectly competent on servers and desktops write their first firmware the way they'd write a desktop app — and then spend a week debugging failures that a desktop never produces. Embedded isn't defined by the chip. A phone's application processor and a pacemaker's microcontroller are both "chips," and the software practices around them could not be more different. What actually defines the field is a fixed set of constraints that never fully goes away, no matter how big or small the target is. Understand the constraints, and the rest of this section — why bare-metal code looks the way it does, why an RTOS exists, why "just add more RAM" isn't always an option — falls out as a consequence rather than a pile of arbitrary rules to memorize.

What "Real-Time" Actually Means

"Real-time" is the most abused word in embedded engineering. It is used to mean fast, to mean interrupt-driven, to mean "there is an RTOS in the build", and occasionally to mean nothing at all beyond marketing. None of those is the definition. A real-time system is one whose correctness depends on when a result is produced as well as on what the result is. A right answer delivered late is a wrong answer. That is the whole idea, and everything else on this page is a consequence of it.

What Hardware to Buy

Firmware is the one branch of software engineering where you genuinely cannot do the work on the machine you write the code on. A simulator will run your main(), but it will not show you that the sensor holds the clock line low for 40 microseconds longer than the datasheet suggests, that your board browns out when the motor starts, or that the pin you thought was an output has been floating since reset. Every important lesson in this section arrives through a physical board, and the reason newcomers stall here is not the money — the whole kit costs less than a mid-range monitor — but the catalogue. There are hundreds of development boards, every tutorial assumes a different one, and nothing on the vendor's site tells you which one the thing you are reading was written against.

What volatile Does and Does Not Do

volatile has a reputation for being either a magic word that makes hardware access work or a deprecated relic that nobody should use. Both readings come from the same place: people learn what it does by observing that adding it fixed a bug, and never learn the boundary of the guarantee. The boundary is narrow, it is written down precisely in the C standard, and knowing exactly where it stops is what separates code that works from code that works on your desk.

Why an RTOS

An RTOS does not make anything faster. It runs the same instructions on the same core at the same clock, and it adds work — a tick interrupt, a scheduler, a context switch — so a system that adopts one gets strictly less CPU time for its application. Anybody selling a kernel on performance is selling the wrong thing.

Worst-Case Execution Time

Every piece of real-time analysis needs one number per task: how long it takes to run. Not how long it usually takes — how long it can possibly take. That number, the worst-case execution time, is the input to response-time analysis, to utilisation budgets, and to every claim that a deadline is met. It is also the hardest number in embedded engineering to obtain honestly, because the obvious way to get it — run the code and time it — cannot produce it even in principle.

Writing a Driver Worth Reusing

Most firmware drivers are written once, for one board, and thrown away at the next project — not because the author was careless but because the driver was never separable from the chip it was born on. It calls HALI2CMaster_Transmit. It hard-codes I2C1. It has a static state variable so there can only ever be one. It blocks forever waiting on a status bit. Each of those is a small local convenience and together they weld the driver to one MCU, one instance, one board and one timing assumption.

Writing Interrupt Handlers in C

An interrupt handler is the only function in your program that nothing calls. You write it, you never reference it, and yet it runs — sometimes millions of times a second, at a moment you did not choose, on top of whatever the main program happened to be doing. That inversion is the whole difficulty. Every rule below follows from it: you cannot pass arguments to something nobody calls, you cannot return a value to nobody, and you cannot assume anything about the state of the code you interrupted.

Your First Bare-Metal Blink

Blinking an LED from an Arduino sketch takes two lines and teaches nothing about the machine. Blinking one with no HAL, no IDE and no library takes about ninety lines spread over five files, and by the end you know where every byte of the image came from, what the processor did before your first instruction, and which two writes in the whole program actually made the light change.

Zephyr in Practice

Everything up to here has treated the RTOS as a library. FreeRTOS is a handful of .c files you add to a project you already own: you keep your linker script, your startup code, your CubeMX-generated HAL_Init(), your main(). The kernel schedules; you do everything else. Swapping FreeRTOS for ThreadX would change the function names and almost nothing structural.