Skip to main content

62 docs tagged with "stm32"

View all tags

A GPIO Driver from Scratch

The blink program drove one pin with two #defines and it was the right amount of code for one pin. The second pin costs another two, the first alternate-function pin costs four, and by the tenth you have shift arithmetic scattered across the codebase with the pin number written out by hand in each place. The fix is not a HAL. It is about eighty lines that name what the hardware already does.

ADC and DAC Drivers

A successive-approximation ADC is, physically, a capacitor with a switch in front of it. Converting a voltage happens in two completely different phases: first the switch closes and the capacitor is allowed to charge towards your signal through the source impedance, and then the switch opens and a comparator plays twelve rounds of twenty questions against the trapped charge. The second phase is fixed by the hardware and takes exactly twelve clocks. The first phase is the one you configure, the one every tutorial leaves at its reset value, and the one that decides whether your reading means anything at all.

Analog Basics: ADC and DAC

An analog-to-digital converter looks, from firmware, like a register you read. That framing hides the two things that actually determine whether the number is right. First, the conversion is a comparison against a reference, so the answer is a ratio, not a voltage — and a wrong or noisy reference is invisible in the result. Second, before any comparing happens the converter must charge a small capacitor through your circuit, and if you did not give it long enough, it will confidently report the voltage it managed to reach rather than the voltage that was there.

Brownout and Power-Loss Safety

A supply does not fail like a light switch. Pull the plug, drop a battery connector, or let a coin cell finally reach the end of its discharge curve, and VDD does not fall to zero — it decays, over a span set by whatever capacitance is still holding charge on the rail and whatever the circuit is still drawing from it. That decay is the entire subject of this page, because it is the only thing standing between "the supply is gone" and "the core has actually stopped running," and the honest answer to "how long do I have" is: a lot less than most designs assume, and the amount is calculable rather than a matter of taste.

C Libraries: newlib, newlib-nano, picolibc

The C standard library is written as though a process exists. printf writes to a file descriptor. malloc asks the kernel for more address space. exit tells a parent that a child finished. fopen needs a filesystem, time needs a clock somebody set, errno needs somewhere thread-local to live. On a Cortex-M4 with no OS, not one of those things is true — and yet #include compiles fine and printf links, because the library was built with the bottom of every one of those paths left as a hole for you to fill.

Clock and Peripheral Gating

Two habits do most of the work in a low-power design that never touches Sleep, Stop, or Standby at all the energy cost of a fixed piece of work is not simply proportional to clock speed.

Clocks and Oscillators

On a desktop machine the clock is somebody else's problem — it was configured by firmware you never see, and by the time your program runs it is a constant. On a microcontroller you are that firmware. The chip comes out of reset running on a cheap internal RC oscillator at a fraction of its rated speed, with almost every peripheral's clock switched off, and the first job your code has is to build the clock tree the rest of the system will run on. Nothing you write behaves as intended until that is done.

CMake for Embedded

CMake's defaults encode one assumption so deeply that it is easy to miss: the machine running the build can also run what the build produces. That is what lets CMake test a compiler by compiling and linking a tiny program, what lets findpackage look in /usr/lib, and what lets checkcsourceruns exist at all. Every one of those assumptions is false for a Cortex-M4 with 128 KB of RAM and no operating system.

CMSIS and Vendor HALs

"Bare metal" does not have to mean "type every address yourself". Between raw pointer casts and a full vendor framework there are three or four distinct layers, each with a different bargain, and the useful skill is knowing which one you are standing on and why — not picking a side.

Configuring the Clock Tree

Out of reset the STM32F411RE runs at 16 MHz on an internal RC oscillator, which is a deliberately conservative choice: it works with no crystal, no configuration, and no risk. It is also one sixth of what the part can do, it is accurate to about ±1 % over temperature rather than the ±20 ppm a crystal gives, and it cannot produce the 48 MHz that USB requires. Somewhere in the first week of a real project you will need to change it.

Cross-Compilation

The compiler on your laptop is not a general-purpose translator that happens to be pointed at x86. It is a program that was built to emit x86-64 instructions, linked against a C library that was built assuming a Linux kernel is underneath it, and wired to a startup object that assumes something already created a process, set up a stack, and handed it argc and argv. Every one of those assumptions is false on a microcontroller. That is the whole reason a separate toolchain exists — not because the MCU is "different hardware", but because three independent layers of the build all encode a machine and an environment, and all three have to change together.

Determinism Killers

Almost every mechanism that makes a computer fast makes it less predictable, and the trade is usually invisible in source code. Caches, prefetchers, branch speculation, bus arbitration and dynamic memory allocation are all bets on history repeating: they are fast when the recent past resembles the present and slow when it does not. In a throughput-oriented system that is an unambiguously good bargain. In a system with deadlines it means the number you have to defend — the worst case — is set by the unlucky path through every one of them, while every measurement you take is dominated by the lucky one.

DMA

A DMA controller is not an accelerator bolted onto a peripheral. It is a second bus master: a small, dumb machine that sits on the same bus matrix as the Cortex-M4 and, when a peripheral raises a request line, performs the load and the store that your interrupt handler would otherwise have performed. It has no idea what the data means. It knows a source address, a destination address, a count, and whether to increment each pointer.

Energy Budgets

Every battery-powered product answers one question before any other design decision matters a firmware team that only measures active current and ignores what the part draws asleep will overestimate lifetime by orders of magnitude, and a team that only reads the nameplate mAh off a battery datasheet without reading the discharge curves will do the same in the opposite direction.

Exceptions and the Vector Table

On most processors, getting from "an interrupt line went high" to "my C function is running" involves software: a dispatcher reads a status register, works out which source fired, and calls the right handler. Cortex-M does none of that. The hardware reads a table of function pointers at a known address, indexes it by exception number, and branches — having already pushed the registers a C function is allowed to clobber. Your handler is an ordinary function with an ordinary prologue, and it starts running a fixed and small number of cycles after the event.

External Memory and QSPI

The moment your data stops fitting on-chip, the interesting question is not which memory — it is whether the processor has to execute from it, or merely read it. Those two requirements lead to completely different hardware. Data you read into a buffer can live behind four wires and a software driver, and a plain SPI port is enough. Code the CPU fetches instructions from must appear in the address map, which means a controller that turns a bus read into a flash transaction with no software involved at all.

Flashing and Programming

On a hosted system, "running the program" means handing a file to a loader. Here it means writing your image into non-volatile memory inside the chip and then resetting it, using a second piece of hardware that talks to the silicon over a two-wire debug port. That second piece of hardware is doing considerably more than copying bytes: it halts the core, drives the flash controller through a sequence the reference manual specifies, verifies, and releases reset.

Floating Point and DSP Extensions

Writing float in C on a microcontroller does not tell you what the hardware will do. The same line of source can compile to one instruction, to a forty-cycle library call, or to a two-hundred-cycle double-precision emulation — and the compiler chooses silently, based on flags you may not have set deliberately. Nothing in the source distinguishes the three cases, which is why "why is my control loop suddenly missing deadlines" is so often a floating-point question.

HardFault Forensics

A HardFault is not an error message. It is the processor announcing that it has already stopped being able to run your program, and that everything it knows about why is sitting in four registers and one stack frame. Nothing is printed, nothing is logged, and the default handler in every startup file ever shipped is b . — an infinite loop that discards all of it.

How a GPIO Pin Really Behaves

GPIOA->ODR |= (1 << 5); looks exactly like every other memory write you have ever done, and that resemblance is the problem. On the far side of that register is not a bit of storage but a pair of transistors, wired to a physical pin, with a maximum current, a maximum switching rate, and a real-world net on the other end that may already be being driven by something else. The register model hides all of it, right up until the moment it matters.

I2C in Depth

I²C is the only one of the three common serial buses whose electrical design is part of its protocol. SPI and UART drive their lines push-pull multiple controllers on the same two wires, targets that can pause the controller, collision detection that costs no extra hardware, and the ability to hang three sensors off two pins.

Input Capture and Encoders

PWM points the timer outward: the counter drives a pin. Input capture points it inward. The counter free-runs, an edge on a pin tells the hardware "now", and the value of CNT at that instant is copied into a capture register before software has had a chance to be late. That last clause is the entire value of the peripheral. A GPIO interrupt can also tell you an edge happened, but by the time your handler reads a counter it has been anywhere from 12 to several hundred cycles — jittering with whatever else the NVIC was doing — and the measurement carries that jitter. The capture unit's latch has no jitter at all.

Instruction and Event Tracing

There is a category of bug that a breakpoint destroys by looking at it, and a log line destroys almost as thoroughly. A race between two interrupts, an occasional priority inversion, a state machine that takes a wrong branch once in ten thousand iterations — halting the core to inspect any of these changes exactly the timing that produced them, and a printf line inserted to catch them costs enough cycles to close the window it was supposed to observe. The Debug Toolbox's perturbation table already ranks every instrument in this folder by how much it disturbs the thing it is measuring; trace is the answer at the bottom of that table, the one built specifically to be close to zero.

Internal Flash and EEPROM Emulation

Flash is not memory that happens to be non-volatile. It is a device with an asymmetric write model, and every design decision about storing settings on an MCU comes out of that asymmetry: you can clear a bit at any time, cheaply, one word at a time — but you cannot set a bit back to 1 without erasing an entire sector, which takes up to two seconds, during which the processor cannot fetch instructions from the same flash it is executing from.

Interrupt Latency

Ask what the interrupt latency of a Cortex-M4 is and you get "12 cycles", which is true and almost never the number you need. Twelve cycles is what the processor contributes when the memory is instant, nothing else is running, and the exception arrives at a convenient moment. The number that decides whether your product works is the one measured from the event at the pin to the first instruction of your handler that does something about it — and on a real board, at a real clock, with the rest of your firmware present, that number is dominated by things the core designer had no say in.

Measuring Power

Every figure in this folder so far has come from a datasheet table or from arithmetic built on one. Neither tells you what your board actually draws. A datasheet's Stop-mode current is measured on ST's own characterisation board, with nothing attached to the GPIOs, a specific silicon revision, and every peripheral in the state ST chose to test it in; your board has a pull-up ST did not model, a status LED that never got gated, and firmware that may or may not be entering the mode it claims to. Energy Budgets's entire arithmetic is only as good as the I_avg fed into it, and the only way to know that number is real is to measure it on the actual hardware running the actual release firmware — the warning on that page about a demo build with logging left on exists precisely because nobody measured until it was too late.

Memory Sections and VMA vs LMA

Every variable and every function in your firmware ends up in one of about six buckets, and which bucket it lands in is decided by two things: whether it is code or data, and whether its initial value is zero. That is nearly the whole rule. int counter; goes in .bss because its initial value is zero. int counter = 5; goes in .data because it is not. const int limit = 5; goes in .rodata because it never changes and can therefore stay in flash. Nobody chose those placements for your variable; the compiler applied that rule and emitted a section name.

Optimization for Size and Speed

Raising the optimization level is the one build change that routinely alters what a firmware does. Not what it does more quickly — what it does. A delay loop disappears. A register write that was there at -O0 is gone at -O2. Code that worked for two years starts failing, and nothing in the source changed.

Polling, Interrupt, or DMA

There are exactly three ways to get a byte out of a peripheral register and into your program's memory. The CPU can ask repeatedly until the answer is yes. The peripheral can raise a line that makes the CPU stop what it was doing. Or a second bus master can do the load and the store on the CPU's behalf and tell it afterwards. Every driver you will ever write picks one of these, and the choice is made badly far more often than it is made wrong — badly, meaning by reflex rather than from a budget.

Postmortem Debugging

Every technique in HardFault Forensics assumes something that is only true on your bench: a debugger attached at the moment of the fault, watching CFSR before anything clears it, holding the stack frame before the stack is reused. A device in the field has none of that. It faults, and unless you decided in advance what to do about it, the only evidence is a b . loop nobody is watching, or a reset that erases everything and starts the firmware running again as if nothing happened.

Power Supplies and Regulators

Firmware is written as though the supply rail were a constant — a number in the datasheet, 3.3 V, always there. The rail is not a constant. It is the output of a control loop with finite bandwidth, fed through traces with real resistance and inductance, feeding a load whose current draw your own code is modulating thousands of times a second. Every time the CPU switches from an idle loop to a burst of floating-point work, every time a GPIO drives an LED, every time the chip wakes from Stop mode, the load steps and the rail moves.

Priorities and Nesting

Every interrupt on a Cortex-M comes out of reset at priority 0 — the most urgent level there is. A program that enables six interrupts and never calls NVIC_SetPriority has six handlers that all sit at the top, none of which can pre-empt any other, serviced in exception-number order when several arrive at once. That configuration is not "no priority scheme". It is a specific, and usually wrong, priority scheme: it says every interrupt in the system is equally urgent and none may interrupt another, which means the worst-case latency of your fastest deadline is the sum of every other handler's execution time.

Privilege Modes and the Two Stacks

A Cortex-M has two independent switches that most bare-metal firmware never touches, and that an RTOS depends on completely. One decides which stack the processor is using; the other decides whether the code running is allowed to change anything important. They are separate — you can be unprivileged on the main stack or privileged on the process stack — and confusing them is the source of a lot of half-right explanations.

PWM

A microcontroller pin has two output voltages and nothing in between. Pulse-width modulation is the trick that gets the third: switch fast enough and whatever is downstream — an LED and your eye, a motor and its inductance, an RC filter and a slow ADC — averages the square wave into a level. The pin is still only ever fully on or fully off, which is why it dissipates almost no power doing it. That is the whole reason PWM won over analogue drive for everything from a status LED to a 10 kW inverter.

Reading a Datasheet

Coming from software, the instinct when you meet a new chip is to look for "the docs" — one document, searchable, that tells you everything. That document does not exist, and looking for it is the reason people bounce off hardware. Silicon vendors ship a set of documents, deliberately separated, because they answer questions that different people ask at different times: the person choosing a part, the person laying out the board, the person writing the firmware, and the person whose product works on the bench but fails one unit in fifty. Each document is written for one of those people and is close to useless for the others.

Reading the Map File

Every firmware project reaches the same afternoon. The build that fitted last week does not fit this week, or it fits but the RAM figure has doubled, and nobody changed anything that should have cost 12 KB. The instinct is to start deleting features. The correct move is to ask the toolchain, which has known the answer the whole time and wrote it down.

Register-Level Programming

A peripheral is a piece of digital logic sitting on the same bus as your RAM. It has no API, no calling convention and no way to be invoked. The only interface it exposes is a small block of addresses: write a word to one of them and some flip-flops change state; read from another and you get the current state of some wires. That is the whole model.

Reset and Boot Configuration

There is a gap between the moment power reaches the chip and the moment your first instruction executes, and firmware engineers habitually treat it as empty. It is not. In that gap the supply supervisor decides whether the rail is trustworthy, a pulse generator stretches whatever event caused the reset into a signal long enough for every block on the die to see it, an option-byte loader runs, boot-mode pins are sampled and latched, an address decoder is reconfigured so that a completely different memory appears at address zero, and only then does the CPU fetch two words and start running.

RTC and Timekeeping

Every other peripheral in this folder lives in your power domain, stops when you stop, and forgets everything when the supply goes away. The RTC does not. It is a small independent machine in a separate power domain with its own oscillator, its own supply pin, and its own reset — and the only things it shares with the rest of the chip are a bus interface and a couple of locks that exist specifically to stop your code from disturbing it by accident.

Sleep Modes

A low-power mode is not one switch, it is a set of trade-offs arranged along a slider first the CPU clock, then every clock and the main regulator, then the regulator's own supply to everything but a small backup island.

SPI in Depth

SPI is a shift register with a wire between two halves of it. That is the whole protocol, and holding it in mind explains everything the peripheral does. The controller has eight bits, the target has eight bits, and the clock the controller generates walks them past each other in a ring: the controller's MSB goes out on MOSI and into the target's LSB position, the target's MSB goes out on MISO and into the controller's. After eight clocks the two registers have swapped contents. There is no addressing, no acknowledgement, no error detection and no notion of a transaction — every one of those has to be built on top by whatever protocol the target's datasheet defines.

Startup Code: Reset to main

C has a set of guarantees that programmers stop noticing after the first month. Initialised globals hold their initialisers. Uninitialised globals are zero. The stack works. printf has somewhere to write. Static C++ objects have had their constructors run before main starts. On a hosted system a program loader and a crt0 object you have never opened deliver all of that before your first line executes.

SysTick and the Core Peripherals

Almost everything on a microcontroller is the vendor's. The timers, the UARTs, the clock tree, the GPIO blocks — all of it is ST's design, at ST's addresses, described in ST's reference manual, and none of it transfers to a part from a different vendor. A handful of peripherals are not like that. They are Arm's, they are built into the processor itself, they sit at the same addresses on every Cortex-M ever made, and code that uses them ports between vendors without changes.

The Anatomy of a Peripheral

An STM32F411RE has a UART, three SPIs, three I²C controllers, eight timers, an ADC, a USB controller, an RTC, two watchdogs and two DMA engines. The reference manual gives each of them thirty to eighty pages, and read front to back they look like thirteen unrelated pieces of hardware. They are not. They are thirteen instances of one design, drawn by the same team, wired onto the same two buses, and configured by the same six steps in the same order every time.

The Cortex-M Family

Arm does not sell chips. It sells processor designs, and a silicon vendor — ST, NXP, Nordic, Raspberry Pi — licenses one, wraps it in memory and peripherals, and sells you the result. That arrangement is why "it's an Arm chip" tells you almost nothing on its own, and why the useful question is always which Arm core, in which configuration, from which vendor.

The Cortex-M Memory Map

A microcontroller has one 4 GB address space and everything lives in it: flash, RAM, every peripheral register, the interrupt controller, the debug hardware. There is no MMU, so what you write in a pointer is the physical address the bus sees. That is the simplification that makes bare-metal firmware tractable — and it means the layout of that space is not a vendor's private business but part of the architecture you program against.

The Linker Script

On a hosted system nobody writes a linker script, because the answer to "where does this code go" is "wherever the loader decides", and the loader is part of the OS. On a microcontroller there is no loader. The addresses in the binary are the addresses the CPU will use, forever, and something has to choose them. That something is a text file, usually about a hundred lines long, that most projects copy once and never read.

The Memory Protection Unit

A null-pointer write on a Cortex-M does not crash. Address 0x00000000 is real memory — it is the start of flash, or the boot alias — so *(uint32t )0 = 42 on a fresh chip silently does nothing at all, and the program carries on. A stack that overflows its intended region does not crash either; it grows down into .bss and corrupts variables that belong to something else, and the failure surfaces minutes later in code that is entirely innocent. Both are the same problem: on a bare Cortex-M, memory has no permissions, so a wrong access is indistinguishable from a right one.*

The NVIC

The vector table answers where. The Nested Vectored Interrupt Controller answers whether, when, and in what order — and it is the only part of the exception model you actively program. Every interrupt line on the chip arrives at the NVIC; the NVIC decides whether that line is enabled, records it as pending if it is not yet serviceable, compares its priority against whatever the processor is currently doing, and hands the winner to the core.

The Oscilloscope for Firmware Engineers

A logic analyzer and an oscilloscope look like they answer the same question — "what happened on this wire" — and firmware engineers who own one tend to reach for it for everything, because a protocol decode is easier to read than a wobbly trace. They are not the same instrument. A logic analyzer's front end does one thing: at every sample instant, compare the voltage to a threshold and emit a 0 or a 1. Everything about how the signal got to that voltage — how fast, how far past it, whether it wobbled on the way — is discarded before the decoder ever sees a bit. An oscilloscope keeps that information. It is not a better logic analyzer; it answers a different class of question, and it is the only instrument in this folder that can.

The Register Model

A Cortex-M shows you sixteen 32-bit registers at any moment, and thirteen of them are genuinely interchangeable scratch space. The other three are the processor's control surface — a stack pointer that is secretly one of two, a link register that sometimes holds a return address and sometimes holds a magic number, and a program counter that is one bit weirder than it looks. Alongside those sit half a dozen special registers that are not in the main file at all: you cannot load or store them, and the only way to touch them is a pair of dedicated instructions.

Thumb-2 and Code Density

On a desktop CPU, instruction encoding is an implementation detail you can go a whole career without thinking about. On a microcontroller it is a budget. The STM32F411RE has 512 KB of flash and 128 KB of RAM, and that flash figure is one of the two or three numbers that set the price of the chip. Every byte an instruction occupies is a byte of product cost, multiplied by the production run.

Tickless Idle

Tasks and Scheduling established that the tick is a periodic interrupt — SysTick, by default 1000 Hz — that the kernel uses as its only window onto elapsed time. That periodicity is exactly what makes an RTOS system a bad candidate for the deep sleep modes earlier in this folder, for a reason that has nothing to do with how much idle time there is: a system that wakes once a millisecond to increment a counter never gets to stay in a low-power mode for longer than a millisecond, no matter how little else it has to do. Sleep Modes showed that Stop mode's wake latency comes from a regulator switch-back and, at the deepest setting, a flash re-power — real time that has to be paid on every entry and exit. Pay it a thousand times a second and the average current never gets anywhere near the deep-sleep floor in the table on that page; it is dominated by the wake-and-sleep-again overhead, not by the sleep current itself.

Timers and Counters

A timer is the least interesting peripheral to describe and the one you will use most. It is a counter that increments on a clock edge, and a comparator that notices when the count reaches a number you chose. Everything else in the chapter — PWM, input capture, encoder decoding, one-pulse output, triggering the ADC — is that counter with different plumbing bolted onto the comparator.

UART in Depth

A UART has no clock wire. That single fact generates every interesting property of the peripheral and every way it fails. SPI and I²C both ship a clock alongside the data, so the receiver is told exactly when to look; a UART receiver is told nothing. It sees a falling edge, starts its own counter, and from that moment guesses where the bit centres are using an oscillator the transmitter has never met. Everything below — the divider arithmetic, the oversampling modes, the tolerance budget, the overrun flag — is machinery built around that one guess.

Voltage Levels and Logic

There is no 1 on a wire. There is a voltage, and there is a receiver that has decided in advance which range of voltages it will call one and which it will call zero. Digital logic is an agreement layered on top of an analogue quantity, and the reason firmware normally gets to ignore that is that the agreement usually holds — the hardware on both ends was designed to the same convention, so the bits you read are the bits that were sent.

Wake Sources and Event-Driven Design

Every page so far in this folder has been about what happens once the core decides to sleep. This one is about the decision that matters more than any of them 90 ms of idle Run mode cost roughly two orders of magnitude more than the same 90 ms in Stop.

Watchdogs

A watchdog does not detect that your software is wrong. It detects that one specific piece of code stopped executing, and it reboots the system when that happens. Everything about designing a watchdog into a product follows from taking that sentence literally: the watchdog proves exactly what you make it prove, and not one thing more.

What Hardware to Buy

Firmware is the one branch of software engineering where you genuinely cannot do the work on the machine you write the code on. A simulator will run your main(), but it will not show you that the sensor holds the clock line low for 40 microseconds longer than the datasheet suggests, that your board browns out when the motor starts, or that the pin you thought was an output has been floating since reset. Every important lesson in this section arrives through a physical board, and the reason newcomers stall here is not the money — the whole kit costs less than a mid-range monitor — but the catalogue. There are hundreds of development boards, every tutorial assumes a different one, and nothing on the vendor's site tells you which one the thing you are reading was written against.

Worst-Case Execution Time

Every piece of real-time analysis needs one number per task: how long it takes to run. Not how long it usually takes — how long it can possibly take. That number, the worst-case execution time, is the input to response-time analysis, to utilisation budgets, and to every claim that a deadline is met. It is also the hardest number in embedded engineering to obtain honestly, because the obvious way to get it — run the code and time it — cannot produce it even in principle.

Writing Interrupt Handlers in C

An interrupt handler is the only function in your program that nothing calls. You write it, you never reference it, and yet it runs — sometimes millions of times a second, at a moment you did not choose, on top of whatever the main program happened to be doing. That inversion is the whole difficulty. Every rule below follows from it: you cannot pass arguments to something nobody calls, you cannot return a value to nobody, and you cannot assume anything about the state of the code you interrupted.

Your First Bare-Metal Blink

Blinking an LED from an Arduino sketch takes two lines and teaches nothing about the machine. Blinking one with no HAL, no IDE and no library takes about ninety lines spread over five files, and by the end you know where every byte of the image came from, what the processor did before your first instruction, and which two writes in the whole program actually made the light change.