Buses & I/O — Overview
A CPU, RAM, storage, and peripherals are separate physical chips — a bus is the shared electrical
A CPU, RAM, storage, and peripherals are separate physical chips — a bus is the shared electrical
Almost every mechanism that makes a computer fast makes it less predictable, and the trade is usually invisible in source code. Caches, prefetchers, branch speculation, bus arbitration and dynamic memory allocation are all bets on history repeating: they are fast when the recent past resembles the present and slow when it does not. In a throughput-oriented system that is an unambiguously good bargain. In a system with deadlines it means the number you have to defend — the worst case — is set by the unlucky path through every one of them, while every measurement you take is dominated by the lucky one.
A DMA controller is not an accelerator bolted onto a peripheral. It is a second bus master: a small, dumb machine that sits on the same bus matrix as the Cortex-M4 and, when a peripheral raises a request line, performs the load and the store that your interrupt handler would otherwise have performed. It has no idea what the data means. It knows a source address, a destination address, a count, and whether to increment each pointer.
Peripherals are slow compared to a CPU: a disk read can take milliseconds, which is millions of CPU
Here is the observation this page exists to explain. You have a fault that happens once every few minutes — a corrupted buffer, a missed deadline, a state machine that ends up somewhere it cannot reach. You add a printf to narrow it down. The fault stops happening. You remove the printf and it comes back.
There are exactly three ways to get a byte out of a peripheral register and into your program's memory. The CPU can ask repeatedly until the answer is yes. The peripheral can raise a line that makes the CPU stop what it was doing. Or a second bus master can do the load and the store on the CPU's behalf and tell it afterwards. Every driver you will ever write picks one of these, and the choice is made badly far more often than it is made wrong — badly, meaning by reflex rather than from a budget.
A UART has no clock wire. That single fact generates every interesting property of the peripheral and every way it fails. SPI and I²C both ship a clock alongside the data, so the receiver is told exactly when to look; a UART receiver is told nothing. It sees a falling edge, starts its own counter, and from that moment guesses where the bit centres are using an oscillator the transmitter has never met. Everything below — the divider arithmetic, the oversampling modes, the tolerance budget, the overrun flag — is machinery built around that one guess.
Every page so far in this folder has been about what happens once the core decides to sleep. This one is about the decision that matters more than any of them 90 ms of idle Run mode cost roughly two orders of magnitude more than the same 90 ms in Stop.
volatile has a reputation for being either a magic word that makes hardware access work or a deprecated relic that nobody should use. Both readings come from the same place: people learn what it does by observing that adding it fixed a bug, and never learn the boundary of the guarantee. The boundary is narrow, it is written down precisely in the C standard, and knowing exactly where it stops is what separates code that works from code that works on your desk.