Skip to main content

4 docs tagged with "latency"

View all tags

Interrupt Latency

Ask what the interrupt latency of a Cortex-M4 is and you get "12 cycles", which is true and almost never the number you need. Twelve cycles is what the processor contributes when the memory is instant, nothing else is running, and the exception arrives at a convenient moment. The number that decides whether your product works is the one measured from the event at the pin to the first instruction of your handler that does something about it — and on a real board, at a real clock, with the rest of your firmware present, that number is dominated by things the core designer had no say in.

Latency, Throughput, and Latency Hiding

Latency and throughput sound like the same idea measured two ways, but a GPU treats them as almost unrelated design targets. Latency is how long one memory request takes to come back; throughput is how many bytes per second the memory system can sustain in steady state. A single DRAM access on a modern GPU takes several hundred nanoseconds — not meaningfully faster than it was a decade ago — yet the same hardware sustains terabytes per second in aggregate. The only way to reconcile a slow individual request with a fast aggregate rate is to have an enormous number of requests outstanding at once, and that single fact is why the CUDA programming model insists you expose thousands of threads instead of a handful.

What "Real-Time" Actually Means

"Real-time" is the most abused word in embedded engineering. It is used to mean fast, to mean interrupt-driven, to mean "there is an RTOS in the build", and occasionally to mean nothing at all beyond marketing. None of those is the definition. A real-time system is one whose correctness depends on when a result is produced as well as on what the result is. A right answer delivered late is a wrong answer. That is the whole idea, and everything else on this page is a consequence of it.