Skip to main content

Reading Assembly Output

Understanding compiler-generated assembly helps verify optimizations, debug performance issues, and understand low-level behavior.

Why Read Assembly?

Verify optimizations, debug weird performance, understand compiler decisions, learn CPU architecture. Not needed daily, but invaluable when needed.

Generating Assemblyโ€‹

# Generate assembly file
g++ -S -O2 program.cpp -o program.s

# With Intel syntax (more readable)
g++ -S -O2 -masm=intel program.cpp -o program.s

# Optimized assembly
g++ -S -O3 -march=native program.cpp -o program.s

# Disassemble existing binary
objdump -d program > program.asm
objdump -d -M intel program > program.asm # Intel syntax

# Compiler Explorer (online)
# https://godbolt.org - instant assembly view

AT&T vs Intel Syntaxโ€‹

# AT&T syntax (GCC default)
movl $5, %eax # Source, Destination
addl %ebx, %eax # Add ebx to eax

# Intel syntax (more readable)
mov eax, 5 # Destination, Source
add eax, ebx # Add ebx to eax

Use Intel syntax: Easier to read.

Basic x86-64 Registersโ€‹

RegisterPurpose64-bit32-bit16-bit8-bit
AccumulatorMath/returnRAXEAXAXAL
BaseBase pointerRBXEBXBXBL
CounterLoop counterRCXECXCXCL
DataDataRDXEDXDXDL
Stack PointerStack topRSPESPSPSPL
Base PointerStack frameRBPEBPBPBPL
Source IndexString opsRSIESISISIL
Dest IndexString opsRDIEDIDIDIL
R8-R15GeneralR8-R15R8D-R15DR8W-R15WR8B-R15B

Special registers:

  • rip = Instruction pointer
  • rflags = Flags (zero, carry, etc.)

Common Instructionsโ€‹

# Data movement
mov rax, rbx # Move: rax = rbx
lea rax, [rbx + 8] # Load effective address
push rax # Push to stack
pop rax # Pop from stack

# Arithmetic
add rax, rbx # rax = rax + rbx
sub rax, rbx # rax = rax - rbx
imul rax, rbx # rax = rax * rbx (signed)
idiv rbx # rax = rdx:rax / rbx (signed)
inc rax # rax++
dec rax # rax--

# Bitwise
and rax, rbx # rax = rax & rbx
or rax, rbx # rax = rax | rbx
xor rax, rbx # rax = rax ^ rbx
not rax # rax = ~rax
shl rax, 2 # rax << 2
shr rax, 2 # rax >> 2 (logical)
sar rax, 2 # rax >> 2 (arithmetic, sign-extend)

# Comparison
cmp rax, rbx # Compare (sets flags)
test rax, rax # AND and set flags (often: is zero?)

# Control flow
jmp label # Unconditional jump
je label # Jump if equal (ZF=1)
jne label # Jump if not equal (ZF=0)
jl label # Jump if less (SFโ‰ OF)
jg label # Jump if greater
call function # Call function
ret # Return

Reading Simple Functionsโ€‹

int add(int a, int b) {
return a + b;
}
# Intel syntax
add:
lea eax, [rdi + rsi] # eax = rdi + rsi (a + b)
ret # Return (result in eax)

# Explanation:
# rdi = first argument (a)
# rsi = second argument (b)
# eax = return value
# lea = load effective address (fast add)

Function Call Convention (x86-64 System V)โ€‹

Arguments (in order):

  1. rdi
  2. rsi
  3. rdx
  4. rcx
  5. r8
  6. r9
  7. Stack (if more than 6)

Return value: rax

Caller-saved: rax, rcx, rdx, r8-r11
Callee-saved: rbx, rbp, r12-r15

int func(int a, int b, int c, int d, int e, int f, int g) {
return a + b + c + d + e + f + g;
}
func:
# a=rdi, b=rsi, c=rdx, d=rcx, e=r8, f=r9
add edi, esi # a += b
add edi, edx # a += c
add edi, ecx # a += d
add edi, r8d # a += e
add edi, r9d # a += f
mov eax, DWORD PTR [rsp+8] # g from stack
add eax, edi # result
ret

Optimization Examplesโ€‹

Loop Unrollingโ€‹

void zero_array(int* arr, int n) {
for (int i = 0; i < n; ++i) {
arr[i] = 0;
}
}
# -O0: Simple loop with increment

# -O3: Vectorized + unrolled
zero_array:
test esi, esi
jle .L1
pxor xmm0, xmm0 # Zero vector register
.L3:
movdqu [rdi], xmm0 # Store 16 bytes at once
movdqu [rdi+16], xmm0 # Store another 16
add rdi, 32 # Advance pointer
sub esi, 8 # Decrement counter (8 ints)
jg .L3 # Loop
.L1:
ret

Optimized: Processes 8 ints per iteration with SIMD.

Constant Foldingโ€‹

int compute() {
return 10 * 20 + 5;
}
# -O0: Actual math at runtime
# -O2: Computed at compile time
compute:
mov eax, 205 # Result precomputed!
ret

Inliningโ€‹

inline int square(int x) {
return x * x;
}

int use_square(int n) {
return square(n) + 1;
}
# -O2: Function inlined, no call
use_square:
imul edi, edi # n * n (inlined)
lea eax, [rdi + 1] # result + 1
ret

Identifying Bottlenecksโ€‹

# Good: Fast operations
mov eax, ebx # 1 cycle
add eax, 5 # 1 cycle
lea eax, [rbx + 8] # 1 cycle

# Slow: Division
idiv ecx # ~20-40 cycles

# Slow: Memory access (if cache miss)
mov eax, [rbx] # 1 cycle (L1), ~200 cycles (RAM)

# Good: Vectorized
movdqu xmm0, [rdi] # Load 16 bytes at once

Stack Frameโ€‹

void function(int x) {
int local = x + 5;
// ...
}
function:
push rbp # Save old base pointer
mov rbp, rsp # New base pointer
sub rsp, 16 # Allocate stack space

mov DWORD PTR [rbp-4], edi # Store x
mov eax, DWORD PTR [rbp-4]
add eax, 5
mov DWORD PTR [rbp-8], eax # Store local

leave # Restore rbp, rsp
ret

Stack layout:

High addresses
+----------------+
| Return address | <- rsp on entry
+----------------+
| Old rbp | <- rbp points here after push
+----------------+
| Local variable | <- rbp-4, rbp-8, etc.
+----------------+
Low addresses

Compiler Explorer (Godbolt)โ€‹

Online tool for instant assembly viewing.

URL: https://godbolt.org

// Paste code, see assembly instantly
int add(int a, int b) {
return a + b;
}

Features:

  • Compare compilers (GCC, Clang, MSVC)
  • Compare optimization levels
  • Color-coded source โ†” assembly mapping
  • Share links

Reading Disassembled Binaryโ€‹

# Disassemble specific function
objdump -d -M intel program | grep -A 20 "^[0-9a-f]* <main>:"

# Disassemble with source interleaved
objdump -S -M intel program

# GDB disassembly
gdb ./program
(gdb) disassemble main
(gdb) disas /m main # With source

Common Patternsโ€‹

Null Checkโ€‹

if (ptr == nullptr) return;
test rdi, rdi # Test if rdi is zero
je .L_return # Jump if zero

Loopโ€‹

for (int i = 0; i < n; ++i)
xor eax, eax # i = 0
.L_loop:
cmp eax, esi # Compare i with n
jge .L_done # Jump if i >= n
# ... loop body ...
inc eax # i++
jmp .L_loop
.L_done:

Function Callโ€‹

int result = func(a, b, c);
mov edi, eax # First arg (a)
mov esi, ebx # Second arg (b)
mov edx, ecx # Third arg (c)
call func # Call function
mov [result], eax # Store return value

Summaryโ€‹

Generate: g++ -S -O2 -masm=intel. Registers: rax (return), rdi/rsi/rdx/rcx/r8/r9 (args). Common: mov (copy), add/sub (math), cmp (compare), jmp/je/jne (branch), call/ret (function). Optimizations: Inlining (no call), constant folding (precompute), vectorization (SIMD), unrolling (fewer iterations). Tools: Compiler Explorer (godbolt.org), objdump -d. Focus on hot paths, not every line.

; Memory aid: "MACJ" (Mov Add Cmp Jmp)
; M = mov (data movement)
; A = add/sub (arithmetic)
; C = cmp/test (comparison)
; J = jmp/je/call (control flow)

; Reading assembly:
; 1. Find function entry
; 2. Identify prologue (push rbp, mov rbp rsp)
; 3. Track register usage (rdi=arg1, rsi=arg2, rax=return)
; 4. Follow control flow (jmp, je, call)
; 5. Find epilogue (leave, ret)