A reading path for CPU performance engineering, from one instruction to production inference. Primary sources only, with a runnable benchmark for every section.
See the code
Making a program fast on a modern CPU means knowing what the core does with each instruction, where the time actually goes, and how to prove a change helped. This is the reading that gets you there, in the order that makes the next piece legible.
Scope. x86 and Arm server parts, from one instruction through to serving a model on CPU. Not language runtimes, database internals, or anything above the socket.
Evidence. Primary sources only: the paper, the specification, the vendor manual, the repository, or a report by the person who did the work. Any number, anywhere in this repository, carries all seven fields set out in What earns a place, or it is not quoted.
Proof. Fourteen of the sections end in a benchmark under misc/benchmarks/: C source, the build line, the machine, the raw numbers and the analysis, all committed. Run them yourself.
Section 1 is a path through the rest; read it top to bottom before using the numbered sections as a reference.
Each entry assumes only the ones before it. The first and sixth are paid books; the third is a manual to open at the chapters its reason names.
Work the exercises in Performance Ninja alongside them; reading alone will not build the instinct.
Reproduce it: misc/benchmarks/04-cache-latency, the cost per dependent load stepping up at each cache boundary, the curve the fourth entry measures.
Vendors name the same structures differently (Intel's decoded ICache is AMD's op cache, and AMD's macro-op is Arm's MOP), so the stage names below are generic.
Reproduce it: misc/benchmarks/02-branch-misprediction, the cost of a mispredicted branch, sorted against unsorted against branchless.
Where a design paper, the vendor manual and a measurement disagree about a core, the measurement is the one to trust and re-run.
Reproduce it: misc/benchmarks/03-latency-vs-throughput, one dependency chain against eight independent accumulators.
Line size and page size are machine parameters, not constants, so every padding and alignment rule below is applied against the target's own values.
Reproduce it: misc/benchmarks/04-cache-latency, dependent-load latency from L1 to DRAM, with and without TLB pressure.
Cloud instances often virtualise the hardware counters away and macOS runs no Linux perf, so perf stat has to show a non-zero cycles count before any counter entry below is trusted.
Reproduce it: misc/benchmarks/05-measurement-pitfalls, dead-code elimination, run-to-run spread, cold against warm.
A roofline is a bound built from measured roofs and counted bytes, so a point above a roof means a wrong roof or a wrong byte count, not fast code.
Reproduce it: misc/benchmarks/06-roofline, measured roofs and three kernels of rising arithmetic intensity.
A loop the compiler reports as vectorised can still run at scalar speed: a float reduction stays one serial chain until reassociation is permitted.
Reproduce it: misc/benchmarks/07-aos-vs-soa-simd, array of structs against structure of arrays, scalar against NEON.
No -O level changes the target instruction set: without -march or -mcpu, every instruction emitted belongs to the default target ISA, so target flags come before any judgement of codegen.
Reproduce it: misc/benchmarks/08-autovectorization-aliasing, the vectoriser with and without restrict.
Every cost below is a cache line moving between cores, so the ordering models and the measured line-transfer cost in the memory hierarchy section come first.
Reproduce it: misc/benchmarks/09-false-sharing, adjacent counters against padded counters across threads.
A page's node is decided at first touch, not when memory is allocated or a policy is set, and every vendor table below depends on the BIOS node mode of the machine it ran on.
Reproduce it: misc/benchmarks/10-first-touch, first touch of fresh pages against the second pass.
A syscall's cost depends on the mitigation state, the governor and the idle state the core was in, three sysfs settings that change after boot, so each is recorded beside any number below.
Reproduce it: misc/benchmarks/11-syscall-cost, the fixed cost of a kernel crossing across request sizes.
A latency figure means nothing without its percentile, its load model and the way it was recorded.
Reproduce it: misc/benchmarks/12-coordinated-omission, closed-loop against open-loop p99 under the same stalls.
A decode step at batch one reads every weight once for a few flops, so it runs at the memory system's rate, and a CPU figure compares with a GPU figure only at the same batch and precision.
Reproduce it: misc/benchmarks/13-sgemm-naive-vs-blas, naive GEMM, a hand microkernel and the vendor BLAS, plus int8 against float32 dot products.
The measurement articles below state no compiler, flags or run count, so each is kept for the structure it exposes and no figure from it is repeated.
Reproduce it: misc/benchmarks/14-pcore-vs-ecore, the same three kernels on a performance core and an efficiency core.
A score means what its suite's run rules say it means, so the rules come before the number.
Reproduce it: misc/benchmarks/15-stream-bandwidth, triad bandwidth by thread count against the vendor figure.
Everything below is real but unproven: no item yet has all three of a written specification, a part you can buy, and a public measurement stating every one of the seven fields. Each line says what would promote it. Vendor multiples never qualify. Last checked 2026-09-15.
Seven fields, and a number without all of them does not appear here:
| 1 | CPU model and microarchitecture |
| 2 | core count used |
| 3 | frequency, with turbo and SMT state |
| 4 | compiler and flags |
| 5 | workload |
| 6 | baseline |
| 7 | measurement method |
Miss one and the number is dropped; if the entry rests on that number, it moves to the watchlist or goes.
An entry itself has to be the thing, not writing about the thing: the paper
that first described a mechanism, the specification or manual that defines
it, the repository the implementation lives in, or a report from whoever did
the work with code and reproducible measurements. Summaries, tutorials,
surveys, marketing pages, mirrors and repackagings do not qualify. Every URL
points at the live canonical copy, and misc/scripts/check_links.py and
misc/scripts/check_format.py prove it on every push and again weekly.
CONTRIBUTING.md has the rules in full.
MIT. Maintained by @usamahz.
Format inspired by the GPU-side list at wafer-ai/gpu-perf-engineering-resources.
C
61.9%
Shell
29.1%
Python
9.1%
A reading path for CPU performance engineering, from one instruction to production inference. Primary sources only, with a runnable benchmark for every section.
See the code
Making a program fast on a modern CPU means knowing what the core does with each instruction, where the time actually goes, and how to prove a change helped. This is the reading that gets you there, in the order that makes the next piece legible.
Scope. x86 and Arm server parts, from one instruction through to serving a model on CPU. Not language runtimes, database internals, or anything above the socket.
Evidence. Primary sources only: the paper, the specification, the vendor manual, the repository, or a report by the person who did the work. Any number, anywhere in this repository, carries all seven fields set out in What earns a place, or it is not quoted.
Proof. Fourteen of the sections end in a benchmark under misc/benchmarks/: C source, the build line, the machine, the raw numbers and the analysis, all committed. Run them yourself.
Section 1 is a path through the rest; read it top to bottom before using the numbered sections as a reference.
Each entry assumes only the ones before it. The first and sixth are paid books; the third is a manual to open at the chapters its reason names.
Work the exercises in Performance Ninja alongside them; reading alone will not build the instinct.
Reproduce it: misc/benchmarks/04-cache-latency, the cost per dependent load stepping up at each cache boundary, the curve the fourth entry measures.
Vendors name the same structures differently (Intel's decoded ICache is AMD's op cache, and AMD's macro-op is Arm's MOP), so the stage names below are generic.
Reproduce it: misc/benchmarks/02-branch-misprediction, the cost of a mispredicted branch, sorted against unsorted against branchless.
Where a design paper, the vendor manual and a measurement disagree about a core, the measurement is the one to trust and re-run.
Reproduce it: misc/benchmarks/03-latency-vs-throughput, one dependency chain against eight independent accumulators.
Line size and page size are machine parameters, not constants, so every padding and alignment rule below is applied against the target's own values.
Reproduce it: misc/benchmarks/04-cache-latency, dependent-load latency from L1 to DRAM, with and without TLB pressure.
Cloud instances often virtualise the hardware counters away and macOS runs no Linux perf, so perf stat has to show a non-zero cycles count before any counter entry below is trusted.
Reproduce it: misc/benchmarks/05-measurement-pitfalls, dead-code elimination, run-to-run spread, cold against warm.
A roofline is a bound built from measured roofs and counted bytes, so a point above a roof means a wrong roof or a wrong byte count, not fast code.
Reproduce it: misc/benchmarks/06-roofline, measured roofs and three kernels of rising arithmetic intensity.
A loop the compiler reports as vectorised can still run at scalar speed: a float reduction stays one serial chain until reassociation is permitted.
Reproduce it: misc/benchmarks/07-aos-vs-soa-simd, array of structs against structure of arrays, scalar against NEON.
No -O level changes the target instruction set: without -march or -mcpu, every instruction emitted belongs to the default target ISA, so target flags come before any judgement of codegen.
Reproduce it: misc/benchmarks/08-autovectorization-aliasing, the vectoriser with and without restrict.
Every cost below is a cache line moving between cores, so the ordering models and the measured line-transfer cost in the memory hierarchy section come first.
Reproduce it: misc/benchmarks/09-false-sharing, adjacent counters against padded counters across threads.
A page's node is decided at first touch, not when memory is allocated or a policy is set, and every vendor table below depends on the BIOS node mode of the machine it ran on.
Reproduce it: misc/benchmarks/10-first-touch, first touch of fresh pages against the second pass.
A syscall's cost depends on the mitigation state, the governor and the idle state the core was in, three sysfs settings that change after boot, so each is recorded beside any number below.
Reproduce it: misc/benchmarks/11-syscall-cost, the fixed cost of a kernel crossing across request sizes.
A latency figure means nothing without its percentile, its load model and the way it was recorded.
Reproduce it: misc/benchmarks/12-coordinated-omission, closed-loop against open-loop p99 under the same stalls.
A decode step at batch one reads every weight once for a few flops, so it runs at the memory system's rate, and a CPU figure compares with a GPU figure only at the same batch and precision.
Reproduce it: misc/benchmarks/13-sgemm-naive-vs-blas, naive GEMM, a hand microkernel and the vendor BLAS, plus int8 against float32 dot products.
The measurement articles below state no compiler, flags or run count, so each is kept for the structure it exposes and no figure from it is repeated.
Reproduce it: misc/benchmarks/14-pcore-vs-ecore, the same three kernels on a performance core and an efficiency core.
A score means what its suite's run rules say it means, so the rules come before the number.
Reproduce it: misc/benchmarks/15-stream-bandwidth, triad bandwidth by thread count against the vendor figure.
Everything below is real but unproven: no item yet has all three of a written specification, a part you can buy, and a public measurement stating every one of the seven fields. Each line says what would promote it. Vendor multiples never qualify. Last checked 2026-09-15.
Seven fields, and a number without all of them does not appear here:
| 1 | CPU model and microarchitecture |
| 2 | core count used |
| 3 | frequency, with turbo and SMT state |
| 4 | compiler and flags |
| 5 | workload |
| 6 | baseline |
| 7 | measurement method |
Miss one and the number is dropped; if the entry rests on that number, it moves to the watchlist or goes.
An entry itself has to be the thing, not writing about the thing: the paper
that first described a mechanism, the specification or manual that defines
it, the repository the implementation lives in, or a report from whoever did
the work with code and reproducible measurements. Summaries, tutorials,
surveys, marketing pages, mirrors and repackagings do not qualify. Every URL
points at the live canonical copy, and misc/scripts/check_links.py and
misc/scripts/check_format.py prove it on every push and again weekly.
CONTRIBUTING.md has the rules in full.
MIT. Maintained by @usamahz.
Format inspired by the GPU-side list at wafer-ai/gpu-perf-engineering-resources.
C
61.9%
Shell
29.1%
Python
9.1%