usamahz/cpu-performance-engineering

A reading path for CPU performance engineering, from one instruction to production inference. Primary sources only, with a runnable benchmark for every section.

C

111

1 commits

updated Sep 27, 2026

See the code

See what people are saying

SourceMessageScoreDate

CPU Performance Engineering

1

Oct 1, 2026

README

CPU Performance Engineering

Links Quality Entries Benchmarks License Stars

Making a program fast on a modern CPU means knowing what the core does with each instruction, where the time actually goes, and how to prove a change helped. This is the reading that gets you there, in the order that makes the next piece legible.

Scope. x86 and Arm server parts, from one instruction through to serving a model on CPU. Not language runtimes, database internals, or anything above the socket.

Evidence. Primary sources only: the paper, the specification, the vendor manual, the repository, or a report by the person who did the work. Any number, anywhere in this repository, carries all seven fields set out in What earns a place, or it is not quoted.

Proof. Fourteen of the sections end in a benchmark under misc/benchmarks/: C source, the build line, the machine, the raw numbers and the analysis, all committed. Run them yourself.

Section 1 is a path through the rest; read it top to bottom before using the numbered sections as a reference.

Contents

1. Start here

Each entry assumes only the ones before it. The first and sixth are paid books; the third is a manual to open at the chapters its reason names.

  1. Computer Architecture: A Quantitative Approach, 7th Edition - Its pipelining appendix and memory chapters define the hazard, speculation and cache vocabulary the list assumes.
  2. Optimizing software in C++ - Maps C++ onto pipeline mechanisms and shows why a loop-carried dependency chain, not instruction count, paces a loop.
  3. Intel Optimization Reference Manual - Its opening chapters show how a shipping x86 core implements the textbook pipeline, each rule tied to a mechanism.
  4. What Every Programmer Should Know About Memory - Measures the step in cost per access at each cache boundary and the gap a prefetcher hides.
  5. Memory Barriers: a Hardware View for Software Hackers - Explains why a second core makes loads and stores reorder and what a barrier drains.
  6. Systems Performance: Enterprise and the Cloud, 2nd Edition - Puts the method before the tools: what to measure, in what order, and how benchmarks mislead.
  7. Roofline: An Insightful Visual Performance Model for Multicore Architectures - Places a loop from a byte count and a datasheet bandwidth alone, before any counter is read.
  8. A Top-Down Method for Performance Analysis and Counters Architecture - Defines the split of pipeline slots into front end, bad speculation, back end and retiring, the tree profilers report.
  9. Performance Analysis and Tuning on Modern CPUs - Walks from a noisy timing to counters to a named bottleneck, applying roofline and top-down to whole programs.
  10. What Has My Compiler Done for Me Lately? Unbolting the Compiler's Lid - Shows how to read emitted assembly against its source, so each mechanism is checked in a listing, not assumed.

Work the exercises in Performance Ninja alongside them; reading alone will not build the instinct.

Reproduce it: misc/benchmarks/04-cache-latency, the cost per dependent load stepping up at each cache boundary, the curve the fourth entry measures.

2. One instruction, end to end

Vendors name the same structures differently (Intel's decoded ICache is AMD's op cache, and AMD's macro-op is Arm's MOP), so the stage names below are generic.

Fetch and decode

Reproduce it: misc/benchmarks/02-branch-misprediction, the cost of a mispredicted branch, sorted against unsorted against branchless.

Rename and issue

Execute

Memory access and retire

3. Microarchitecture

Where a design paper, the vendor manual and a measurement disagree about a core, the measurement is the one to trust and re-run.

Limits of ILP and SMT

Branch prediction and speculation

Vendor estimates and measured tables

What the manuals leave out

Reproduce it: misc/benchmarks/03-latency-vs-throughput, one dependency chain against eight independent accumulators.

4. Memory hierarchy

Line size and page size are machine parameters, not constants, so every padding and alignment rule below is applied against the target's own values.

Cache geometry, replacement and misses in flight

Reproduce it: misc/benchmarks/04-cache-latency, dependent-load latency from L1 to DRAM, with and without TLB pressure.

TLBs, page walks and prefetchers

Store buffers, ordering and cache-line contention

Struct layout, software prefetch and page size

5. Measurement

Cloud instances often virtualise the hardware counters away and macOS runs no Linux perf, so perf stat has to show a non-zero cycles count before any counter entry below is trusted.

Method and the whole-system view

Counters, events and precise sampling

CPU profilers and flame graphs

Microbenchmarks that lie

Reproduce it: misc/benchmarks/05-measurement-pitfalls, dead-code elimination, run-to-run spread, cold against warm.

6. Models

A roofline is a bound built from measured roofs and counted bytes, so a point above a roof means a wrong roof or a wrong byte count, not fast code.

Roofline and the execution-cache-memory model

Reproduce it: misc/benchmarks/06-roofline, measured roofs and three kernels of rising arithmetic intensity.

Top-down analysis

Scaling laws

Queueing

7. Single-thread optimisation

A loop the compiler reports as vectorised can still run at scalar speed: a float reduction stays one serial chain until reassociation is permitted.

Data layout and loop transforms

Reproduce it: misc/benchmarks/07-aos-vs-soa-simd, array of structs against structure of arrays, scalar against NEON.

SIMD instruction sets

SIMD libraries and measured kernels

Branchless code and bit manipulation

8. Compilers and codegen

No -O level changes the target instruction set: without -march or -mcpu, every instruction emitted belongs to the default target ISA, so target flags come before any judgement of codegen.

Reading emitted code

  • Compiler Explorer - Shows how a source change alters the emitted instructions across compilers, versions and flags, with nothing installed.
  • What Every C Programmer Should Know About Undefined Behavior - Explains how the signed-overflow and aliasing rules let a trip count be known and a store loop become memset.
  • llvm-objdump - Reads the binary that shipped, with source lines and symbolised branch targets, rather than a recompiled snippet.
  • llvm-mca - Predicts loop throughput and port pressure from the scheduling model, and states it models neither front end nor caches.
  • llvm-exegesis - Measures instruction latency and throughput with counters, so the model llvm-mca predicts from is checked, not trusted.
  • Options That Control Optimization (GCC) - Lists what each -O level turns on, the inlining limits, and that -Ofast admits transforms invalid for conforming code.
  • There Are No Zero-cost Abstractions (CppCon 2019) - Shows with real codegen that an abstraction is free only when inlining and the ABI allow it, and the cost when either refuses.
  • Itanium C++ ABI - Fixes the rule that a non-trivial class goes by reference to a caller-made temporary, the cost a wrapped pointer pays.
  • How To Write Shared Libraries - States what PLT calls and interposition cost, and the visibility controls a library needs to inline its own exports.
  • LTO Overview (GCC Internals) - Defines whole-program LTO against partitioned WHOPR, and the LGEN, WPA and LTRANS stages that run -flto in parallel.
  • ThinLTO - Defines the thin link, summaries analysed whole-program then parallel backends, and the cache for incremental rebuilds.

Target flags and auto-vectorisation

Reproduce it: misc/benchmarks/08-autovectorization-aliasing, the vectoriser with and without restrict.

Profile-guided and post-link optimisation

9. Concurrency

Every cost below is a cache line moving between cores, so the ordering models and the measured line-transfer cost in the memory hierarchy section come first.

Memory models and atomics

Locks, contention and allocators

Reproduce it: misc/benchmarks/09-false-sharing, adjacent counters against padded counters across threads.

Lock-free structures and RCU

Thread pools and work stealing

10. NUMA and multi-socket

A page's node is decided at first touch, not when memory is allocated or a policy is set, and every vendor table below depends on the BIOS node mode of the machine it ran on.

NUMA and Linux memory placement

  • NUMA (Non-Uniform Memory Access): An Overview - The one account tying first touch, policy scope, zone reclaim and page movement together from the implementer's side.
  • What is NUMA? - Defines nodes, zonelists and the distance-ordered fallback that places an allocation once local memory runs out.
  • NUMA Memory Policy - The normative statement of policy scopes, every mode including weighted interleave, and the cpuset intersection rule.
  • Numa policy hit/miss statistics - Defines numa_hit, numa_miss and numa_foreign, the counters that show whether a policy put pages where it said.
  • numactl - Reference implementation of the policy API, prints the distance table and binds a binary that cannot be rebuilt.

Reproduce it: misc/benchmarks/10-first-touch, first touch of fresh pages against the second pass.

Topology and interconnects

Migration, balancing and measured effects

11. OS and I/O

A syscall's cost depends on the mitigation state, the governor and the idle state the core was in, three sysfs settings that change after boot, so each is recorded beside any number below.

Syscalls and asynchronous I/O

Reproduce it: misc/benchmarks/11-syscall-cost, the fixed cost of a kernel crossing across request sizes.

Scheduling, affinity and isolation

  • EEVDF Scheduler - Defines lag and virtual deadline, which the default class schedules by, and the slice request in sched_setattr.
  • The Linux Scheduler: a Decade of Wasted Cores - Proves cores sit idle while runnable threads queue, and gives the invariant checker that found the load-balancer bugs.
  • Control Group v2 - Defines cpu.max throttling, cpu.weight and the cpusets that bound affinity, the controls behind every container limit.
  • CPU Performance Scaling - Defines the governors, driver and boost switch that set a core's frequency, the sysfs state a measurement records.
  • CPU Isolation - Ties isolcpus, nohz_full, IRQ affinity, RCU offload and cpusets into one recipe, and lists the jitter it leaves.

Interrupts and kernel bypass

  • NAPI - Defines the polling, software coalescing, busy polling and IRQ suspension knobs that trade interrupts against latency.
  • DPDK Programmer's Guide - Defines the full bypass model, pinned poll-mode cores with no interrupts, that every kernel path is measured against.
  • The eXpress Data Path - Measures an in-kernel programmable path against DPDK and the stack per core, with the full configuration published.
  • Kernel vs. User-Level Networking: Don't Throw Out the Stack with the Interrupts - Separates direct and indirect NIC interrupt cost, measures the stack against bypass, and is where IRQ suspension began.
  • AF_XDP - Defines the socket and UMEM rings handing XDP frames to user space, and the zero-copy and need-wakeup modes.

Cache and bandwidth partitioning

12. Tail latency and production systems

A latency figure means nothing without its percentile, its load model and the way it was recorded.

Measuring the tail

  • The Tail at Scale - Shows why fan-out makes a rare slow server a common slow request, and names the techniques that tolerate variance.
  • Attack of the Killer Microseconds - Defines the stall band that out-of-order hardware cannot hide and a context switch cannot amortise.
  • How NOT to Measure Latency - Shows that a summary without a max discards the samples that define the tail, and closed-loop load never records them.
  • Coordinated Omission - The original definition of the recording error, with arithmetic for how far a reported percentile sits from the truth.
  • HdrHistogram - Keeps the whole distribution at fixed relative precision in constant time, so the far percentiles and max survive.

Reproduce it: misc/benchmarks/12-coordinated-omission, closed-loop against open-loop p99 under the same stalls.

Where jitter comes from

  • rt-tests - The reference wakeup-latency measurement for Linux, whose README states that an unloaded run proves nothing.
  • osnoise tracer - Counts the noise a spinning thread suffers and attributes each event to NMI, IRQ, softirq, thread or hardware.
  • Tales of the Tail - Derives the queueing-ideal tail and attributes the excess to scheduling, interrupt placement, power saving and NUMA.
  • Latency Implications of Virtual Memory - Measures with code the page-fault, TLB-shootdown and writeback stalls that memory mapping hides from the caller.
  • The KVM halt polling system - Defines the host-side polling after a vCPU halt that trades idle host CPU for guest wakeup time, unseen by the guest.

Load generation and production workloads

Mechanical sympathy

  • Inter Thread Latency - Measures with code the floor for handing a cache line between cores, which every queue and lock is built on.
  • Single Writer Principle - States the design rule that removes write contention outright, using a contended increment's cost as the argument.
  • Optimizing a Ring Buffer for Throughput - Adds cached indices to a single-producer single-consumer ring and shows with counters the coherence traffic removed.
  • LMAX Disruptor - Applies the single writer rule and cache-line padding to a ring buffer, with the queue comparison that motivated it.
  • Aeron - Carries the single writer and batching rules through a whole transport, the reference beyond one in-process queue.

13. Inference on CPU

A decode step at batch one reads every weight once for a few flops, so it runs at the memory system's rate, and a CPU figure compares with a GPU figure only at the same batch and precision.

GEMM and BLAS

Reproduce it: misc/benchmarks/13-sgemm-naive-vs-blas, naive GEMM, a hand microkernel and the vendor BLAS, plus int8 against float32 dot products.

Runtimes

  • oneDNN - Generates VNNI and AMX kernels at run time under PyTorch, TensorFlow and OpenVINO, and calls Compute Library on Arm.
  • ggml - Defines the quantized block formats and the per-ISA dot-product kernels over them that every llama.cpp type rests on.
  • llama.cpp - Where new quantization types, kernels and thread pools land first, each with the perplexity and speed table behind it.
  • ONNX Runtime MLAS - Holds the CPU provider's GEMM, int8 and int4 MatMul kernels, dispatched per ISA at run time from SSE to AMX and SME.
  • OpenVINO CPU Device - States precision defaults per ISA, the int8 path through oneDNN and the streams model that turns cores into throughput.

Quantization

Matrix extensions

Threading for inference

  • Thread management - Sets the physical-core default, the affinity it implies and the spin-wait controls that trade idle CPU for latency.
  • Performance Hints and Thread Scheduling - States the vendor defaults, one thread per core, SMT siblings off, core type by precision and one socket for latency.
  • Threadpool: take 2 - Defines the explicit thread pool with CPU masks, strict placement, priority and polling that ggml runs without OpenMP.
  • llama-bench - Defines the prompt and generation tests, repetitions and mean with deviation behind any comparable llama.cpp number.
  • Dual Epyc Genoa/Turin token generation performance bottlenecks - Traces poor decode scaling across sockets to remote NUMA access from weight placement, with numatop counts as evidence.

When CPU beats GPU

14. Hardware generations

The measurement articles below state no compiler, flags or run count, so each is kept for the structure it exposes and no figure from it is repeated.

Intel Xeon

AMD EPYC

Arm Neoverse server parts

Independent measurement across vendors

Reproduce it: misc/benchmarks/14-pcore-vs-ecore, the same three kernels on a performance core and an efficiency core.

15. Benchmarks

A score means what its suite's run rules say it means, so the rules come before the number.

Standard suites

Microbenchmark suites

Reproduce it: misc/benchmarks/15-stream-bandwidth, triad bandwidth by thread count against the vendor figure.

Methodology and what suites miss

16. Watchlist

Everything below is real but unproven: no item yet has all three of a written specification, a part you can buy, and a public measurement stating every one of the seven fields. Each line says what would promote it. Vendor multiples never qualify. Last checked 2026-09-15.

ISA extensions without a shipped server part

Parts without a public measurement

  • 6th Gen AMD EPYC Server CPUs - The Zen 6 server family, so far a press release with no shipped part, pending shipment and a public run against Zen 5.
  • Intel Xeon 6+ Processors - The E-core-only sockets after Sierra Forest, shipped with vendor multiples footnoted off the page, pending a public run.
  • NVIDIA Vera CPU - Custom Arm cores with statically partitioned SMT and no architecture document, pending a specification and a public run.
  • Arm Neoverse V3 Core Software Optimization Guide - Vendor timing tables for the core shipped in Graviton 5 and previewed in Cobalt 200, pending a public run on the core.

Memory and interconnect

Kernel paths and generated code

  • Extensible Scheduler Class - Lets a BPF program schedule at run time with safe fallback, once a run against the default scheduler states every field.
  • io_uring zero copy Rx - Lands payloads straight in user memory on header-splitting NICs, pending the implementer's epoll run naming every field.
  • T-MAC - Table-lookup kernels for low-bit weights, with a baseline stated but no frequency, compiler or flags, pending those.
  • Faster sorting algorithms discovered using deep reinforcement learning - Generated small sorts shipped in libc++, timed by CPU family with no model, compiler or flags stated, pending those.

What earns a place

Seven fields, and a number without all of them does not appear here:

1CPU model and microarchitecture
2core count used
3frequency, with turbo and SMT state
4compiler and flags
5workload
6baseline
7measurement method

Miss one and the number is dropped; if the entry rests on that number, it moves to the watchlist or goes.

An entry itself has to be the thing, not writing about the thing: the paper that first described a mechanism, the specification or manual that defines it, the repository the implementation lives in, or a report from whoever did the work with code and reproducible measurements. Summaries, tutorials, surveys, marketing pages, mirrors and repackagings do not qualify. Every URL points at the live canonical copy, and misc/scripts/check_links.py and misc/scripts/check_format.py prove it on every push and again weekly.

CONTRIBUTING.md has the rules in full.

License

MIT. Maintained by @usamahz.

Format inspired by the GPU-side list at wafer-ai/gpu-perf-engineering-resources.

arm
awesome-list
benchmarks
computer-architecture
cpu
edge-ai
inference
low-latency
optimization
performance
performance-engineering
simd
systems-programming
x86

usamahz/cpu-performance-engineering

A reading path for CPU performance engineering, from one instruction to production inference. Primary sources only, with a runnable benchmark for every section.

C

111

1 commits

updated Sep 27, 2026

See the code

See what people are saying

SourceMessageScoreDate

CPU Performance Engineering

1

Oct 1, 2026

README

CPU Performance Engineering

Links Quality Entries Benchmarks License Stars

Making a program fast on a modern CPU means knowing what the core does with each instruction, where the time actually goes, and how to prove a change helped. This is the reading that gets you there, in the order that makes the next piece legible.

Scope. x86 and Arm server parts, from one instruction through to serving a model on CPU. Not language runtimes, database internals, or anything above the socket.

Evidence. Primary sources only: the paper, the specification, the vendor manual, the repository, or a report by the person who did the work. Any number, anywhere in this repository, carries all seven fields set out in What earns a place, or it is not quoted.

Proof. Fourteen of the sections end in a benchmark under misc/benchmarks/: C source, the build line, the machine, the raw numbers and the analysis, all committed. Run them yourself.

Section 1 is a path through the rest; read it top to bottom before using the numbered sections as a reference.

Contents

1. Start here

Each entry assumes only the ones before it. The first and sixth are paid books; the third is a manual to open at the chapters its reason names.

  1. Computer Architecture: A Quantitative Approach, 7th Edition - Its pipelining appendix and memory chapters define the hazard, speculation and cache vocabulary the list assumes.
  2. Optimizing software in C++ - Maps C++ onto pipeline mechanisms and shows why a loop-carried dependency chain, not instruction count, paces a loop.
  3. Intel Optimization Reference Manual - Its opening chapters show how a shipping x86 core implements the textbook pipeline, each rule tied to a mechanism.
  4. What Every Programmer Should Know About Memory - Measures the step in cost per access at each cache boundary and the gap a prefetcher hides.
  5. Memory Barriers: a Hardware View for Software Hackers - Explains why a second core makes loads and stores reorder and what a barrier drains.
  6. Systems Performance: Enterprise and the Cloud, 2nd Edition - Puts the method before the tools: what to measure, in what order, and how benchmarks mislead.
  7. Roofline: An Insightful Visual Performance Model for Multicore Architectures - Places a loop from a byte count and a datasheet bandwidth alone, before any counter is read.
  8. A Top-Down Method for Performance Analysis and Counters Architecture - Defines the split of pipeline slots into front end, bad speculation, back end and retiring, the tree profilers report.
  9. Performance Analysis and Tuning on Modern CPUs - Walks from a noisy timing to counters to a named bottleneck, applying roofline and top-down to whole programs.
  10. What Has My Compiler Done for Me Lately? Unbolting the Compiler's Lid - Shows how to read emitted assembly against its source, so each mechanism is checked in a listing, not assumed.

Work the exercises in Performance Ninja alongside them; reading alone will not build the instinct.

Reproduce it: misc/benchmarks/04-cache-latency, the cost per dependent load stepping up at each cache boundary, the curve the fourth entry measures.

2. One instruction, end to end

Vendors name the same structures differently (Intel's decoded ICache is AMD's op cache, and AMD's macro-op is Arm's MOP), so the stage names below are generic.

Fetch and decode

Reproduce it: misc/benchmarks/02-branch-misprediction, the cost of a mispredicted branch, sorted against unsorted against branchless.

Rename and issue

Execute

Memory access and retire

3. Microarchitecture

Where a design paper, the vendor manual and a measurement disagree about a core, the measurement is the one to trust and re-run.

Limits of ILP and SMT

Branch prediction and speculation

Vendor estimates and measured tables

What the manuals leave out

Reproduce it: misc/benchmarks/03-latency-vs-throughput, one dependency chain against eight independent accumulators.

4. Memory hierarchy

Line size and page size are machine parameters, not constants, so every padding and alignment rule below is applied against the target's own values.

Cache geometry, replacement and misses in flight

Reproduce it: misc/benchmarks/04-cache-latency, dependent-load latency from L1 to DRAM, with and without TLB pressure.

TLBs, page walks and prefetchers

Store buffers, ordering and cache-line contention

Struct layout, software prefetch and page size

5. Measurement

Cloud instances often virtualise the hardware counters away and macOS runs no Linux perf, so perf stat has to show a non-zero cycles count before any counter entry below is trusted.

Method and the whole-system view

Counters, events and precise sampling

CPU profilers and flame graphs

Microbenchmarks that lie

Reproduce it: misc/benchmarks/05-measurement-pitfalls, dead-code elimination, run-to-run spread, cold against warm.

6. Models

A roofline is a bound built from measured roofs and counted bytes, so a point above a roof means a wrong roof or a wrong byte count, not fast code.

Roofline and the execution-cache-memory model

Reproduce it: misc/benchmarks/06-roofline, measured roofs and three kernels of rising arithmetic intensity.

Top-down analysis

Scaling laws

Queueing

7. Single-thread optimisation

A loop the compiler reports as vectorised can still run at scalar speed: a float reduction stays one serial chain until reassociation is permitted.

Data layout and loop transforms

Reproduce it: misc/benchmarks/07-aos-vs-soa-simd, array of structs against structure of arrays, scalar against NEON.

SIMD instruction sets

SIMD libraries and measured kernels

Branchless code and bit manipulation

8. Compilers and codegen

No -O level changes the target instruction set: without -march or -mcpu, every instruction emitted belongs to the default target ISA, so target flags come before any judgement of codegen.

Reading emitted code

  • Compiler Explorer - Shows how a source change alters the emitted instructions across compilers, versions and flags, with nothing installed.
  • What Every C Programmer Should Know About Undefined Behavior - Explains how the signed-overflow and aliasing rules let a trip count be known and a store loop become memset.
  • llvm-objdump - Reads the binary that shipped, with source lines and symbolised branch targets, rather than a recompiled snippet.
  • llvm-mca - Predicts loop throughput and port pressure from the scheduling model, and states it models neither front end nor caches.
  • llvm-exegesis - Measures instruction latency and throughput with counters, so the model llvm-mca predicts from is checked, not trusted.
  • Options That Control Optimization (GCC) - Lists what each -O level turns on, the inlining limits, and that -Ofast admits transforms invalid for conforming code.
  • There Are No Zero-cost Abstractions (CppCon 2019) - Shows with real codegen that an abstraction is free only when inlining and the ABI allow it, and the cost when either refuses.
  • Itanium C++ ABI - Fixes the rule that a non-trivial class goes by reference to a caller-made temporary, the cost a wrapped pointer pays.
  • How To Write Shared Libraries - States what PLT calls and interposition cost, and the visibility controls a library needs to inline its own exports.
  • LTO Overview (GCC Internals) - Defines whole-program LTO against partitioned WHOPR, and the LGEN, WPA and LTRANS stages that run -flto in parallel.
  • ThinLTO - Defines the thin link, summaries analysed whole-program then parallel backends, and the cache for incremental rebuilds.

Target flags and auto-vectorisation

Reproduce it: misc/benchmarks/08-autovectorization-aliasing, the vectoriser with and without restrict.

Profile-guided and post-link optimisation

9. Concurrency

Every cost below is a cache line moving between cores, so the ordering models and the measured line-transfer cost in the memory hierarchy section come first.

Memory models and atomics

Locks, contention and allocators

Reproduce it: misc/benchmarks/09-false-sharing, adjacent counters against padded counters across threads.

Lock-free structures and RCU

Thread pools and work stealing

10. NUMA and multi-socket

A page's node is decided at first touch, not when memory is allocated or a policy is set, and every vendor table below depends on the BIOS node mode of the machine it ran on.

NUMA and Linux memory placement

  • NUMA (Non-Uniform Memory Access): An Overview - The one account tying first touch, policy scope, zone reclaim and page movement together from the implementer's side.
  • What is NUMA? - Defines nodes, zonelists and the distance-ordered fallback that places an allocation once local memory runs out.
  • NUMA Memory Policy - The normative statement of policy scopes, every mode including weighted interleave, and the cpuset intersection rule.
  • Numa policy hit/miss statistics - Defines numa_hit, numa_miss and numa_foreign, the counters that show whether a policy put pages where it said.
  • numactl - Reference implementation of the policy API, prints the distance table and binds a binary that cannot be rebuilt.

Reproduce it: misc/benchmarks/10-first-touch, first touch of fresh pages against the second pass.

Topology and interconnects

Migration, balancing and measured effects

11. OS and I/O

A syscall's cost depends on the mitigation state, the governor and the idle state the core was in, three sysfs settings that change after boot, so each is recorded beside any number below.

Syscalls and asynchronous I/O

Reproduce it: misc/benchmarks/11-syscall-cost, the fixed cost of a kernel crossing across request sizes.

Scheduling, affinity and isolation

  • EEVDF Scheduler - Defines lag and virtual deadline, which the default class schedules by, and the slice request in sched_setattr.
  • The Linux Scheduler: a Decade of Wasted Cores - Proves cores sit idle while runnable threads queue, and gives the invariant checker that found the load-balancer bugs.
  • Control Group v2 - Defines cpu.max throttling, cpu.weight and the cpusets that bound affinity, the controls behind every container limit.
  • CPU Performance Scaling - Defines the governors, driver and boost switch that set a core's frequency, the sysfs state a measurement records.
  • CPU Isolation - Ties isolcpus, nohz_full, IRQ affinity, RCU offload and cpusets into one recipe, and lists the jitter it leaves.

Interrupts and kernel bypass

  • NAPI - Defines the polling, software coalescing, busy polling and IRQ suspension knobs that trade interrupts against latency.
  • DPDK Programmer's Guide - Defines the full bypass model, pinned poll-mode cores with no interrupts, that every kernel path is measured against.
  • The eXpress Data Path - Measures an in-kernel programmable path against DPDK and the stack per core, with the full configuration published.
  • Kernel vs. User-Level Networking: Don't Throw Out the Stack with the Interrupts - Separates direct and indirect NIC interrupt cost, measures the stack against bypass, and is where IRQ suspension began.
  • AF_XDP - Defines the socket and UMEM rings handing XDP frames to user space, and the zero-copy and need-wakeup modes.

Cache and bandwidth partitioning

12. Tail latency and production systems

A latency figure means nothing without its percentile, its load model and the way it was recorded.

Measuring the tail

  • The Tail at Scale - Shows why fan-out makes a rare slow server a common slow request, and names the techniques that tolerate variance.
  • Attack of the Killer Microseconds - Defines the stall band that out-of-order hardware cannot hide and a context switch cannot amortise.
  • How NOT to Measure Latency - Shows that a summary without a max discards the samples that define the tail, and closed-loop load never records them.
  • Coordinated Omission - The original definition of the recording error, with arithmetic for how far a reported percentile sits from the truth.
  • HdrHistogram - Keeps the whole distribution at fixed relative precision in constant time, so the far percentiles and max survive.

Reproduce it: misc/benchmarks/12-coordinated-omission, closed-loop against open-loop p99 under the same stalls.

Where jitter comes from

  • rt-tests - The reference wakeup-latency measurement for Linux, whose README states that an unloaded run proves nothing.
  • osnoise tracer - Counts the noise a spinning thread suffers and attributes each event to NMI, IRQ, softirq, thread or hardware.
  • Tales of the Tail - Derives the queueing-ideal tail and attributes the excess to scheduling, interrupt placement, power saving and NUMA.
  • Latency Implications of Virtual Memory - Measures with code the page-fault, TLB-shootdown and writeback stalls that memory mapping hides from the caller.
  • The KVM halt polling system - Defines the host-side polling after a vCPU halt that trades idle host CPU for guest wakeup time, unseen by the guest.

Load generation and production workloads

Mechanical sympathy

  • Inter Thread Latency - Measures with code the floor for handing a cache line between cores, which every queue and lock is built on.
  • Single Writer Principle - States the design rule that removes write contention outright, using a contended increment's cost as the argument.
  • Optimizing a Ring Buffer for Throughput - Adds cached indices to a single-producer single-consumer ring and shows with counters the coherence traffic removed.
  • LMAX Disruptor - Applies the single writer rule and cache-line padding to a ring buffer, with the queue comparison that motivated it.
  • Aeron - Carries the single writer and batching rules through a whole transport, the reference beyond one in-process queue.

13. Inference on CPU

A decode step at batch one reads every weight once for a few flops, so it runs at the memory system's rate, and a CPU figure compares with a GPU figure only at the same batch and precision.

GEMM and BLAS

Reproduce it: misc/benchmarks/13-sgemm-naive-vs-blas, naive GEMM, a hand microkernel and the vendor BLAS, plus int8 against float32 dot products.

Runtimes

  • oneDNN - Generates VNNI and AMX kernels at run time under PyTorch, TensorFlow and OpenVINO, and calls Compute Library on Arm.
  • ggml - Defines the quantized block formats and the per-ISA dot-product kernels over them that every llama.cpp type rests on.
  • llama.cpp - Where new quantization types, kernels and thread pools land first, each with the perplexity and speed table behind it.
  • ONNX Runtime MLAS - Holds the CPU provider's GEMM, int8 and int4 MatMul kernels, dispatched per ISA at run time from SSE to AMX and SME.
  • OpenVINO CPU Device - States precision defaults per ISA, the int8 path through oneDNN and the streams model that turns cores into throughput.

Quantization

Matrix extensions

Threading for inference

  • Thread management - Sets the physical-core default, the affinity it implies and the spin-wait controls that trade idle CPU for latency.
  • Performance Hints and Thread Scheduling - States the vendor defaults, one thread per core, SMT siblings off, core type by precision and one socket for latency.
  • Threadpool: take 2 - Defines the explicit thread pool with CPU masks, strict placement, priority and polling that ggml runs without OpenMP.
  • llama-bench - Defines the prompt and generation tests, repetitions and mean with deviation behind any comparable llama.cpp number.
  • Dual Epyc Genoa/Turin token generation performance bottlenecks - Traces poor decode scaling across sockets to remote NUMA access from weight placement, with numatop counts as evidence.

When CPU beats GPU

14. Hardware generations

The measurement articles below state no compiler, flags or run count, so each is kept for the structure it exposes and no figure from it is repeated.

Intel Xeon

AMD EPYC

Arm Neoverse server parts

Independent measurement across vendors

Reproduce it: misc/benchmarks/14-pcore-vs-ecore, the same three kernels on a performance core and an efficiency core.

15. Benchmarks

A score means what its suite's run rules say it means, so the rules come before the number.

Standard suites

Microbenchmark suites

Reproduce it: misc/benchmarks/15-stream-bandwidth, triad bandwidth by thread count against the vendor figure.

Methodology and what suites miss

16. Watchlist

Everything below is real but unproven: no item yet has all three of a written specification, a part you can buy, and a public measurement stating every one of the seven fields. Each line says what would promote it. Vendor multiples never qualify. Last checked 2026-09-15.

ISA extensions without a shipped server part

Parts without a public measurement

  • 6th Gen AMD EPYC Server CPUs - The Zen 6 server family, so far a press release with no shipped part, pending shipment and a public run against Zen 5.
  • Intel Xeon 6+ Processors - The E-core-only sockets after Sierra Forest, shipped with vendor multiples footnoted off the page, pending a public run.
  • NVIDIA Vera CPU - Custom Arm cores with statically partitioned SMT and no architecture document, pending a specification and a public run.
  • Arm Neoverse V3 Core Software Optimization Guide - Vendor timing tables for the core shipped in Graviton 5 and previewed in Cobalt 200, pending a public run on the core.

Memory and interconnect

Kernel paths and generated code

  • Extensible Scheduler Class - Lets a BPF program schedule at run time with safe fallback, once a run against the default scheduler states every field.
  • io_uring zero copy Rx - Lands payloads straight in user memory on header-splitting NICs, pending the implementer's epoll run naming every field.
  • T-MAC - Table-lookup kernels for low-bit weights, with a baseline stated but no frequency, compiler or flags, pending those.
  • Faster sorting algorithms discovered using deep reinforcement learning - Generated small sorts shipped in libc++, timed by CPU family with no model, compiler or flags stated, pending those.

What earns a place

Seven fields, and a number without all of them does not appear here:

1CPU model and microarchitecture
2core count used
3frequency, with turbo and SMT state
4compiler and flags
5workload
6baseline
7measurement method

Miss one and the number is dropped; if the entry rests on that number, it moves to the watchlist or goes.

An entry itself has to be the thing, not writing about the thing: the paper that first described a mechanism, the specification or manual that defines it, the repository the implementation lives in, or a report from whoever did the work with code and reproducible measurements. Summaries, tutorials, surveys, marketing pages, mirrors and repackagings do not qualify. Every URL points at the live canonical copy, and misc/scripts/check_links.py and misc/scripts/check_format.py prove it on every push and again weekly.

CONTRIBUTING.md has the rules in full.

License

MIT. Maintained by @usamahz.

Format inspired by the GPU-side list at wafer-ai/gpu-perf-engineering-resources.

arm
awesome-list
benchmarks
computer-architecture
cpu
edge-ai
inference
low-latency
optimization
performance
performance-engineering
simd
systems-programming
x86

Languages

C

61.9%

Shell

29.1%

Python

9.1%