Understand every cycle · Prove every speedup
A CPU performance engineering brain for any codebase.
Making a program fast on a modern CPU means knowing what the core does with each instruction, where the time actually goes, and how to prove a change helped. This is the reading that gets you there, in the order that makes the next piece legible.
- sections
- 16
- primary sources
- 304
- benchmarks
- 14
- 01
Pasted output
perf stat, top-down, compiler remarks, assembly
- 02
Metrics
computed as perf computes them, each with its formula
- 03
Entries and reasons
the list's own sources, in reading order
- 04
Source passages
from the linked papers and manuals, with page numbers
- 05
Benchmark
the matching committed experiment
- 06
Editorial record
what was left out, and the rule it failed
Scope
x86 and Arm server parts, from one instruction through to serving a model on CPU. Not language runtimes, database internals, or anything above the socket.
Evidence
Primary sources only: the paper, the specification, the vendor manual, the repository, or a report by the person who did the work. Any number, anywhere in this repository, carries all seven fields set out in What earns a place, or it is not quoted.
Proof
Fourteen of the sections end in a benchmark under misc/benchmarks/: C source, the build line, the machine, the raw numbers and the analysis, all committed. Run them yourself.
Plug it in
The MCP server in misc/mcp/ turns this list into a CPU performance brain inside any AI client: the reading order, every rejected candidate with the rule it failed, every benchmark, and an index of the linked sources built on the reader's own machine. Point it at real work and the answers quote those sources and cite them. With uv installed, one command adds it.
The reading path
16 sections, read down and then out
§1
Start here
Ten numbered items, read top to bottom, each assuming only the ones before it.
§2–6
Down
The core, the memory hierarchy, measurement, models.
§7–15
Out
One thread, the compiler, many threads, NUMA, the kernel boundary, tail latency, CPU inference, the parts themselves, the benchmark suites.
MCP · cpu-perf
Why does my multithreaded counter stop scaling past two threads?
The MCP server in misc/mcp/ turns this list into a CPU performance brain inside any AI client: the reading order, every rejected candidate with the rule it failed, every benchmark, and an index of the linked sources built on the reader's own machine. Point it at real work and the answers quote those sources and cite them. With uv installed, one command adds it.
Claude Desktop, Cursor and VS Code take a few lines of config, given in Connect it.
How to connect it →claude mcp add --scope user cpu-perf -- uvx cpu-perf
codex mcp add cpu-perf -- uvx cpu-perf
Featured benchmark · committed data
Dependent-load latency across the memory hierarchy
Full analysisDependent-load latency steps at each cache level, and page-random access adds TLB cost
Line mode: 128-byte nodes packed 128 bytes apartPage mode: one 128-byte node per page
Data behind this chart
| Series | working set (line mode) or span (page mode) | Value | Unit | RESULT key |
|---|---|---|---|---|
| Line mode: 128-byte nodes packed 128 bytes apart | 4 KiB | 0.665 | ns/load | line_4096_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 8 KiB | 0.665 | ns/load | line_8192_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 16 KiB | 0.665 | ns/load | line_16384_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 32 KiB | 0.665 | ns/load | line_32768_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 64 KiB | 0.665 | ns/load | line_65536_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 128 KiB | 0.666 | ns/load | line_131072_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 256 KiB | 6.167 | ns/load | line_262144_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 512 KiB | 5.927 | ns/load | line_524288_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 1 MiB | 5.947 | ns/load | line_1048576_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 2 MiB | 6.048 | ns/load | line_2097152_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 4 MiB | 7.224 | ns/load | line_4194304_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 8 MiB | 7.974 | ns/load | line_8388608_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 16 MiB | 16.98 | ns/load | line_16777216_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 32 MiB | 60.04 | ns/load | line_33554432_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 64 MiB | 112.2 | ns/load | line_67108864_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 128 MiB | 117.3 | ns/load | line_134217728_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 256 MiB | 120 | ns/load | line_268435456_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 512 MiB | 121.5 | ns/load | line_536870912_ns |
| Line mode: 128-byte nodes packed 128 bytes apart | 1 GiB | 123.7 | ns/load | line_1073741824_ns |
| Page mode: one 128-byte node per page | 512 KiB | 0.665 | ns/load | page_524288_ns |
| Page mode: one 128-byte node per page | 1 MiB | 0.665 | ns/load | page_1048576_ns |
| Page mode: one 128-byte node per page | 2 MiB | 0.665 | ns/load | page_2097152_ns |
| Page mode: one 128-byte node per page | 4 MiB | 2.01 | ns/load | page_4194304_ns |
| Page mode: one 128-byte node per page | 8 MiB | 2.019 | ns/load | page_8388608_ns |
| Page mode: one 128-byte node per page | 16 MiB | 2.105 | ns/load | page_16777216_ns |
| Page mode: one 128-byte node per page | 32 MiB | 6.115 | ns/load | page_33554432_ns |
| Page mode: one 128-byte node per page | 64 MiB | 14.03 | ns/load | page_67108864_ns |
| Page mode: one 128-byte node per page | 128 MiB | 14.74 | ns/load | page_134217728_ns |
| Page mode: one 128-byte node per page | 256 MiB | 15.03 | ns/load | page_268435456_ns |
| Page mode: one 128-byte node per page | 512 MiB | 15.05 | ns/load | page_536870912_ns |
| Page mode: one 128-byte node per page | 1 GiB | 15.76 | ns/load | page_1073741824_ns |
| Reference line | Value | Source |
|---|---|---|
| L1d | 128 KiB | raw.txt header, P-core L1d |
| L2 | 16 MiB | raw.txt header, P-core L1d |
- 1 CPU model and microarchitecture
- Apple M4 Pro, arm64; Apple publishes no microarchitecture name. P-core: 128 KiB L1d, 16 MiB L2 shared by a cluster of 5 cores; 128-byte cache lines, 16 KiB pages, 24 GiB unified memory (all from
sysctl, printed at the top ofresults/raw.txt). - 2 Core count used
- one thread at default QoS, which macOS schedules on a P-core; nothing is pinned because macOS has no affinity API. Machine condition: load average at start 8.01 7.60 7.85 (the
load average at startline inresults/raw.txt), so other processes belonging to the user were running on the 14 logical cores; the benchmark is single-threaded, so it competed for a core only where the table's cv says so. - 3 Frequency, with turbo and SMT state
- 4.49 GHz from
common/clock_estimate(dependent one-cycle add chain, min of 7 runs), DVFS on, no SMT on this part. - 4 Compiler and flags
- Apple clang 17.0.0 (clang-1700.4.4.1),
cc -std=c11 -O2 -Wall -Wextra -o bench bench.c. - 5 Workload
- single-threaded pointer chase over a random single-cycle permutation of 128-byte nodes, working sets 4 KiB to 1 GiB doubling, in line placement (128-byte stride) and page placement (one node per 16 KiB page, line offsets balanced across the page).
- 6 Baseline
- the 32 KiB line-mode row (an L1 hit) and the 1 MiB line-mode row (an L2 hit) for the level ratios, and for each page-mode span the line-mode working set with the same line count.
- 7 Measurement method
now_ns()(CLOCK_MONOTONIC_RAW) around 6710886 dependent loads; 10 runs after one discarded warmup run; min and cv reported; cycles are ns times the estimated clock.
The measurement standard
Seven fields, or the number is not quoted
- 1CPU model and microarchitecture
- 2Core count used
- 3Frequency, with turbo and SMT state
- 4Compiler and flags
- 5Workload
- 6Baseline
- 7Measurement method
Watchlist · real but unproven
What would promote it
- ISA extensions without a shipped server part4 items
- Parts without a public measurement4 items
- Memory and interconnect3 items
- Kernel paths and generated code4 items
Run it, check it, extend it.
Re-measure on another machine, report a dead link, or propose a primary source.