Understand every cycle · Prove every speedup

A CPU performance engineering brain for any codebase.

Making a program fast on a modern CPU means knowing what the core does with each instruction, where the time actually goes, and how to prove a change helped. This is the reading that gets you there, in the order that makes the next piece legible.

sections
16
primary sources
304
benchmarks
14
  1. 01

    Pasted output

    perf stat, top-down, compiler remarks, assembly

  2. 02

    Metrics

    computed as perf computes them, each with its formula

  3. 03

    Entries and reasons

    the list's own sources, in reading order

  4. 04

    Source passages

    from the linked papers and manuals, with page numbers

  5. 05

    Benchmark

    the matching committed experiment

  6. 06

    Editorial record

    what was left out, and the rule it failed

01

Scope

x86 and Arm server parts, from one instruction through to serving a model on CPU. Not language runtimes, database internals, or anything above the socket.

02

Evidence

Primary sources only: the paper, the specification, the vendor manual, the repository, or a report by the person who did the work. Any number, anywhere in this repository, carries all seven fields set out in What earns a place, or it is not quoted.

03

Proof

Fourteen of the sections end in a benchmark under misc/benchmarks/: C source, the build line, the machine, the raw numbers and the analysis, all committed. Run them yourself.

04

Plug it in

The MCP server in misc/mcp/ turns this list into a CPU performance brain inside any AI client: the reading order, every rejected candidate with the rule it failed, every benchmark, and an index of the linked sources built on the reader's own machine. Point it at real work and the answers quote those sources and cite them. With uv installed, one command adds it.

The reading path

16 sections, read down and then out

All sections

§1

Start here

Ten numbered items, read top to bottom, each assuming only the ones before it.

  1. 01Start here

§2–6

Down

The core, the memory hierarchy, measurement, models.

  1. 02One instruction, end to end
  2. 03Microarchitecture
  3. 04Memory hierarchy
  4. 05Measurement
  5. 06Models

§7–15

Out

One thread, the compiler, many threads, NUMA, the kernel boundary, tail latency, CPU inference, the parts themselves, the benchmark suites.

  1. 07Single-thread optimisation
  2. 08Compilers and codegen
  3. 09Concurrency
  4. 10NUMA and multi-socket
  5. 11OS and I/O
  6. 12Tail latency and production systems
  7. 13Inference on CPU
  8. 14Hardware generations
  9. 15Benchmarks

§16

Watchlist

Dated, for things whose evidence is still moving.

  1. 16Watchlist

MCP · cpu-perf

Why does my multithreaded counter stop scaling past two threads?

The MCP server in misc/mcp/ turns this list into a CPU performance brain inside any AI client: the reading order, every rejected candidate with the rule it failed, every benchmark, and an index of the linked sources built on the reader's own machine. Point it at real work and the answers quote those sources and cite them. With uv installed, one command adds it.

Claude Desktop, Cursor and VS Code take a few lines of config, given in Connect it.

How to connect it →
Claude Code
claude mcp add --scope user cpu-perf -- uvx cpu-perf
Codex
codex mcp add cpu-perf -- uvx cpu-perf

Featured benchmark · committed data

Dependent-load latency across the memory hierarchy

Full analysis

Dependent-load latency steps at each cache level, and page-random access adds TLB cost

The measurement standard

Seven fields, or the number is not quoted

  1. 1CPU model and microarchitecture
  2. 2Core count used
  3. 3Frequency, with turbo and SMT state
  4. 4Compiler and flags
  5. 5Workload
  6. 6Baseline
  7. 7Measurement method
Evidence rules →

Watchlist · real but unproven

What would promote it

Open the watchlist →

Run it, check it, extend it.

Re-measure on another machine, report a dead link, or propose a primary source.