Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

ch13 · Writing a Paged Attention Kernel in Triton [DRAFT]

The problem

ch08’s PyTorch gather materialises far too much and launches too many kernels. To fix it we have to write the kernel.

[To write: open with the measurement from the previous chapter that is unacceptable. Show it, do not assert it.]

The idea

[To write: develop the mechanism from first principles, with the arithmetic that predicts the result.]

The build

[To write: incremental implementation. Quote code from llmserve/ with {literalinclude} and :start-at: / :end-at: anchors — never paste it inline.]

The measurement

Kernel-level microbenchmark plus the end-to-end scorecard row.

[To write: re-run the harness, append the scorecard row, and explain in one paragraph why the number moved. Read the figures from bench/results/ — never type a number into prose (PLAN.md §6.3).]

The cost

Optional by design (PLAN.md §13, decision 4). Raises the hardware floor and is the least portable code in the book. Nothing after this chapter assumes it.

[To write: what did this make worse? Complexity, latency variance, quality, memory, operational burden. This section is mandatory — the chapter is not done without it (PLAN.md §12.3).]

Key takeaways

Looking ahead

[To write: the transition that sets up the next chapter’s problem.]

Further reading

[To write: primary sources, cited with MyST citation syntax against references.bib.]