Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

ch14 · Quantisation for Serving [DRAFT]

The problem

Per ch03, decode reads all the weights every token. The cheapest way to read fewer bytes is to store fewer bytes.

[To write: open with the measurement from the previous chapter that is unacceptable. Show it, do not assert it.]

The idea

[To write: develop the mechanism from first principles, with the arithmetic that predicts the result.]

The build

[To write: incremental implementation. Quote code from llmserve/ with {literalinclude} and :start-at: / :end-at: anchors — never paste it inline.]

The measurement

Three axes at once: quality, speed, memory. A table the reader can use to choose, not a single headline number.

[To write: re-run the harness, append the scorecard row, and explain in one paragraph why the number moved. Read the figures from bench/results/ — never type a number into prose (PLAN.md §6.3).]

The cost

The first change that makes output worse. Every quantisation claim needs a quality number beside it, and ‘negligible’ is not a number.

[To write: what did this make worse? Complexity, latency variance, quality, memory, operational burden. This section is mandatory — the chapter is not done without it (PLAN.md §12.3).]

Key takeaways

Looking ahead

[To write: the transition that sets up the next chapter’s problem.]

Further reading

[To write: primary sources, cited with MyST citation syntax against references.bib.]