Sizing and TCO

ch12 · The sizing model

Builds on ch09, ch10 and ch11.

The question

What does the whole chain produce, and how much of it would you defend?

Part I described a workload. Part II found the ceilings. Part III has turned both into machines. This chapter puts them together, arrives at a number, and then makes the number look at itself.

The material

The whole model in one graph

A web service and its data, on a fleet of Linux hosts — The whole sizing, before a price — dependency graph, showing only what feeds hosts the model recommendsinputderivedmeasuredceilingyou decidethroughput the fleet canactually reachannual growth factorcores busy at the busyhourlimit on working setagainst memorycache marginrequests in the systemcontentionlimit on utilisation,counting coordinationcores in the fleetcores per hostcrosstalklimit on disk fill athorizondisk margindisk per hosteffective utilisationlimit on utilisation withone host downmemory the service canuse, whole fleetfraction of the peakalready builthorizonhorizon periodshost counthosts in the fleethosts left when one dieshosts for memoryhosts for requestshosts for storagehosts the model recommendsshare of records touchedin a busy hourrequests in flight, ifnone waitedindex overheadinstalled diskthroughput if scaling werefreemean request rate athorizonone coreone hostone requestone yearhow much the queueing viewunderstated itos reservewhere adding hosts stopshelpingpeak request rate athorizonpeak request rate, day onepeak-to-mean ratiolimit on utilisation atthe busy hourqueueing marginmemory the service canuse, per hostram per hostraw disk needed at horizonraw bytes per stored byterecord compression ratioreplication factorresidence timescaling efficiencylimit on fraction of thefleet doing nothing usefulCPU time per requestservice timethroughput of one hostalonerecords held at horizonrecords held, day oneutilisationutilisation with one hostdownutilisation the model willadmit toutilisation, countingcoordinationtime spent queueingworking set at horizon

Every node in that sub-graph has appeared in a chapter:

Follow it left to right and there is nothing surprising in it. Sizing models are not clever. They are a dozen multiplications anybody could check, and the difficulty has never been the arithmetic. Appendix E lists every one of them, with the chapter that added it.

What the model recommends

OutputPoint estimate90% intervalUnit
hosts the model recommends5420 to 231host
hosts in the fleet54fixedhost
hosts for requests5013 to 193host
hosts for memory5414 to 194host
hosts for storage4916 to 153host

Source — web_service_sizing-reference · every input on a slider

The first row is the answer, evaluated at every input’s point estimate. It is what a competently built spreadsheet would give you. It is also what this book’s reference fleet was bought against, which is the second row, and the fleet’s provenance says so in as many words. The three rows below are the chains it was the largest of.

The number looks at itself

Distribution of hosts the model recommendshosts the model recommends — 100,000 samples90% interval 20 to 231 · median 68p5pointp9531993960.9% of samples run on to 1,679

The red line is where the point estimate falls. Everything else is the same model, the same chains and the same margins, with the inputs allowed to be as uncertain as the people who wrote them down are. Notice where the line sits: below the middle of the bars. The largest of three uncertain counts is usually larger than the largest of their three point estimates. So the spreadsheet’s answer is not merely uncertain. It is low.

And here is what that fleet does against the ceilings the last six chapters declared:

CeilingAt the planHeadroomAllowedLimitVerdictOver allowedOver limit
working set against memory0.7525%0.751.00ok49%35%
utilisation, counting coordination0.9530%0.701.00into the margin65%48%
disk fill at horizon0.6725%0.751.00ok44%29%
utilisation with one host down0.6630%0.701.00ok46%30%
utilisation at the busy hour0.6430%0.701.00ok46%30%
fraction of the fleet doing nothing useful0.3250%0.501.00ok1%0%

Source — web_service_sizing-reference · every input on a slider

At the point estimate, every ceiling but one is comfortable. Of course they are. The fleet was sized from those point estimates against those chains, so it satisfies them by construction. The one that is not comfortable is the one no chain was sized against: ch07’s utilisation counting coordination. The request chain ignores it, because the chain divides by the fleet’s processors as if each worked alone. The fleet is inside that margin before a single input has moved. A model that reported only the verdict column would be marking its own homework, and this one has marked it wrong in one place already.

The last two columns ask a different question. Read the utilisation at the busy hour row. Buy the fleet the arithmetic recommends, and across everything this model thinks could happen, it is over the knee at the busy hour in a substantial share of the futures. The working set has outgrown memory in more of them still. The last column says how often, and it is not an extreme-scenario number.

Nothing went wrong to produce that. Every input was defensible and every multiplication was correct. The result is a fleet that stands a real chance of not lasting its horizon under the knee. That is what sizing from point estimates does.

Here is all of Part III in one graph, with a slider on every input. Drag hosts in the fleet and watch every ceiling’s verdict at once. That is the decision this chapter is about.

The sizing model, complete. Nothing arrives in this chapter: it is every earlier one, together.

The same file, running. It is the file ch02 started, eleven chapters on.

So what is the answer?

There is not one. Part III has been building to that.

A sizing model does not produce a number. It produces a relationship between a number and a risk, and somebody has to choose a point on it. Problem 12.1 is that choice made explicitly: pick a breach probability you are willing to be accountable for, and ask the model what it costs in machines.

That is a different conversation from “how many hosts do we need”, and a better one, because it is answerable. Here is one other point on that curve: the same model, the same ceilings, with a fleet bought for the growth case rather than the expected one:

CeilingAt the planHeadroomAllowedLimitVerdictOver allowedOver limit
working set against memory0.3125%0.751.00ok13%6%
utilisation, counting coordination0.6730%0.701.00ok50%34%
disk fill at horizon0.2825%0.751.00ok8%3%
utilisation with one host down0.2730%0.701.00ok12%5%
utilisation at the busy hour0.2730%0.701.00ok12%5%
fraction of the fleet doing nothing useful0.6050%0.501.00into the margin90%0%

Source — web_service-sized_for_growth · every input on a slider

Every figure in the last two columns falls, most of them to a few per cent. The one that falls least is the coordination ceiling, because more hosts spend more of themselves on each other. That is ch07’s argument, arriving in a sizing table. What the bigger fleet costs is ch21’s table rather than this one. But the pair, what it costs beside how often it breaks, is the only form in which this decision can be handed to somebody.

Problem 12.2 is the shape of the trade, counted in hosts, because hosts are all Part III has. Removing risk costs machines, and not at a steady rate: the last few percentage points cost more hosts than the ones before them. Having that count is the difference between an argument and a preference. ch21 prices the same pair of fleets, and puts the price to the person whose decision it is.

The decision is an input

The number of hosts in the fleet is an input, not a derived node, and that is the detail a spreadsheet hides. It has a provenance and a source like any other. Sizing produces a recommendation. A person then decides, once, before the five years happen. Everything downstream, every dollar, every watt and every ceiling, follows from what they chose, not from what the model would recommend in hindsight.

Deriving it instead would make the ceilings tautologies. A fleet sized to sit under the knee sits under it in every sample, and the model would cheerfully report no chance at all of queueing. Keeping it an input lets the ceilings ask the only question worth asking: given what we bought, how often does the world break it?

Key takeaways

  • A sizing model is a dozen multiplications anybody could check. The difficulty has never been the arithmetic.

  • The spreadsheet’s answer is not merely uncertain. It is low. The largest of three uncertain counts is usually larger than the largest of their three point estimates.

  • A fleet sized from point estimates satisfies its ceilings by construction, and still breaks. Across the futures the model thinks plausible it is over the knee at the busy hour in a substantial share of them, with nothing having gone wrong.

  • A sizing model produces a relationship between a number and a risk, not a number. Somebody has to pick a point on it, and the only form the choice can be handed over in is what it costs beside how often it breaks.

  • The fleet is an input, because the decision is. Keeping the host count an input lets the ceilings ask the only question worth asking: given what was bought, how often does the world break it?

What this cannot tell you

Whether the structure is right. Everything above takes the chains as given and asks what the inputs are worth. A missing chain, a database’s connection limit, a cache’s eviction rate or the network between the hosts, is invisible from inside. Nothing in the output distinguishes a model that is complete from one that is not. That is ch20 · The missing node.

Whether the ceilings are where they were declared. All six were declared by somebody with a reason (ch11). The probabilities in the last two columns are exact statements about where the model’s samples fall relative to lines that are judgements.

Where the uncertainty comes from. The interval is wide, and this chapter has not said which input makes it wide. That is the only actionable question about a wide interval, and ch19 answers it. The answer will not surprise you if you read ch04 · Peak, mean and growth.

What any of it costs. Part III has sized a fleet and said nothing about money. Part V is cost, and it comes after sizing because it consumes sizing’s output, including, if anybody is careful, its uncertainty.

How any of these numbers were produced. The last two columns of every ceiling table have been appearing since ch06 without explanation. ch13 is the explanation, and it is next because this is the chapter where a number appeared that you cannot defend.

Problems

Three, in tests/the_sizing_model/. The first two have tests. The last does not, and says why.

12.1 — Size to a risk, not to a point estimate. Find the smallest fleet whose queueing ceiling is breached in at most some fraction of samples. Bisect rather than step, and turn the number of draws down while searching. A search nobody runs twice is a search nobody runs.

tests/the_sizing_model/stubs.py · hosts_for_riskyours to edit
def hosts_for_risk(risk_at: Callable[..., float], target: float) -> int:
    """Problem 12.1 - size to a risk, not to a point estimate.

    Return the smallest number of hosts for which the web service's queueing ceiling is breached
    in at most ``target`` of its samples.

    ``risk_at(hosts)`` asks the model: it works the web service through with that many hosts in
    the fleet and returns the share of its samples in which the queueing ceiling is breached.
    ``risk_at(hosts, samples=n)`` does the same over ``n`` draws instead of the full number.
    Search - the relationship is monotonic, so a bisection is a few lines and is much faster
    than stepping, which matters because each call samples the whole graph.

    Reduce the number of draws while you search and put it back for the final answer. A search
    that takes a minute is a search nobody runs twice, and the precision you need to compare
    candidates is much lower than the precision you need to report one.

    This is what sizing is. ch12's point estimate recommends a number; this asks the model a
    question somebody can be accountable for.
    """
    raise NotImplementedError("problem 12.1")

The same check at a desk: python3 -m pytest tests/the_sizing_model/test_problem_1_risk.py -m problem

12.2 — What a percentage point of risk costs, in hosts. Count the hosts between the fleets two risk targets need, then look at the shape as the target tightens. The last few points cost more machines than the ones before them, and knowing how many is the difference between an argument and a preference. Part III has no prices, so the answer is a count of hosts; ch21 prices the same pair.

tests/the_sizing_model/stubs.py · hosts_for_riskyours to edit
def hosts_for_risk(risk_at: Callable[..., float], target: float) -> int:
    """Problem 12.1 - size to a risk, not to a point estimate.

    Return the smallest number of hosts for which the web service's queueing ceiling is breached
    in at most ``target`` of its samples.

    ``risk_at(hosts)`` asks the model: it works the web service through with that many hosts in
    the fleet and returns the share of its samples in which the queueing ceiling is breached.
    ``risk_at(hosts, samples=n)`` does the same over ``n`` draws instead of the full number.
    Search - the relationship is monotonic, so a bisection is a few lines and is much faster
    than stepping, which matters because each call samples the whole graph.

    Reduce the number of draws while you search and put it back for the final answer. A search
    that takes a minute is a search nobody runs twice, and the precision you need to compare
    candidates is much lower than the precision you need to report one.

    This is what sizing is. ch12's point estimate recommends a number; this asks the model a
    question somebody can be accountable for.
    """
    raise NotImplementedError("problem 12.1")
tests/the_sizing_model/stubs.py · cost_of_certaintyyours to edit
def cost_of_certainty(risk_at: Callable[..., float], from_risk: float, to_risk: float) -> int:
    """Problem 12.2 - what a percentage point of risk costs, in hosts.

    Return the additional hosts between the fleet your ``hosts_for_risk`` finds for ``from_risk``
    and the one it finds for ``to_risk``, asking the same ``risk_at`` both times. Moving to a
    smaller risk costs hosts; moving the other way gives them back, so the sign matters.

    Then look at the shape of the answer as ``to_risk`` falls. Removing risk does not cost the same
    at every point: the last few percentage points cost more hosts than the ones before them, and
    by more than a little. Part III has said nothing about money, so the answer is counted in
    machines. ch21 prices the same pair of fleets, and that price is the argument it has to put to
    the person whose decision it is. The count is where the argument starts.
    """
    raise NotImplementedError("problem 12.2")

The same check at a desk: python3 -m pytest tests/the_sizing_model/test_problem_2_cost_of_certainty.py -m problem

12.3 — Your own model, as far as it goes. No test: it is your chain, and the marks on it are judgements.

Take the quantities from your workload and assemble them into a chain that ends in a count of machines. Not in a file, unless you want to; on paper is fine. The point is to get from what arrives to what you buy without skipping a step.

Then find the two things that make it a sizing model rather than a cost model: a constant somebody measured on a particular version of a particular piece of software, and a limit your system runs into. Mark each one.

Then turn to the fleet you have, rather than the one the chain recommends. Which of its limits gives first as the load grows, how often would you expect that to happen over the horizon, and who accepted that? This chapter’s argument is that somebody did, whether or not they knew it.

A good answer reaches a number, has at least one mark on it, and names the risk the bought fleet accepts and the person who accepted it. A chain with no marks is a cost model. Either your system genuinely has no measured constants and no ceilings, which is rare, or you have not found them yet, which is the more likely reading and the more expensive one. A risk nobody accepted is the commoner finding, and it is the one to take to whoever signs for the fleet.

Where to go next

ch13 is where the last two columns came from. It sits here rather than at the front of the book for the reason this chapter has just demonstrated: the method is no use to you until you have a number you cannot defend, and can feel that you cannot defend it.

Appendix E is this model in full, through every output the toolkit produces.