ch21 · A TCO for a finance audience
The question
How do you present an interval to somebody who has asked you for a number?
Every chapter so far has been about getting the answer right. This one is about the twenty minutes in which it is either used or ignored. Those are not the same skill. A model that nobody acts on has the same value as a model that was never built.
The material
The two things the room wants
The instinct on the engineering side is to protect the interval. It was hard won. It is the most honest thing in the document. Collapsing it to a single figure feels like throwing away the work.
The instinct on the finance side is to get one number into a spreadsheet. That is not laziness. A budget line is a single number by construction. A commitment is a single number. The question “can we afford this” cannot be answered by a range until someone, somewhere, picks a point in it.
Both instincts are right. The argument between them is usually conducted as though only one of them can be. No compromise is needed, because the two sides want different things. Finance wants a number to commit to. Engineering wants to say what the number hides. Those two fit together.
The number, and the sentence
So the deliverable is a pair.
One number, and one sentence saying what it leaves out.
Not a range. Not a range with a recommendation attached. A number, chosen deliberately and named as a choice, and then a single sentence that says the thing the number cannot.
Which number you pick is a decision. You can defend any of these three out loud:
The median. Half the futures cost more. Honest, easy to explain, and the one most people mean when they say “the estimate”. It is also the number that will be wrong half the time in the direction that hurts.
A high percentile. The figure you would be comfortable committing to. Useful when overrunning is expensive and underspending is not. Its failure mode is quiet. You will be held to it, the money will be allocated, and when the cost lands lower nobody will thank you for the accuracy.
A round number above the median. Not as unprincipled as it sounds. Rounding to a precision the model can support is more honest than quoting a figure to the dollar. ch14 gives the arithmetic for what precision that is: the run-to-run wobble has to be below the digit you are prepared to defend.
Here is the shape all three are chosen from. It is the same picture ch01 opened with, now as something to pick a number off rather than something to be alarmed by:
The red line is the point estimate, and it is not the middle. Whichever of the three you choose, choose it off that chart and say which one it was.
The number must not arrive without the sentence. The sentence is all of the engineering position, and there is only room for one. So it has to name something specific: a percentile, an omission, an assumption the total rests on. “There is some uncertainty” names nothing, and it will be heard as “no”.
Problem 21.2 is that pair, and it is graded on the sentence.
A decision, not an interval
The single most effective change to a sizing document is this. Stop presenting one design with an interval. Start presenting two designs with a price.
| Output | Reference scenario | Sized for the growth we might get, not the growth we expect |
|---|---|---|
| hosts the model recommends | 54 20 to 230 | 54 20 to 230 |
| hosts in the fleet | 54 | 131 |
| five-year total cost of ownership | $2,002,083 $1,508,230 to $2,923,724 | $3,646,206 $2,781,458 to $5,393,738 |
| cost per million requests | $1.82 $0.55 to $5.98 | $3.32 $1.02 to $10.96 |
| cost per stored TB per month | $839.64 $305.33 to $2,108 | $1,529 $561.50 to $3,859 |
| capex | $421,214 $272,130 to $658,773 | $1,021,835 $660,167 to $1,598,135 |
| annual opex | $316,174 $227,660 to $484,257 | $524,874 $386,622 to $822,825 |
| annual energy | 206,269 159,636 to 269,269 | 500,394 387,264 to 653,228 |
| utilisation at the busy hour | 0.644 0.156 to 2.48 | 0.265 0.0643 to 1.02 |
| utilisation with one host down | 0.656 0.159 to 2.53 | 0.267 0.0648 to 1.03 |
| working set against memory | 0.746 0.194 to 2.67 | 0.307 0.0798 to 1.10 |
| disk fill at horizon | 0.670 0.213 to 2.12 | 0.276 0.0879 to 0.874 |
| fraction of the fleet doing nothing useful | 0.321 0.213 to 0.457 | 0.604 0.464 to 0.753 |
| utilisation, counting coordination | 0.948 0.233 to 3.81 | 0.671 0.163 to 2.90 |
| utilisation | 0.644 0.156 to 2.48 | 0.265 0.0643 to 1.02 |
| residence time | 0.0367 0.0128 to 0.857 | 0.0178 0.0110 to 0.619 |
| time spent queueing | 0.0236 0.0021 to 0.839 | 0.0047 0.0008 to 0.606 |
| requests in the system | 1,562 160 to 107,336 | 757 144 to 107,336 |
| requests in flight, if none waited | 556 135 to 2,147 | 556 135 to 2,147 |
| how much the queueing view understated it | 1.47 1.27 to 1.84 | 2.53 1.86 to 4.05 |
| fraction of the peak already built | 0.328 0.186 to 0.584 | 0.797 0.450 to 1.42 |
| working set against memory — over its limit | 35% | 6% |
| utilisation, counting coordination — over its limit | 48% | 34% |
| disk fill at horizon — over its limit | 29% | 3% |
| utilisation with one host down — over its limit | 30% | 5% |
| utilisation at the busy hour — over its limit | 30% | 5% |
| fraction of the fleet doing nothing useful — over its limit | 0% | 0% |
Source — web_service-reference and web_service-sized_for_growth · every input on a slider
The table has two columns. The left one buys what the model recommends at the point estimate. The right one buys the same fleet sized for the growth we might get rather than the growth we expect. That is the decision of ch11, taken deliberately instead of by default.
Read across the rows and the conversation changes shape. The question is no longer “is this estimate right”, which nobody in the room can answer. It is “is the difference between these two columns worth the difference in the last two rows”. That is the sort of question the people being asked are good at.
An interval is a statement about the world. A decision table is a statement about what you can buy. Only the second can be acted on by someone who cannot change the world but can sign for the extra machines.
Pricing a risk
Lead with the last two rows of that table, because they are the only ones with a consequence in them.
| Ceiling | At the plan | Headroom | Allowed | Limit | Verdict | Over allowed | Over limit |
|---|---|---|---|---|---|---|---|
| working set against memory | 0.75 | 25% | 0.75 | 1.00 | ok | 49% | 35% |
| utilisation, counting coordination | 0.95 | 30% | 0.70 | 1.00 | into the margin | 65% | 48% |
| disk fill at horizon | 0.67 | 25% | 0.75 | 1.00 | ok | 44% | 29% |
| utilisation with one host down | 0.66 | 30% | 0.70 | 1.00 | ok | 47% | 30% |
| utilisation at the busy hour | 0.64 | 30% | 0.70 | 1.00 | ok | 46% | 30% |
| fraction of the fleet doing nothing useful | 0.32 | 50% | 0.50 | 1.00 | ok | 1% | 0% |
Source — web_service-reference · every input on a slider
Over limit is how often, across the sampled futures, the fleet is asked to do something it cannot. The busy-hour row is a plain English sentence: this is how often we buy the fleet and it cannot serve the busy hour we said it would.
That sentence is worth more than any amount of argument about the growth rate. Nobody in the room has an opinion about a lognormal. Everybody in the room has an opinion about the service being slow on its busiest day.
The second scenario reduces that number, and the table says by how much and what it costs. Put the two together and you have the only sentence in the document that is a recommendation: this much additional capital buys this much less chance of that happening. If the answer is obviously yes, the meeting is over. If it is obviously no, the meeting is also over, and you have the decision in writing rather than in somebody’s memory.
Where each number came from
Then comes the appendix nobody asks for until they do.
| Input | Provenance | Source | |
|---|---|---|---|
| ○ | annual growth factor | assumption | ch04 — lognormal because growth compounds and cannot be negative. The p10/p90 say: surprised below 12% a year, surprised above 60%. One factor for requests and for records, because users drive both. |
| ○ | cache margin | assumption | ch08 — how much of the fleet’s memory is kept free of the working set, for its own daily swing and for the process heaps. Declared once and used by the chain and by the ceiling |
| ○ | contention | assumption | ch07 — fitted from two measurements, where two exist. Triangular because a fit gives a central value and a range rather than a shape, and its maximum is the claim to distrust: a serial fraction can always be worse than the one you measured |
| ◐ | cores per host | vendor claim | spec sheet: physical cores. A hyperthread is not a core, and a sheet that counts threads doubles this number without doubling the work a host does (ch07) |
| ○ | crosstalk | assumption | ch07 — fitted, and the harder of the two to fit. Lognormal because it spans an order of magnitude and cannot be negative, and because a shape should not assert a hard upper bound on a quantity this weakly fitted |
| ○ | disk margin | assumption | ch11 — room to re-replicate a dead host’s records onto the survivors, plus what a filesystem needs to keep allocating well. Declared once and used by the chain and by the ceiling |
| ◐ | disk per host | vendor claim | spec sheet: one local drive. Decimal TB, not TiB — appendix D, and it is a 10% difference |
| ○ | electricity price | assumption | all-in delivered rate including transmission. Lognormal: it cannot go negative and its history is multiplicative |
| ○ | fully loaded salary | assumption | salary, employer costs, tooling and overhead. Lognormal because pay is right-skewed and cannot go negative; the p90 is a senior engineer in an expensive city |
| ○ | horizon | assumption | the refresh cycle this fleet is bought against |
| ◐ | host power | vendor claim | typical draw under load, per host as configured. Triangular, and one of the few inputs here whose bounds are physical rather than editorial: a host cannot draw less than it idles at, or more than its supply will give it |
| ◐ | host price | vendor claim | chassis, CPU, memory, drives and boot media, as configured. Lognormal like any price, and wide because a host is a configuration rather than a commodity |
| ○ | hosts in the fleet | assumption | the sizing decision, taken the way it is usually taken: hosts_recommended evaluated at every input’s point estimate. Change this number and watch the ceilings move — that is the exercise of ch12 |
| ○ | share of records touched in a busy hour | assumption | ch08 — the working set as a share of everything held. Triangular; nobody measures this and everybody has an opinion, and the maximum is a service whose users all look at the same week’s data |
| ● | hours per year | fact | by definition, 365.25 x 24. The quarter-day is worth about a fifth of a per cent over five years — less than this model’s other errors, and free to get right |
| ○ | index overhead | assumption | indexes, the write-ahead log and journals as a multiplier on stored bytes. Triangular, and the bounds are for a service with a few indexes per table — a search-heavy one is off the top of this range |
| ◐ | licence per core | vendor claim | the platform software’s per-core licence, as quoted. Lognormal: a price, and a wide one, because it is the line item most often negotiated. It is what makes the core count a cost as well as a capacity (ch17) |
| ◐ | licence per host | vendor claim | the price list: per core, with no per-host charge. A quote that licenses per host instead puts its figure here and zero in licence_per_core, and the two totals are then compared like for like |
| ○ | one-off cost of moving to this design | assumption | nothing to migrate: the platform the plan already runs on, on the hosts already quoted. A challenger’s scenario overrides this with the team’s own estimate of the move, marked as what it is |
| ○ | network price per host | assumption | switch ports, optics and cabling, amortised per host. Triangular rather than lognormal, although it is a price: it is a bill of materials divided by a host count somebody chose, so the bounds are the plausible designs rather than a market |
| ● | one core | fact | definition |
| ● | one host | fact | definition |
| ● | one request | fact | definition |
| ● | one year | fact | definition |
| ○ | os reserve | assumption | the share of memory the kernel, the agents and the page cache floor keep before the service sees any. Triangular: a floor, a usual figure, and a host with too many agents on it |
| ○ | peak request rate, day one | assumption | ch04 — the busy hour, not the daily mean. Triangular because this is an engineer’s min/likely/max and pretending to more shape than that would be invention. |
| ○ | peak-to-mean ratio | assumption | ch04 — the busy hour against the daily mean. Triangular, and it is a property of your traffic that belongs to the estate target: nothing here can measure it |
| ○ | PUE | assumption | ch16 — facility overhead. A multiplier on IT load, and the single number a colocation contract is most likely to disagree with you about. Triangular: the minimum is a good building, the maximum is a poor one, and below one is impossible |
| ○ | queueing margin | assumption | ch06 — how far under the knee the fleet is sized to run at the busy hour. Declared once, here, and used by the sizing chain and by the ceiling that checks it, so the two cannot drift apart |
| ◐ | ram per host | vendor claim | spec sheet: the modules fitted. The sheet says 64 GB and means GiB — appendix D — and the operating system will report less than either, which is os_reserve’s job |
| ○ | replication factor | assumption | three copies of every record, so that a host can die and take its disks with it. A different durability scheme substitutes its own factor here and the rest of the model is unchanged, which is the point of it being a node |
| ● | seconds per year | fact | by definition, 365.25 x 86,400 |
| ○ | CPU time per request | assumption | held as an assumption because no reference machine is declared in rig/machine.yml; make measure-rig on a declared machine replaces this node with a measured one. Triangular because it has not been measured: once it is, the shape becomes a normal around the measurement, which is a change of claim and not only of numbers (ch13) |
| ○ | engineers, full-time equivalent | assumption | engineers this fleet occupies, full-time equivalent. Triangular, and the shape cannot express what actually happens: people are not divisible, so the real distribution is lumpy in the way ch08 calls a regime change |
| ○ | records held, day one | assumption | stated workload (ch02) — what the service holds today: its database and the objects users have uploaded, before replication, indexes or compression |
| ◐ | support rate | vendor claim | annual support as a fraction of capital cost. Triangular because it is negotiated inside a band the market sets rather than drawn from one: the spread is what different buyers get, not what varies from year to year |
| ○ | utilisation the model will admit to | assumption | where this model stops being about queues (ch06) |
| 37 inputs | 6 fact, 8 vendor claim, 23 assumption |
Source — web_service-reference · every input on a slider
The table carries three marks. The one that matters in this room is vendor claim: a number supplied by the party being paid. It may well be right. It has not been checked here, and it is coloured differently in every figure in this book for that reason (ch03). The finance audience is entitled to know which of the inputs to a capital request came from the supplier.
Handing this over unprompted makes the rest of the document more believable, and little else does. A model that volunteers which of its inputs are guesses is not a model trying to win an argument.
Three ways to lose the room
Presenting the interval instead of the decision. “Somewhere between these two figures” with no recommendation is not caution. It is a refusal to do the last part of the job. The person across the table is being asked to absorb uncertainty that you understand and they do not.
Presenting a high percentile as the cost. It gets approved, the money is set aside, and the actual spend comes in well under. That looks like success exactly once. The second time, the number is discounted before you have finished saying it, and the discount is applied by somebody who does not know which parts of it were conservative.
Presenting the median as though it were the plan. The first two fail in how they look. This one fails in what happens: half the futures cost more, and nothing has been said about them.
Each of the three is a way of not saying the sentence.
What to hand over
One page:
The number, and the sentence.
The decision table: what each design costs, and how often each one breaks.
The ceilings, in the language of what fails rather than the language of utilisation.
The provenance table, marked.
A link to the model file, because it re-runs. They can change an input and see what happens.
And the cash flow, separated by year rather than summed. ch15 says why this book does not discount: the rate is a policy decision, not an engineering one. The first thing a finance team will do with a five-year total is discount it. They can only do that if the years have not already been added together.
Key takeaways
Finance wants a number to commit to. Engineering wants to say what the number hides. Both are right, and they fit together.
The deliverable is one number and one sentence saying what it leaves out. The number is a choice made out loud: the median, a high percentile, or a round figure above the median. The sentence names something specific.
Present two designs with a price, not one design with an interval. The question becomes whether the difference between the columns is worth the difference in how often each one breaks, and that is a question the room can answer.
Lead with the rows that have a consequence in them. How often the fleet cannot serve the busy hour it was bought for is a sentence everybody in the room has an opinion about.
Volunteer where every number came from. A model that says which of its inputs are the supplier’s is not a model trying to win an argument, and it is believed more for it.
What this cannot tell you
What running out is worth. Every figure in the decision table is a cost of building. There is no term anywhere in this model for what happens when the queueing ceiling is breached: the busy hour spent turning users away, the emergency purchase at list price, the quarter spent on it, the conversation with whoever was promised the service. The right-hand column’s extra capital buys a reduction in that risk. The model prices the capital precisely and the risk not at all. Anybody who says the extra machines are not worth it is making a claim about a number this book has not measured.
Your organisation’s appetite for it. How much should a real chance of the busiest hour going over the knee cost to avoid? That is not an engineering quantity, and there is no defensible way to derive it from the model. It belongs to the people who carry the consequence. That is one more reason to put the ceiling row in front of them rather than resolving it yourself.
What the money is worth. The totals here add dollars from different years as though they were the same dollar. They are not. Applying a discount rate would change the comparison between a design that spends capital up front and one that spends it over time, and the two columns above differ in that way. The model hands over the shape of the spend so that somebody can apply theirs. It does not pretend the undiscounted total is the answer.
Whether the structure is complete. ch20 is the standing limitation, and it does not stop applying because the audience has changed. The decision table is a comparison between two designs inside one model, and both columns inherit whatever that model is missing. The comparison is more robust than either total, because a missing cost line that scales with host count hurts both columns. But “more robust” is not “unaffected”.
Whether it worked. There is no measurement in this repository of whether a document shaped like this gets a better decision than one shaped some other way. This chapter is the one place in the book arguing from experience rather than from a stamped result.
Problems
Three, in tests/a_tco_for_finance/. The first two have tests and neither is arithmetic. The last
has no test, and says why.
21.1 — The decision table. The test hands you the stamped result of each of the two designs. Build the two-design comparison from them, with the columns that answer the question and no others. The test grades the numbers against the stamped results and the shape against what fits in somebody’s head.
def decision_table(summaries: dict[str, dict]) -> list[dict]:
"""Problem 21.1 - two designs, priced, with their risk.
``summaries`` holds the stamped result of each of the chapter's two designs, keyed by scenario
name. Both are scenarios of the web service model: ``reference`` buys what the model
recommends at the point estimate, and ``sized_for_growth`` buys the same fleet sized for the
growth we might get rather than the growth we expect.
A summary is the build's record of one run, as data. ``summary["nodes"]`` has an entry per
node of the model, and the figures the table needs live at:
``nodes["hosts"]["point"]`` hosts purchased
``nodes["tco"]["summary"]["p50"]`` and ``["p95"]`` the five-year total's median and
its 95th percentile
``nodes["queueing_headroom"]["ceiling"]["p_over_limit"]`` how often the queueing ceiling
is breached
Return a list of rows, one for each design. Each row is a dictionary with exactly these keys:
``"scenario"`` the scenario's name
``"hosts"`` hosts purchased
``"tco_p50"`` the median five-year total
``"tco_p95"`` the 95th percentile of it
``"p_over_the_knee"`` how often the queueing ceiling is breached
Build it from the stamped results rather than by re-running anything, so that the table
somebody is shown is the table the build produced.
Then read it as the person on the other side of the table will. They have one question - what
am I buying and what am I buying it *instead of* - and every column is there to answer it or it
should not be there. A handful of columns is not a constraint; it is what fits in somebody's
head while they decide.
"""
raise NotImplementedError("problem 21.1")The same check at a desk: python3 -m pytest tests/a_tco_for_finance/test_problem_1_decision_table.py -m problem
21.2 — They have asked for one number. Give it. Then write the sentence. Any defensible choice of number passes. The test checks that it came out of the model, that it is rounded to a precision the model can support, and that the sentence names something specific.
def one_number() -> tuple[float, str]:
"""Problem 21.2 - they have asked for a single number.
They will. Refusing is not an option that exists, and answering with a range is a way of
refusing that annoys people without informing them.
Return the number you would give for the reference design's five-year total, and a string
saying what it hides.
Any defensible choice passes: the median, a percentile, a rounded figure. The test
checks that it came out of the model, that you rounded it to a precision the model can support,
and that your sentence names a specific thing - the percentile you chose, the structural
omission, the assumption the whole thing rests on. A sentence that says "it is uncertain" is
not naming anything.
The point is that the sentence is the deliverable. The number is what gets written down; the
sentence is what makes it honest, and you get one.
"""
raise NotImplementedError("problem 21.2")The same check at a desk: python3 -m pytest tests/a_tco_for_finance/test_problem_2_one_number.py -m problem
21.3 — Write the page, and hand it over. No test: a page is graded by the person it is for.
Write the one-page version for your own system and give it to whoever signs for it. A recommendation, what it rests on, what would change it, and the cost of being wrong in each direction.
Then do the part that is not writing: watch what they ask. The question they ask first is the thing your page failed to answer, and it is almost never the one you expected to spend a page on.
A good answer is one page and gets a decision. If it gets a request for more detail, the detail they asked for belongs on the page and something currently on it does not. This book has no measurement of whether a document like this works, which is why the only test available is handing it to somebody.
Where to go next
ch22 is the last step of the argument. When the two columns are two quotes rather than two sizes of one fleet, the difference between them is what somebody is deciding, and it has an interval of its own.
ch23 is what happens afterwards: three years later, when one of the futures in that table turned out to be the one you got.
What remains after it is reference material: Appendix A for the model file format in full, Appendix B for the sampler read end to end, Appendix C for choosing a shape, Appendix E and Appendix F for the web service and observability models in full, and Appendix G for the vocabulary, including the terms this book refuses.
If you read one thing again, make it ch20. Everything in this chapter is about presenting what the model knows, and the hardest sentence to write is still the one about what it does not.