Sizing and TCO

ch15 · Capex, opex and where the total stops

Builds on ch12.

The question

What do you pay once, what do you pay every month, and what does this book deliberately not model?

ch12 chose the fleet. This chapter prices it. The arithmetic is easy. Two things are hard. The first is the split between the invoice somebody signs and the bills that arrive afterwards. The second is knowing where a cost model should stop.

The material

The split

AmountShare of the total
Capital, paid once$421,21421%
Hosts$353,07118%
Network$68,1443%
Running, over 5 years$1,580,86979%
Energy$137,7927%
Licences$345,60017%
Support$248,41512%
People$849,06242%
Total$2,002,083

Source — web_service-reference · every input on a slider

Read the two headline rows first. Over the declared horizon, the running cost is several times the capital cost. Only the capital got argued about.

Capital cost, or capex, is what you pay once. Running cost, or opex, is what you pay every month for as long as you keep the fleet. The tables use those words, and so will the room. The split between them is not an accounting formality.

Capital arrives as one invoice with somebody’s signature on it. So it gets a meeting, a comparison and a negotiation. Running cost arrives in pieces, every month, from several directions: an electricity bill, a support renewal, a fraction of a salary. Nobody in particular decided any of them. Each piece is too small to argue about. The total is not.

Problem 15.1 is that division. Run it against the table above. The running cost overtakes the purchase well inside the horizon the fleet was bought for. The argument was had about the smaller part.

What is in each

The capital rows split two ways. The hosts dominate. The network port each one needs is the smaller share. It also grows with fleet size in a way this model does not capture: a fleet twice the size needs more than twice the switching.

The four running rows are four different kinds of number.

Energy is physics: watts times hours times price, with a facility multiplier on top. It is the one line in the model with no judgement in it. ch16 is about what happens when it stops being a line item and becomes the binding constraint.

Licences are a vendor’s price per core per year. Every core the fleet has is a core somebody pays for annually, busy or idle. So a decision about hosts becomes a recurring bill nobody remembers agreeing to. Here it is the second-largest running line.

Support is a percentage of capital. It is a deferred part of the purchase, indexed to the purchase, rather than a running cost. Negotiate the capital down and the support falls with it. That is worth knowing before you negotiate.

People is the line most models omit, and the hardest to defend either way. The model carries a fraction of an engineer at a fully loaded rate. Both numbers are assumptions. The line is in the model because leaving it out is a decision too, and a less honest one. Here it is the largest line in the total. That is the usual state of affairs, and the usual reason it is left out.

What moves the running cost

InputKindannual opex at its p10at its p90Swing
engineers, full-time equivalentinput$262,247$387,863$125,616
licence per coreinput$281,614$385,294$103,680
fully loaded salaryinput$276,768$367,486$90,718
host priceinput$303,191$335,038$31,847
support rateinput$304,811$330,621$25,810
electricity priceinput$306,148$331,932$25,784
host powerinput$311,695$321,549$9,854
PUEinput$313,346$319,934$6,588

Source — web_service-reference · every input on a slider

The running cost moves with the things you would expect, in an order most people get wrong. The top three bars are about people and licences: how many engineers, what a licence costs per core, and what an engineer costs. The electricity price is the line everybody expects to argue about. It is near the bottom, with the support rate beside it.

The refresh, and the convention nobody writes down

Hardware is bought against a refresh cycle. Where that cycle lands relative to the horizon decides how many purchases the total contains.

Take a five-year horizon with a five-year refresh. That is either one purchase or two, depending on whether the refresh at the end counts. Nobody writes the convention down. The difference is an entire fleet of hosts, the largest capital line in the model. Two people can produce two totals from the same inputs, and both are right.

Problem 15.2 is that boundary. Which convention is correct is a modelling choice, not a fact. So the test accepts either, and checks only that you are consistent. Saying which you chose is the part that is not optional.

What this book does not model, and why

These are named here so nobody has to guess whether they were forgotten.

Tax and depreciation. How capital is written down, over what period and against what depends on the jurisdiction and the company. All of it changes the answer a great deal. A book that guessed at it would be giving financial advice. So this chapter produces a cash total: what leaves the bank, and when. Hand that to somebody who knows your tax position.

Discount rates. Money later is worth less than money now. The rate is a policy decision, not an engineering one. The model produces undiscounted cash flows on purpose, so that whoever applies a discount rate applies theirs.

Procurement reality. Lead times, minimum orders, the discount you get for asking, the price that changes between the quote and the purchase order. All real, and none of it a modelling problem.

Anything with a contract in it. Commitments, reserved capacity, early termination. Those terms decide whether a plan can change. They are not quantities.

The line this chapter draws is cash out, over time, from physics and prices. Everything on the other side of it belongs to somebody whose job it is.

What the model says

OutputPoint estimate90% intervalUnit
hosts the model recommends5420 to 230host
hosts in the fleet54fixedhost
five-year total cost of ownership$2,002,083$1,508,230 to $2,923,724USD
cost per million requests$1.82$0.55 to $5.98USD/megarequest
cost per stored TB per month$839.64$305.33 to $2,108USD / TB / month
capex$421,214$272,130 to $658,773USD
annual opex$316,174$227,660 to $484,257USD / year
annual energy206,269159,636 to 269,269kWh / year
utilisation at the busy hour0.6440.156 to 2.48
utilisation with one host down0.6560.159 to 2.53
working set against memory0.7460.194 to 2.67
disk fill at horizon0.6700.213 to 2.12
fraction of the fleet doing nothing useful0.3210.213 to 0.457
utilisation, counting coordination0.9480.233 to 3.81
utilisation0.6440.156 to 2.48
residence time0.03670.0128 to 0.857second
time spent queueing0.02360.0021 to 0.839second
requests in the system1,562160 to 107,336request
requests in flight, if none waited556135 to 2,147request
how much the queueing view understated it1.471.27 to 1.84
fraction of the peak already built0.3280.186 to 0.584

Source — web_service-reference · every input on a slider

Key takeaways

  • The running cost is several times the capital, and only the capital gets argued about. Capital arrives as one signed invoice. Running cost arrives in pieces that are each too small to fight.

  • The four running lines are four different kinds of number. Energy is physics, licences are a price per core, support is a deferred share of the purchase, and people is the largest line and the one most models leave out.

  • Support is indexed to the capital. Negotiate the purchase down and the support falls with it, which is worth knowing before you negotiate.

  • Where the refresh lands against the horizon can add or remove a whole fleet. Which convention is right is a choice. Saying which one you chose is not optional.

  • The model produces cash out over time, from physics and prices, and stops there. Tax, depreciation, discount rates and procurement belong to somebody whose job they are.

What this cannot tell you

What the model’s structure omits. Every cost line above is one somebody thought of. There is no line for rack space, cross-connects, backup, disaster recovery, the database’s own licence if it has one, the network gear between racks, or the cost of the migration that fills the fleet. The model reports each absent line as zero, confidently. Neither ch13 nor ch14 can see it. That is ch20 · The missing node.

Whether the prices are yours. Every price in the model is marked as a vendor’s claim or an assumption (ch03). None was measured. A price is not the sort of thing this repository can measure.

What happens if the fleet is wrong. The cost model takes the host count as given. If ch12’s fleet goes over the knee partway through its horizon, the real total includes an unplanned purchase. No line here represents it.

Anything about when the money is spent. The split above is a total over a horizon. Whether the capital lands in one quarter or three changes nothing in this model, and a great deal in somebody’s budget.

Problems

Three, in tests/capex_opex_and_lifecycle/. The first two have tests. The last does not, and says why.

15.1 — When does running cost overtake the purchase? One division. Then look at where it falls relative to the horizon, and at which of the two parts got the meeting.

tests/capex_opex_and_lifecycle/stubs.py · crossover_yearyours to edit
def crossover_year(capex: float, annual_opex: float) -> float:
    """Problem 15.1 - when does the running cost overtake the purchase?

    Return the number of years at which cumulative running cost equals the capital cost. Fractions
    are fine; nobody's horizon lands on a birthday.

    One division, and the point is what it does to a conversation. Capital is the number that gets
    argued about, because it arrives as a single invoice with somebody's signature on it. Running
    cost arrives in pieces, monthly, from several directions, and is nobody's decision in
    particular. For most infrastructure the second overtakes the first well inside the horizon it
    was bought for, and the argument was had about the smaller part.
    """
    raise NotImplementedError("problem 15.1")

The same check at a desk: python3 -m pytest tests/capex_opex_and_lifecycle/test_problem_1_crossover.py -m problem

15.2 — The refresh, and the convention nobody writes down. Decide what happens when a refresh lands exactly on the end of the horizon. Defend it in a comment, and be consistent. The difference is a whole fleet.

tests/capex_opex_and_lifecycle/stubs.py · lifecycle_totalyours to edit
def lifecycle_total(capex: float, annual_opex: float, years: float, refresh_years: float) -> float:
    """Problem 15.2 - what a refresh cycle does to the total.

    Over ``years``, with the whole capital cost paid again every ``refresh_years``, return the
    total spend.

    The first purchase happens at the start. Decide what happens when a refresh falls exactly on
    the end of the horizon - whether you have bought a fleet you will not use - and defend
    whichever you choose in a comment. The test accepts either, and checks only that you are
    consistent about it, because this is a modelling choice rather than a fact.

    That is the honest state of the question. A five-year horizon with a five-year refresh is
    either one purchase or two depending on a convention nobody wrote down, and the difference is
    the largest single line in the model.
    """
    raise NotImplementedError("problem 15.2")

The same check at a desk: python3 -m pytest tests/capex_opex_and_lifecycle/test_problem_2_refresh.py -m problem

15.3 — Prices you can get. No test: the prices are the ones you can get, and nobody else can get them.

Every price in this book’s model is a vendor’s claim or an assumption, and the chapter says so. Try to get yours. Find the real figure for each cost line: hardware, power, licences, support and the people. Record where each came from.

Count how many you could obtain. In most organisations the hardware price is easy. The power price is held by a facilities team who have never been asked. The cost of the people is either unavailable or politically impossible to write down.

A good answer has a source per line and an honest count of the gaps. The gaps are the finding. A five-year total built from two real prices and four guesses is not a cost model. Knowing which is which is the difference between a number and a negotiating position.

Where to go next

ch16 takes the one line in this chapter that is physics rather than negotiation, and asks what changes when it stops being a line item and becomes the constraint.

ch17 turns the total into a number somebody outside the team can compare against something.