ch22 · Comparing two TCOs
The question
Two quotes for the same workload, each with an interval on it. Which is cheaper, and how often would that turn out to be wrong?
ch21 put two designs side by side and called the difference between them a decision. This chapter is about the difference itself. Two totals with intervals on them overlap, and a reader shown the overlap concludes that nothing can be said. Something can. The two designs face the same future, and a difference taken future by future is far narrower than the totals it came from.
It is also the chapter for a comparison somebody else built. A vendor’s comparison arrives with one column already won. The last section says what such a comparison has to show before you believe it, and each item points at the chapter that says why.
The material
Two quotes for one workload
The incumbent is the design the plan was built on: the platform the team already runs, on the hosts ch12 sized and ch18 priced. The challenger is a second quote for the same workload. Its hosts are denser and cost more each, it licenses per host rather than per core, its support rate is higher, and moving to it costs something the incumbent does not pay.
Both are scenarios of the running example, and both hold every line of the quote exactly.
| Line | The incumbent’s quote | The challenger’s quote |
|---|---|---|
| cores per host | ◐ 16 cores per host | ◐ 64 cores per host |
| ram per host | ◐ 64 GiB per host | ◐ 256 GiB per host |
| disk per host | ◐ 2 TB per host | ◐ 8 TB per host |
| host price | ◐ $6,538 per host | ◐ $26,000 per host |
| host power | ◐ 310 W per host | ◐ 700 W per host |
| licence per core | ◐ $80.00 per core per year | ◐ $0 per core, declared |
| licence per host | ◐ $0 per host, declared | ◐ $2,600 per host per year |
| support rate | ◐ 11.8% of capital, per year | ◐ 13.0% of capital, per year |
| hosts in the fleet | ○ 54 hosts | ○ 14 hosts |
| one-off cost of moving to this design | ○ $0, declared | ○ $250,000 |
| ● traceable to a stamped measurement, an invoice or a published specification · ◐ stated by someone selling it; plausible, unverified, and never promoted · ○ a decision this model makes, which a reviewer may disagree with |
Source — comparison and web_service-incumbent and web_service-challenger · sizing.evaluate
Read the marks first. Almost every line on both sides is a vendor’s claim, whichever vendor it is. The model gave the host price a band because nobody had quoted one (ch03). A quote is a number, so the band goes and the number stays. What is left uncertain is what no quote fixes: the workload, its growth, the electricity, the people.
Two lines are not the vendor’s. The host count is the model’s own sizing rule, applied to each host at the point estimate, so both designs are sized to the same margins by the same rule (ch12). The cost of the move is the team’s estimate, marked as an assumption, because the challenger did not quote it and the incumbent has nothing to move.
Here is the challenger’s quote as the toolkit holds it. A comparison starts from a file like this one, and the file is where the reasons for its numbers live.
scenario: challenger
title: The challenger's quote, held exactly
because: >-
A second quote for the same workload: denser hosts at a higher price each, drawing more power
each, licensed per host rather than per core, with a higher support rate, and a one-off cost
of moving that the incumbent does not pay. Every line is pinned, because a quote is a number.
The host count is what the model's own sizing rule recommends at the point estimate for this
host, so both designs are sized to the same margins by the same rule; the licence per core is
a declared zero because this quote has no such line; and the move is the team's estimate, not
the vendor's. Same seed and the same pinned inputs as the incumbent, so that the two totals
can be subtracted future by future (ch22).
samples: 100000
seed: 20260916
overrides:
cores_per_host: 64
ram_per_host: 256
disk_per_host: 8
host_price: 26000
host_power: 700
licence_per_core: 0
licence_per_host: 2600
support_rate: 0.13
hosts: 14
migration_cost: 250000
And here it is running. Drag hosts in the fleet up by one and watch the five-year total pass the incumbent’s.
The challenger’s quote, with every input on a slider. The incumbent’s is linked under every table on this page.
Every line on both sides
The challenger licenses per host. The incumbent’s quote has no such line, and until this chapter the model had no such node. A line that one side does not have is a cost that side never pays, and a comparison with a line missing from one column is won by omission.
So the model carries two lines it did not need before, and both are zero on the incumbent’s side. The zero is declared, with a source, rather than left out:
licence_per_host:
kind: input
decided: world
unit: USD/host/year
value: 0
label: licence per host
note: >-
A declared zero, not an absent line. This quote licenses the platform software per core,
so its per-host charge is nil and the model says so, with a source. A comparison needs
the same lines on both sides: a line that one quote does not have is a cost that quote
never pays, and the only honest way to leave it out is to put it in at zero (ch22).
provenance:
kind: vendor_claim
source: "the price list: per core, with no per-host charge. A quote that licenses per host
instead puts its figure here and zero in licence_per_core, and the two totals are then
compared like for like"
range: [0, 12000]
The second is the cost of the move. It is the line an incumbent’s comparison never shows and a challenger’s never includes:
migration_cost:
kind: input
decided: you
unit: USD
value: 0
label: one-off cost of moving to this design
note: >-
A declared zero. This is the design the plan was built on, so there is nothing to move.
A design that would replace it carries the cost of the move here: the months of running
both, the data copied twice, the people doing that instead of their jobs. It is the line
an incumbent's comparison never shows and a challenger's never includes, which is why it
is a node with a source rather than a footnote (ch22).
provenance:
kind: assumption
source: "nothing to migrate: the platform the plan already runs on, on the hosts already
quoted. A challenger's scenario overrides this with the team's own estimate of the move,
marked as what it is"
range: [0, 1000000]
With both lines on both sides, the two totals contain the same things and can be subtracted line by line:
| Over five years | Incumbent | Challenger | Challenger minus incumbent |
|---|---|---|---|
| Hosts | $353,071 | $364,000 | +$10,929 |
| Network | $68,144 | $17,667 | −$50,477 |
| The move | $0 | $250,000 | +$250,000 |
| Energy | $137,792 | $80,548 | −$57,244 |
| Licences | $345,600 | $182,000 | −$163,600 |
| Support | $248,415 | $248,083 | −$331 |
| People | $849,062 | $849,062 | none |
| Five-year total | $2,002,083 | $1,991,361 | −$10,722 |
Source — comparison · sizing.evaluate
Read down the last column. The challenger wins on licences, on energy and on network, and loses on the move. The hosts cost about the same in total, because fewer of them cost more each. Support is a share of capital on both quotes and the capital is about equal, so it barely moves.
The largest line in either total is people, and it does not appear in the difference at all. Both designs are run by the same engineers at the same cost, so the line is the same on both sides and cancels. That is the general rule. What both designs pay alike drops out of the comparison, however large it is. A comparison is about the lines that differ. The total is only where they are added up.
The last row is the whole comparison at the point estimate. The challenger is cheaper, by a sliver of either total. That is the number a spreadsheet stops at.
Subtract futures, not intervals
Both totals have intervals, and the two intervals are nearly the same interval.
That is not a coincidence. Most of what either total is uncertain about is shared. The electricity price, the building’s overhead, what an engineer costs and how many are needed are the same inputs on both sides, and the same draw of each reaches both designs. When the electricity price comes out high, it comes out high for both fleets. When an engineer costs more, both totals rise together.
So the difference between the totals is far less uncertain than either total. The toolkit takes it the only honest way: the same future, both fleets, subtract. Both scenarios pin the same inputs and share a seed, so every input neither quote fixes is drawn once and reaches both designs, and the toolkit checks that before it subtracts anything.
The interval on the difference is a fraction of the interval on either total. The challenger is cheaper in most of the futures the model thinks plausible, and dearer in a large minority of them. Both of those are the answer. A comparison that reports only the first is the one the challenger would show you. A comparison that reports neither is the one a spreadsheet shows you.
| Challenger minus incumbent, five-year total | |
|---|---|
| At the point estimate | −$10,722 |
| Across the same futures, middle nine in ten | −$72,389 to +$28,521 |
| With each design in a future of its own | −$734,933 to +$701,940 |
| Subtracting the ends of the two intervals | −$1,013,903 to +$986,004 |
| Futures in which the challenger is cheaper | 69% |
| Futures in which the incumbent is cheaper | 31% |
Source — comparison · sizing.evaluate
The third and fourth rows are what goes wrong when the futures are not shared. Draw each design in a future of its own and subtract, and the interval on the difference is many times wider, because the shared uncertainty is counted twice, once with each sign. Subtract the ends of the two intervals, which is what two totals side by side invite a reader to do, and it is wider still. Neither is a statement about the decision. Both are statements about how the arithmetic was done.
ch14 made the same point from the other side. Two inputs that move together, drawn independently, understate the width of a total. Two totals that move together, subtracted independently, overstate the width of their difference. It is one fact about shared uncertainty, and which way it cuts depends on which side of the subtraction it sits.
Problem 22.1 is that subtraction, and its test can tell which way you did it.
Cheaper at what risk
A total is the price of a fleet that works. Two fleets that cope different amounts of the time are not the same purchase at two prices, so a comparison carries the ceilings as well as the money.
| Ceiling | Incumbent | Challenger |
|---|---|---|
| working set against memory | 35% over its limit 49% past the allowed line | 33% over its limit 47% past the allowed line |
| utilisation, counting coordination | 48% over its limit 64% past the allowed line | 32% over its limit 48% past the allowed line |
| disk fill at horizon | 29% over its limit 44% past the allowed line | 27% over its limit 42% past the allowed line |
| utilisation with one host down | 30% over its limit 47% past the allowed line | 31% over its limit 47% past the allowed line |
| utilisation at the busy hour | 30% over its limit 46% past the allowed line | 28% over its limit 44% past the allowed line |
| fraction of the fleet doing nothing useful | 0% over its limit 1% past the allowed line | 0% over its limit 0% past the allowed line |
Source — comparison and web_service-incumbent and web_service-challenger · sizing.evaluate
They cope about equally often, and that is by construction. Both were sized by the same rule to the same margins, so both are over the knee at the busy hour in about the same share of futures. A comparison that sized one design to the daily mean and the other to the busy hour would show a cheaper fleet that copes less often, and would not say so.
Two rows differ, in opposite directions. The challenger’s hosts spend less of their capacity agreeing with each other, because there are fewer of them (ch07), so it is over the coordination ceiling less often. And losing one host of a small fleet costs more of the fleet than losing one of a large one, so the challenger is over the one-host-down ceiling slightly more often (ch11). Neither shows up in the money. Both are part of what is being bought.
What would flip it
The difference is small, so the next question is what would reverse it. For each line the challenger could argue about, and for each input both designs share, the toolkit finds the value at which the two totals tie, with everything else at its point estimate.
| Input | Whose | As quoted, or at the point | The totals tie at | Verdict |
|---|---|---|---|---|
| one-off cost of moving to this design | the challenger’s quote | $250,000 | $260,722 | +4.3% on the quote |
| host price | the challenger’s quote | $26,000 per host | $26,464 per host | +1.8% on the quote |
| licence per host | the challenger’s quote | $2,600 per host per year | $2,753 per host per year | +5.9% on the quote |
| hosts in the fleet | the challenger’s quote | 14 hosts | 14.2 hosts | +1.2% on the quote |
| engineers, full-time equivalent | a claim about your people | 1.13 | 1.15 | +1.3% of an engineer’s time, on the challenger’s side |
| electricity price | shared by both | $0.134 per kWh | $0.109 per kWh | inside the middle eighty per cent of what it could be |
| PUE | shared by both | 1.4 | 1.14 | outside the middle eighty per cent, inside the range the model admits |
| network price per host | shared by both | $1,262 per host | $1,091 per host | inside the middle eighty per cent of what it could be |
| annual growth factor | shared by both | 1.34 | — | no value of it moves the difference |
Source — comparison · sizing.evaluate
Read the challenger’s lines first. Every one of them ties within a few per cent of the quote. The move can cost a little more than the team estimated, the host price and the licence can each rise a little, and the fleet can grow by less than one host before the ordering reverses. Carry that last row into the meeting. How many hosts a workload needs on hardware nobody has run it on is the kind of number ch03 says to distrust, and this comparison turns on a fraction of one.
The row about people is a different kind of claim. Both quotes hold the engineers equal, because a vendor’s figure for how many people you will need is a claim about your organisation and not about their product. If the challenger cost a little more of an engineer’s time, the totals would tie. If it claimed to save some, the comparison would rest on that claim and on nothing else, since the people line is larger than every line the quotes differ on. A comparison that turns on a headcount the vendor supplied has a thumb on the scale, and this model refuses to take that number from a quote.
The shared inputs are next. Cheaper electricity favours the incumbent, because the challenger’s advantage on energy shrinks with the price, and the tie sits inside the middle eighty per cent of what the price could be. So does the tie on the network price. The building’s overhead would have to be better than the model thinks plausible before it flipped anything. And growth cannot flip it at all. Both fleets were bought before the growth arrived, so the input that dominates every other tornado in this book does not move the difference by a dollar.
The chart is a few bars and a count. Everything the two designs pay alike has a bar of no length, and that includes the growth, the busy hour, the salary and the head count. The comparison is sensitive to the two prices that scale with the host count and to almost nothing else. That is what a comparison between two quotes for one workload is usually sensitive to, and it is rarely what the argument in the room is about.
Problem 22.2 finds one of those break-evens by hand.
What a competitive comparison has to show
A comparison somebody else built arrives with one column already won. This is what it has to contain before you believe it, where the book covers each item, and how the comparison is tilted when the item is missing.
| What to look for | Where it is taught | How the comparison is tilted without it |
|---|---|---|
| The workload it was sized for, stated, and both designs sized by the same rule to the same margins | ch02, ch12 | one design sized to the daily mean and the other to the busy hour |
| The same horizon and the same refresh convention on both sides | ch15 | a three-year quote against a five-year one |
| Every line on both sides, with a declared zero where a quote has no such line | this chapter | hardware only on one side, fully loaded on the other |
| Whose number each line is, marked on both sides, the presenter’s own included | ch03 | list price for the other side, the negotiated price for yours |
| The cost of the move, on the side that has to move | this chapter | the incumbent never shows it; the challenger never includes it |
| The difference as an interval over shared futures, and the share of futures in which it flips | this chapter | two intervals side by side, which say nothing, or one number, which says too much |
| Each design’s ceilings, as how often it fails to cope | ch11, ch21 | cheaper because it copes less often |
| The break-evens, and whether each lies inside a range anybody would defend | this chapter, ch19 | a comparison that turns on an input nobody measured |
| The people line held equal unless there is evidence, and marked as a claim if not | this chapter | a headcount the vendor supplied |
| Unit costs on the same denominator and the same period | ch17 | per terabyte-month against per terabyte-year, the busy hour against the mean |
| What was left out, named | ch20 | a line neither total has, which favours whichever side it would have hurt |
| Cash by year, undiscounted, so that finance can apply its own rate | ch15, ch21 | a discount rate chosen to flatter the side that spends later |
| A file that re-runs, so the reader can change an input and watch | the introduction | a bar chart, with nothing in it to change |
Check the first item twice. Two quotes that were sized differently are not two prices for the same thing, and no amount of care over the lines below it repairs that.
Key takeaways
A comparison is about the lines that differ. What both designs pay alike drops out of the difference however large it is, so every line goes on both sides, with a declared zero where a quote has no such line.
Subtract futures, not intervals. The two totals share most of their uncertainty, so the difference taken future by future is far narrower than either total. Subtracting independent draws, or the ends of two intervals, says how the arithmetic was done and nothing about the decision.
A total is the price of a fleet that works. Two fleets that cope different amounts of the time are not the same purchase at two prices, so the comparison carries the ceilings as well as the money.
Find what would flip it. The break-evens say how far each line and each shared input can move before the ordering reverses. Usually that is a few per cent on the prices that scale with the host count, and nothing on the inputs the argument in the room is about.
A comparison somebody else built arrives with one column already won. The checklist says what it has to show, and the item to check twice is whether both designs were sized by the same rule to the same margins.
What this cannot tell you
What either vendor would charge. Both quotes are held exactly. A quote is a number, and this chapter takes it as one. What you would pay after negotiation is a different number on both sides, and a discount that is the same percentage on both is not the same money, because the two capital lines differ.
What the move will cost. The migration line is the team’s estimate, marked as an assumption, and it is the kind of number that is usually low. The break-even says how wrong it can be before the ordering flips. Nothing here says how wrong it is.
Anything about performance. Both quotes share one CPU time per request, because it is the same software on different hosts. A quote for a different platform would carry a different service demand, and that number is a rig measurement (ch03) which no vendor’s benchmark replaces. The host-count break-even is the hedge. It says how much slower than the spec sheet the challenger’s hosts could be before the comparison reverses, and on this pair the answer is not much.
What the structure omits. Both totals come from one model, so a line neither quote has is a line neither total has (ch20). Rack space, the second environment, the cost of running both platforms during the move, the exit cost at the end of the horizon and the price at renewal are in neither column. A missing line favours whichever side it would have hurt, and the comparison cannot say which side that is.
Whether the futures are shared. They are here by construction, because both designs are scenarios of one model. Two quotes built as two models and joined by figures in documents lose that (ch18), and the interval on their difference comes back as wide as the independent one above.
Which side you are on. A comparison built by the party that wins it is still a comparison, and the checklist above is the same whoever built it. What the checklist cannot supply is the number a prospect is entitled to and rarely gets: the share of futures in which the presenter’s own design loses.
Problems
Three, in tests/comparing_two_tcos/. The first two have tests. The last does not, and says why.
22.1 — Subtract futures, not intervals. The test evaluates both quotes on the same draws and hands you the two five-year totals, one entry per future. Subtract them sample by sample. Return the interval on the difference and the share of futures in which the challenger is cheaper. The test can tell whether you subtracted two independent draws instead, because that answer is several times wider.
def paired_difference(incumbent: np.ndarray, challenger: np.ndarray) -> dict:
"""Problem 22.1 - subtract futures, not intervals.
Two quotes for one workload, the incumbent's and the challenger's. Each has a five-year total
with an interval on it, and the two intervals overlap. That is not the question. The question
is what the *difference* between them looks like, future by future.
The test evaluates both quotes on the same draws and hands you the two totals: ``incumbent``
and ``challenger``, one five-year total per sampled future, in the same order on both sides.
The i-th entry of one and the i-th entry of the other are the same future - the same draw of
the electricity price, the same growth, the same salary - priced once for each design.
Subtract the incumbent's total from the challenger's, sample by sample, and return a
dictionary with exactly these keys:
``"p5"``, ``"p50"``, ``"p95"`` percentiles of the difference, challenger minus incumbent
``"share_challenger_cheaper"`` the fraction of futures in which the challenger costs less
A negative difference is the challenger being cheaper.
If your interval comes out several times wider than the book's, you have subtracted two
independent draws - the electricity price from one future against the price from another -
or the ends of the two intervals, and thrown away the fact that both designs live in the same
world. The test says so.
"""
raise NotImplementedError("problem 22.1")The same check at a desk: python3 -m pytest tests/comparing_two_tcos/test_problem_1_paired_difference.py -m problem
22.2 — How expensive can the move be? The test hands you the difference between the two totals at the point estimate, as a function of the one-off cost of moving. Find the cost at which the difference is zero. The break-even table prints the tie for the fleet the incumbent runs. The test grades that fleet and two the page does not print, with a function for each. An answer that moves when the fleet does cannot have been copied off the page. The total is linear in the cost of the move, so two evaluations fix the line.
def break_even_migration(difference_at: Callable[[float], float]) -> float:
"""Problem 22.2 - how expensive can the move be before the challenger stops being cheaper?
At the point estimate the challenger's quote is cheaper than the incumbent's before the cost
of moving is counted. The move is the ``migration_cost`` input, which the incumbent holds at
zero and the challenger's scenario overrides with the team's own estimate.
``difference_at(migration)`` is the model's own point evaluation. Hand it a one-off cost of
moving and it returns the challenger's five-year total minus the incumbent's at the point
estimate, with the incumbent on the fleet the test is grading and everything else as the two
scenarios hold it. Return the cost of moving at which that difference is zero: the cost at
which the two five-year totals are equal.
The fleet is the test's to choose because the chapter prints the tie for the fleet the
incumbent runs. The test builds one ``difference_at`` for that fleet and two for fleets the
page does not print, and an answer that moves when the fleet does cannot have been copied off
the page.
The total is linear in the cost of the move, so two evaluations determine the line. Bisection
works too, and is the method to reach for when the next break-even is not linear.
"""
raise NotImplementedError("problem 22.2")The same check at a desk: python3 -m pytest tests/comparing_two_tcos/test_problem_2_break_even.py -m problem
22.3 — A comparison somebody showed you. No test: the comparison is theirs, and nothing here has seen it.
Take a comparison that was put in front of you, by a vendor or by a colleague, and hold it against the checklist above. Write down which items it has and which it lacks. Then repair the one that matters most and see what it does to the ordering: put the missing line on the side that lacked it, size both designs by the same rule, or hold the people line equal.
A good answer names the items the comparison failed and says what the difference did once the worst of them was fixed. It would be falsified by a comparison that passed every item and whose ordering survived every repair. That happens, and then you know the comparison was sound. If the ordering flipped on the first repair, you have found what the comparison was for.
Where to go next
Kahn and Marshall [Kahn & Marshall (1953)] set out the trick this chapter rests on, in the paper that first treated Monte Carlo as a problem in experimental design: sample two things you mean to compare on the same draws, so that what they share cancels. Wright and Ramsay [Wright & Ramsay (1979)] show when it fails. The shared draws have to reach both designs the same way, and a design that responds to a draw in the opposite direction gets a wider difference, not a narrower one.
ch23 is what happens after one of the two quotes is chosen, and the future that arrives is one of the ones in which it was the wrong choice.