Aura Digital
Sofia · 42.69°N 23.32°E
All articles
Guides · July 26, 2026 · updated September 7 · 12 min read

Renting vs Buying GPU Compute: the Real TCO in 2026

Rental rates from 1,396 live offers, the electricity arithmetic nobody publishes, and where the break-even actually sits. Plus why the spread within one GPU model beats the gap between models.

GPU server in a rack cabinet - owned hardware versus rented cloud capacity

Every team that starts training or serving models hits the same fork: rent GPU time by the hour, or buy the hardware. The honest answer is not a preference. It is a single number — how many hours per month you will actually keep the GPUs busy — and the arithmetic below tells you which side of the line you are on.

This article uses live market data rather than illustrative figures. The rental prices come from 1,396 offers sampled on 3 August 2026 and are refreshed weekly; the electricity numbers are worked out to the kilowatt-hour so you can substitute your own tariff.

The short version

Renting wins when your usage is spiky, exploratory, or below roughly a third of the month. Owning wins when a machine would run most of the time, and the gap widens fast after that.

But there is a second finding that surprised us when we started sampling the market, and it changes how you should shop for rented capacity. More on that below.

What renting actually costs right now

Median hourly rates, with the cheapest and dearest offers for the same card:

GPU Median $/h Spread Cheapest Dearest VRAM Power
B200 6.63 2.3× 4.38 10.00 179 GB 1000 W
H200 4.15 2.5× 2.37 5.87 140 GB 700 W
H100 SXM 2.27 5.0× 1.34 6.67 80 GB 700 W
RTX PRO 6000 S 1.40 1.7× 0.93 1.60 96 GB 600 W
RTX PRO 6000 WS 1.33 2.5× 0.76 1.87 96 GB 600 W
A100 SXM4 0.95 8.7× 0.13 1.16 80 GB 400 W
L40S 0.57 1.9× 0.47 0.87 45 GB 350 W
RTX 4090 0.38 6.7× 0.12 0.80 24 GB 450 W
RTX 5090 0.44 6.9× 0.21 1.47 32 GB 575 W
RTX 3090 0.14 4.8× 0.04 0.20 24 GB 350 W

Source: 1,283 offers on Vast.ai, sampled 7 September 2026. USD, per GPU-hour.

What these rates are, and what they are not

One caveat before you use the numbers above, because it changes the conclusion for some readers.

Those are marketplace rates — capacity rented out by independent operators and small datacentres. They are not what a hyperscaler charges. Compare like for like:

1× GPU, on-demand Marketplace AWS on-demand Ratio
B200 $5.96/h $14.24/h 2.4×
H100 $2.10/h $6.88/h 3.3×
A100 $0.93/h $4.10/h 4.4×

AWS list prices, us-east-1, divided by GPU count: p6-b200.48xlarge $113.93/h, p5.48xlarge $55.04/h, p4d.24xlarge $32.77/h. Checked August 2026 — AWS raised GPU instance prices in July, so treat these as a moving target.

So a hyperscaler costs three to four times the marketplace rate for the same silicon. You are not paying that for the GPU — you are paying for an SLA, enterprise support, compliance paperwork, redundancy and someone to sue. For regulated workloads that premium is not optional, and for a bank or a hospital the marketplace is not a candidate at all.

Which way does this cut? In favour of owning. Every comparison in this article uses the cheaper marketplace rate, so if your realistic alternative is AWS, the case for your own hardware is three to four times stronger than the arithmetic below suggests.

AWS publishes no per-GPU list price for the H200 or L40S, which is why they are absent above. The larger point is the one that does not fit in a table: half the cards in this article have no hyperscaler equivalent at all — the RTX PRO 6000, 4090, 5090 and 3090 are not rented by AWS or Azure in any form. Workstation and consumer silicon is not something AWS or Azure rent you, at any price. If your workload fits on a 4090 or an RTX PRO 6000, the hyperscaler comparison does not exist — your choice is marketplace or ownership, and that is precisely why these marketplaces have a business.

The finding: which offer beats which GPU

Look at the spread column rather than the price.

An RTX 4090 costs between $0.05 and $0.71 an hour depending on whose machine you rent — a 13.7× range for identical silicon. An RTX 5070 ranges 18×. Meanwhile the median 4090 and the median 5090 are within two cents of each other.

In other words: which offer you pick matters more than which GPU you pick. A team that shops carelessly can pay more for a 4090 than a careful team pays for a 5090.

The spread is also a reliability signal. The bottom of a 14× range is not a bargain — it is usually an oversubscribed host, a machine on a domestic connection, or hardware that will be reclaimed mid-job. The cards with narrow spreads (RTX PRO 6000 WS at 1.8×, L40 at 1.3×) are the ones where the market has settled on what the capacity is worth.

If you are going to rent, budget time for vetting hosts, not just for the hours. We wrote up what to check before trusting a rented machine separately.

The cost everybody forgets: electricity

The ownership column is never just the hardware. Here is the arithmetic, at €0.18/kWh — a business tariff figure you should replace with your own — for a single card running around the clock:

GPU Draw kWh/month Electricity, 1 card
RTX PRO 6000 600 W 438 €79
H100 SXM 700 W 511 €92
RTX 5090 575 W 420 €76

Then multiply by roughly 1.5 for power-supply losses, cooling, and the CPUs, RAM and drives around the cards. An eight-GPU box therefore costs in the region of €900–1,100 a month to run before anyone touches the capital cost.

That sounds like a lot until you put it next to the rental column:

8× GPU, 24/7 Rent Own (power only)
RTX PRO 6000 WS $7,761/mo ~€946/mo
H100 SXM $13,239/mo ~€1,104/mo
RTX 5090 $2,575/mo ~€907/mo

The nuance in that table

The three rows behave completely differently, and this is the part most comparisons miss.

For datacentre cards, ownership is not close. Renting eight H100s around the clock costs eleven times what the electricity does. Every month of genuine 24/7 use retires a large chunk of the hardware cost.

For consumer-class cards, the gap narrows sharply. Eight rented 5090s cost about twice their own electricity — so the payback period stretches, and renting stays defensible far longer.

The reason is simple: rental pricing for consumer cards is fiercely competitive because the supply is enormous. Datacentre capacity is scarcer and priced accordingly. If your workload fits on consumer VRAM, renting is a stronger option than the usual advice suggests.

The break-even formula

Compare like with like — cost per productive GPU-hour:

rented    = rate/hour × hours used
owned     = (hardware ÷ months of life) + power + hosting + maintenance
break-even = hardware cost ÷ (monthly rental − monthly running cost)

Using the RTX PRO 6000 row: every month of 24/7 operation saves roughly $5,000 against renting. Divide your hardware cost by that figure and you have the payback period in months.

We do not publish server prices, because the number is meaningless without the configuration — VRAM, interconnect, storage and power headroom move it by a factor of three. The configurator produces a real quote in 3 working days, and the ROI calculator does this arithmetic against the same live rates used above.

The four questions that actually decide it

Cost is one axis. These are the others, and any one of them can override the arithmetic:

How sensitive is the data? If you are processing personal, medical or financial records, renting from an anonymous host is often simply not available to you as an option, whatever it costs. Owned hardware in a known rack, or a provider with a signed data-processing agreement, is the only lawful path.

How steady is the load? Sustained inference is the strongest case for owning. Occasional experiments and one-off training runs are the strongest case for renting. Be honest about which you are — most teams overestimate how busy they will keep a machine.

What is the budget shape? Renting is operating expenditure with no commitment. Buying is capital expenditure with a payback period. A company that cannot commit capital this year should rent even when the arithmetic favours owning, and there is no shame in that.

What happens if you need to double? Rented capacity scales in minutes. Owned capacity scales in weeks — procurement, assembly, burn-in, deployment. If your demand curve is unpredictable upward, that lead time has a real cost.

When renting is clearly right

  • You are still finding out whether the model works at all
  • Usage is under roughly a third of the month
  • You need a card class for a fortnight and never again
  • You cannot commit capital, or the project has no funding beyond a pilot

If that is where you land, the machine still has to sit somewhere and be looked after. GPU hosting and colocation covers the half that starts after the decision: rack, power, monitoring, and the income from spare capacity.

When buying is clearly right

  • A machine would run most of the month, most months
  • The data cannot leave premises you control
  • You have been renting the same class of card for over six months
  • Latency or bandwidth to your own systems matters

The case that breaks the tie: renting out your own idle time

This is the option most comparisons ignore. A machine you own does not have to sit idle between your own jobs — the same platforms you would have rented from will rent your spare capacity to someone else.

That changes the arithmetic in a way pure ownership does not. The payback period is no longer driven only by what you save, but by what the machine earns while you are not using it. We build for clients on exactly this basis: design, assemble, stress test, deploy, and put the idle hours to work.

How much it earns depends on the card class, the datacentre and how much of the month you actually leave free — which is why we model it per case rather than quoting a rate of return.

Run your own numbers

Two inputs decide almost everything: hours per month and card class. Everything else is second-order.

The ROI calculator runs the comparison against the same weekly-refreshed rates used in this article. If you would rather talk it through with the workload in front of you, tell us what you are running and we will do the arithmetic with you.

A note on the middle path

You do not have to choose once and for good. A common and sensible pattern is to rent while the workload is still shifting, watch the monthly hours, and buy once the machine would clearly be busy. The rented months are not wasted — they are how you learn the number that makes the decision.

#gpu server mieten#gpu server kaufen#ki server#tco#gpu сървър#roi#ai инфраструктура#buy gpu server#rent gpu power

We build custom AI servers

From a single GPU workstation to a warranted server - we design, assemble and maintain. Including SEV-SNP builds with RTX PRO 6000 / H200.

Get in touch
Call Us