Since early 2025, the computers behind cloud AI have improved at a startling rate. They gained far more memory, became much faster and made each AI request cheaper to process.

Meanwhile, the best consumer graphics card you can buy is still the one launched in January 2025. It now costs about 80% more.

That gap is the story. Cloud hardware and consumer cards rely on the same factories and memory suppliers. One improved rapidly while the other stood still. Physics did not make that choice. Business did.

The gap is memory, and it is widening

Memory per accelerator, flagship data centre part vs flagship consumer card (GB)

Data centre flagship Consumer flagship

Data centre: A100 80GB to H100 80GB to MI300X 192GB to B200 192GB to MI355X 288GB. Consumer: RTX 3090 24GB to RTX 4090 24GB to RTX 5090 32GB. The ratio moved from roughly 3x to 9x.

View as table

It is tempting to call this a shortage, as though it were weather that happened to the market. But somebody still decides where scarce components go.

Follow the Money

Nvidia's data centre business generated roughly $185 billion in its most recent fiscal year. Gaming generated around $14 billion.

When supply is tight, that difference answers the allocation question. Nvidia does not need a policy to starve consumers. It only needs a spreadsheet. A component used in a gaming card cannot also become a far more profitable data centre product.

The allocation question answers itself

Nvidia revenue by segment ($bn)

Data centre Gaming

FY2026 is an estimate. Gaming has grown in absolute terms. It has simply become irrelevant to the allocation decision.

View as table

The result is thin supply and higher shop prices. Nvidia's flagship rose from £1,939 at launch to roughly £3,465. Retailers collect much of that increase, but Nvidia created the conditions and has little reason to change them.

Memory Is the Limit

The more revealing limit is not speed. It is memory.

Think of an AI model as a large set of books that must be opened on a desk before any work can begin. The card's memory is the desk. If the model does not fit, the card cannot run it properly, however fast the rest of the hardware may be.

The processing cores are the workers around that desk. More cores can perform more calculations at the same time, so they largely determine how quickly the model responds once it has been loaded. Memory bandwidth is the speed at which information moves between the desk and the workers. A card needs all three, but capacity comes first: a fast team cannot work on books it has nowhere to open.

That is what makes Nvidia's RTX 5090 consumer card and RTX PRO 6000 professional card such a useful comparison. They use the same underlying chip and have identical memory bandwidth. The professional card has 192 groups of processing cores compared with 170 in the 5090, a modest increase. Its memory, however, jumps from 32GB to 96GB.

The 5090 is not too slow for larger AI models. It simply has nowhere to put them. The PRO 6000 can load models the 5090 cannot, and that difference helps turn closely related hardware into products sold at very different prices.

Same silicon, three times the memory, nearly four times the price

Two cards built from the same core chip: memory capacity (GB)

Both cards use Nvidia's GB202 chip and have the same memory bandwidth. The RTX 5090 enables 170 of the chip's 192 processing groups; the PRO 6000 enables all 192. Retail prices are approximate.

View as table

The 32GB ceiling is not a manufacturing limit. It is a product decision. It makes the consumer card useful for local AI, but stops it loading the larger models that many businesses want to run. For those, buyers must move to a much more expensive tier.

Companies separate products by price and capability all the time. The important point is that local AI's ceiling was chosen in a meeting, not imposed by physics.

The Memory Makers Chose Too

The wider memory shortage is real, and Nvidia does not make memory itself. But this shortage also followed a commercial choice.

Samsung, SK Hynix and Micron moved production towards the high-bandwidth memory used in data centres because it earns better margins. Making it also consumes more factory capacity than ordinary computer memory. More supply for cloud AI therefore means less supply for everyone else.

DRAM contract prices across 2025+172%conventional DRAM
Memory as a share of PC build cost35%up from 15-18%
RTX 5090 UK street price vs launch+80%£1,939 to approximately £3,465

The consequences arrived quickly. Apple withdrew one high-memory Mac Studio option and raised the price of another. Nvidia also raised the price of its DGX Spark AI computer, citing supply.

The three memory manufacturers now face a US antitrust case alleging they deliberately restricted ordinary memory production. That is an allegation, not a finding. The argument does not depend on a conspiracy anyway. Scarcity is profitable, and nobody in the chain has a strong reason to end it quickly.

No Villain Required

Three groups made individually rational decisions. Nvidia sent scarce chips towards its most profitable market. It limited consumer memory to protect expensive professional products. Memory manufacturers prioritised the parts with the best margins.

Nobody needed to set out to kill local AI. Everyone with the power to make it more accessible simply benefited from not doing so.

That is why promises of more factory capacity from 2027 offer only limited comfort. Supply may improve. The incentives will remain.

The Cloud Benefits As Well

What follows is conjecture rather than documented strategy. The incentive is real and worth naming.

The companies building the largest AI models spend billions training them and recover that money by charging people to use them. That works best when AI runs in their cloud, with every request metered and billed.

A business running an open model on its own machine pays no usage fee. That local work is invisible to the cloud provider. Businesses are also learning to send fewer requests, use smaller models and keep more work on their own premises.

The contrast is stark. By 2030, training a frontier model may cost $18-38 billion. Producing a smaller model that is good enough for many everyday tasks may fall towards $5 million.

If a much cheaper model is good enough for routine business work, the expensive frontier models must recover their costs from a smaller group of premium tasks. We can already see that split: expensive models capture much of the revenue while cheaper open models handle much of the volume.

This does not mean AI laboratories are secretly suppressing local AI. They do not need to. Expensive local hardware suits the chipmaker, the memory suppliers and the cloud providers for different reasons.

Different motives. Same direction.

That makes the price of a desktop graphics card more important than it looks. It helps decide whether AI remains a service rented from a handful of large companies or becomes something businesses can own and run for themselves.

What Should Buyers Do?

Local AI is still worthwhile. The right choice depends on what you need.

For modest workloads, older hardware remains good value. A used RTX 3090 costs around $900 and can run capable smaller models. Software improvements mean these machines achieve far more than they did two years ago.

If capacity matters more than speed, there are options. AMD's Ryzen AI Halo and Nvidia's DGX Spark can hold models that a faster consumer graphics card cannot, although they run them more slowly.

What has disappeared is the middle: one reasonably priced consumer machine with both speed and enough memory for large models. That combination now costs professional money.

Renting also deserves an honest mention. If you use powerful hardware only occasionally, renting is far cheaper. If it runs most of the day, or your data cannot leave the building, buying still makes sense.

Local AI remains viable. It just stopped improving on its own, and started depending on decisions made by people who benefit when it doesn't.