tech_surveillance2106 wordsRead on Arc Codex

How Data Center AI Can Keep Growing, Despite Supply Chain Bottlenecks

Efficiency is key for the next two to five years. But what happens when supply eventually catches up? Last month we discussed: Recently, all the major hyperscalers announced earnings and increased CapEx plans (partly due to rising costs). On Amazon’s July 30th earnings call, CEO Andy Jassy said demand is strong into 2028. But making back its investment takes a few years, so “as we get a few years out and the revenue growth outpaces the incremental CapEx growth, which will happen at some point, the resulting revenue, free cash flow, and return on invested capital is very compelling.” Like the Gold Rush, those who moved first and fastest seized the most and best supply capacity. The largest players, like Nvidia and Google, have massive capital and strong supply chain management. The biggest companies have likely locked in much of what they need. The number one strategy for coping with supply chain bottlenecks is efficiency — get the most results out of what you have and what you can get. But there are some other options, as well. Fig. 1: AI Growth faces high hurdles to growth (Source: ChatGPT prompt) If you are lucky, you may find companies that acquired more resources than they can use. SpaceX merged with xAI, whose models have not generated sufficient demand. Leading up to their recent IPO they did multiple AI infrastructure deals with their excess capacity: Meta recently announced it will sell/rent excess AI computing power, including a potential $10 billion deal with Anthropic. Customer usage of AI is growing, although still at early stages of wide adoption. Agentic AI brings more benefits, but uses many more tokens. For example, Allstate in December described how its virtual assistants are able to handle 4 of 10 customer calls, freeing up salaried staff. In companies I work with, I see numerous examples of high returns on modest AI spend. Frontier model companies are substantially increasing their efficiency. SemiAnalysis reports that Anthropic’s ARR (annualized revenue run rate) per MW (megawatt) is 4x higher in just 9 months. The best frontier models have immense capability, but they are overkill for many tasks. Using simpler and cheaper models is a good strategy for getting more results using fewer resources. For example, GPT-5.5 costs $30/million tokens; GPT-5.4 is half that at $15/million tokens. DeepSeek V4 is less than $1/million tokens. (Source: OpenRouter). The more expensive models give better results on the toughest tasks, but not all tasks need the best models. Coinbase CEO Brian Armstrong revealed on X in June that his company cut its internal AI spending nearly in half by setting the default model for its engineers to open models from China (Zhipu AI GLM 5.2 and Moonshot AI Kimi K2.7). OpenRouter, valued recently at $1.3 billion, provides a unified interface for a wide range of LLMs. You can use OpenRouter rather than developing your own internal model router software. OpenRouter’s LLM Leaderboard, as of late July, had Chinese models as 8 of the top 10 by tokens used (the other two are Nemotron 3 by Nvidia and Claude Opus 4.8 by Anthropic). Chinese models accounted for 30% of tokens used by U.S. firms since February, according to investment firm William Blair. There are other options to Chinese models. Thinking Machines, run by former OpenAI technology chief Mira Murati, released its first model recently, aiming at balancing cost and performance over raw power. The model has nearly 1 trillion parameters, but only 41 billion are “active,” so only a fraction of the model is used to deal with any query. That makes it cheaper and faster to use (Source: Wall Street Journal). It also has a customization tool, Tinker, to easily fine-tune the model for the customer’s specific needs. Another option for companies with significant volume and repetitive use cases is an SLM (small language model). LLMs are trained to do all sorts of things, including communicating in dozens of languages. Let’s say your application is a virtual assistant answering phone calls in English and Spanish, but doing some relatively simple tasks. You can set up an SLM that is trained just in what you need, which will result in a far smaller model size (number of weights), but do what you want just fine. An SLM might even be small enough to run on a local server or a future iPhone or Mac (Mac Studio with M7 Ultra in 2028 is expected to support models with more than 1 trillion parameters). Noeri is another new AI lab helping businesses customize AI models for their specific needs. Each new generation of Nvidia GPUs is much more efficient than predecessors. Anyone who can acquire Vera Rubin will get more tokens/$ and tokens/W for advanced workloads. Vera, the CPU, is optimized for agentic AI by Nvidia. GPUs are general-purpose and optimized for very large models. Not all applications need this. Nvidia paid $20 billion to license Groq’s technology, and Cerebras has gone public ($42 billion market cap as of late July) because it developed architectures without HBM, but with lots of SRAM, which works well with leading-edge models that need the highest responsiveness. AMD and Cerebras recently announced a cooperation deal. Google has been developing TPUs for a decade, and Amazon has been developing Trainium for years. Both are substantially ramping production. Amazon, on its recent earnings call, said Trainium’s revenue run rate exceeds $25 billion/year. Amazon claims its internal workloads get better throughput/$ and throughput/watt with its custom silicon. Both are now selling or planning to sell their custom silicon to external partners. Google has announced deals with Anthropic and Meta, and is talking to neoclouds. Amazon has announced deals with both Anthropic and OpenAI. Google is developing a chip, code-named Frozen v2, that hard-codes parts of the Gemini model into silicon, targeting release as early as 2028. The goal is to process 6x to 10x more tokens per watt than current hardware (Source: The Information). For the portion of Google’s total workload that runs on its Gemini architecture, this will dramatically cut cost and power. And this maximizes the compute throughput Google gets from its limited TSMC wafer allocation. OpenAI is developing its own AI accelerator using Broadcom. Anthropic is also rumored to be exploring doing its own chip. As the developers of the most popular frontier models, one can expect they’ll optimize their own chips to more efficiently run their models. The big players are developing their own CPUs, too (Nvidia Vera, Amazon Graviton, etc.), which are claimed to be 20% to 50% more efficient than x86 CPUs. There are numerous startups developing architectural variations on AI acceleration — FuriosaAI, Nuvacore, D-Matrix, Etched (which claims to have $1 billion of pre-orders), Tenstorrent, Sambanova, Mat-X, Positron, Majestic Labs, and more. They might have a hard time getting wafer allocations if they get traction, but they could be acquired by a big player with capacity. Coming soon is optical compute from Neurophos, in Austin, which has developed a special optical structure that is being integrated with an unnamed foundry on a CMOS process. Neurophos claims it will soon deliver superior AI compute performance, up to 250x that of Nvidia Blackwell 200 (see the chart below). The numbers are impressive, but we need to see real systems running. Fig. 2: Neurophos CEO presenting in San Francisco recently, comparing three of its optical processing units to the Blackwell 200 architecture. Another optical compute company, Lumai in the UK, uses lenses to do matrix multiply acceleration for AI inference. Lumai is running Llama today as a demo. Early optical compute companies failed because the cost and time to convert from the digital domain to optical, then back, exceeded the benefit of the optical speed. We’ll see if Lumai can move a large enough amount of continuous compute into the optical domain to overcome this. Everyone that has data center AI compute, memory, power, and lasers will make a lot of money while supply is short. But eventually (2028? 2029? 2030?), supply will catch up with demand for data center AI. What then? As Warren Buffett said, “You only find out who is swimming naked when the tide goes out.” The largest and financially strongest companies with the strongest hold over their customers are best positioned. Their access to the supply chain and their ability to lock in capacity is superior. These companies are: These companies have strong franchises, are tough to displace, and are very profitable. Except for AMD, these “fortress” companies have P/E ratios of 18 to 36, which are consistent with their expected rapid earnings growth. AMD has a much higher P/E ratio, well over 100, because investors expect they’ll grow their GPU share significantly over the next few years. All these companies are heavily investing in capital expansion and have raised their plans to keep up with demand. The leading AI custom silicon providers, Broadcom and Marvell, have seen big increases in market cap and now have P/E ratios in the 60 range (as of July 31). These companies do a fabulous job of executing and have tremendous IP, such as the high-speed SerDes required for AI compute, but they are the agents of the hyperscalers that own the cloud, own some of the models, and set the AI accelerator architecture. So far, these hyperscalers have gone with the biggest, best suppliers, but they have begun using other suppliers such as MediaTek either to cut costs, or perhaps because of TSMC wafer allocations. Consequently, there is a growing risk of commoditization as the main players have increasing vendor choice. And over time, it seems likely the hyperscalers will move downstream from architecture/design to silicon implementation, perhaps acquiring some of the smaller custom silicon providers. The stock market recently has seen 10x price increases on DRAM companies (Micron, SK hynix, Samsung), all of which have incredible demand with backlogs that extend for years. SK hynix made more profit last quarter than in the last five years. In recent weeks, these stocks have seen significant drops, though they are still far above their prices of a year ago. Yet they are trading at P/E ratios around 20 (as of July 31), which is similar to the “fortress” companies above, but much higher than normal because historically they have been in a boom-bust cycle. The DRAM companies may be able to get out of the commodity cycle by developing customized HBM optimized for each AI accelerator — the need for efficiency will result in AI accelerators tuned to certain workloads which may benefit from HBMs tuned for that architecture. This requires a business model shift from JEDEC commodity to custom ASIC. If the HBM becomes a customized component, this could make the DRAM business a high-value, differentiated business throughout the business cycle. Laser companies (Coherent, Lumentum) also have seen 10x stock price surges. Michael Hurlston, CEO of Lumentum, recently said on CNBC, “We are sold out for the next 5 years.” Both companies are trading at P/E ratios of well over 100 (as of July 31) because of the expectation of huge revenue growth as InP (indium phosphide) lasers are required in much higher volumes to power the shift from copper to optical interconnect, first in scale-out and then in scale-up. There are other manufacturers of InP, and though the technology is challenging, the cost and complexity of the fabrication process is much less than 2nm ASICs or HBM. The challenge for Coherent and Lumentum will be to develop highly differentiated lasers that add value to avoid a commodity price crash in two to five years, or to use their cash flow to develop or acquire related businesses, like Optical Circuit Switches, to get value-add and differentiation to soften the business cycles. Finally, the neoclouds, such as Coreweave and Nebius, have large valuations despite being unprofitable. They have secured allocations of value AI compute and have landed large contracts from the major players who are sucking up compute wherever they can get it. The risk for the neoclouds, when supply exceeds demand, is that the major hyperscalers will have installed huge amounts of their own capacity and will prioritize that over renting from the neoclouds. Data center AI growth will likely continue, despite severe supply shortages, because the biggest players have secured their supply; optimizing models to match need will reduce demand; and innovative architectures will get more performance from less silicon and memory. Players enjoying huge windfall profits from shortages need to plan for the time when supply again exceeds demand and endeavor to develop product offerings that are differentiated and sticky. Otherwise, they face price crashes when commodity products are in oversupply. Leave a Reply

How it works

Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.

Questions are cached — you'll always get the same 5 for this article.