tech_surveillance3249 wordsRead on Arc Codex

800VDC Pushes AI Power Design From Grid To Gate

The next power bottleneck is no longer just inside the accelerator — it is the full conversion path from medium-voltage AC to sub-1V silicon. Key Takeaways: The power issues facing AI data centers are spilling over onto individual dies, where engineers are working to minimize loss in the form of heat while maximizing the amount of power that makes it to the gates. As server racks climb from hundreds of kilowatts to megawatts, the efficiency of every conversion stage — from medium-voltage AC to 800VDC, and from there to sub-1V rails — will impact everything from performance and reliability to thermal margins. “The power on the server racks is around 125 kilowatts today,” said Peter Wawer, division president for green industrial power at Infineon Technologies. “The next step is going to be around 500 kilowatts. Then we will go to a 1 megawatt, 800-volt high-voltage DC-DC architecture moving forward.” The challenge now is how to improve the efficiency of the energy conversion from the AC grid to DC gates, and from high voltages to the very low voltages needed by GPUs and other AI accelerators. “Typically, you need to connect high-energy consumers like AI data centers to the medium voltage grid,” said Wawer. “In terms of electricity transmission lines, that means 35 KV. Today, you need a transformer to transform that high-voltage AC to low-voltage AC, and then to the AC-DC conversion.” Before the AI explosion, data centers included more steps in the power conversion process. “In the traditional multi-stage architecture, before AI, all the traditional data centers work off the 48V architecture, or 12V,” said Steven Lee, product manager of power electronics design software at Keysight Technologies. “Data center power demands were not nearly as high as the AI stuff is driving up. If you look at the different stages, from the power lines all the way to the chips, it goes from AC to DC, then you step down to a couple of volts, 1.2 or 0.7 volts, in the GPU. This was okay before AI.” New architectures reduce the number of steps. “As hyperscalers evolve from traditional 415VAC electrical architecture to 800VDC electrical architectures in the near future, the component suppliers to data centers/hyperscalers need to rethink the component design,” said Pavani Gottipati, director of applications engineering at Synopsys. “This includes moving to solid-state transformers, which can reduce the number of power conversion stages and the losses associated with each power conversion stage to create an energy-efficient data center.” 800VDC means power doesn’t need to go from A to B, B to C, C to D. “You can go from A to D really quickly,” said Lee. “With power converters, the fewer stages you have, the better the overall efficiency. The loss comes from heat, from the material switching losses and the like.” The AC portion of the 48V flow needed more steps, not the DC-DC part. “The architecture for these older types of data centers has a UPS (uninterruptible power supply) unit, and then they have a PDU (power conditioning unit), and then that goes to 48V,” Lee explained. “But in the new 800V architecture, those parts are condensed into this single AC-DC front end using silicon carbide, so these two stages essentially become one.” Fig. 1: Power delivery trends. Source: Keysight’s Lee, using an AI prompt At each energy conversion level, the architecture to deliver power needs to be changed. “It is required because you cannot physically get that amount of power into an IT rack when you stay at around 50V,” said Pradeep Shenoy, compute power technologist at Texas Instruments. “800V is needed because when you increase the voltage, you can decrease the amount of current for the same power level — they’re inversely proportional. So, even if you’re going to liquid-cool the bus bars running down an IT rack today, it wouldn’t be enough if you stayed at 50V. By necessity, 800V is definitely happening. It’s now a question of how to get ready for it.” In terms of 400V versus 800V, there may be little point in stopping halfway. “People are seeing that we need to move probably quicker than we think,” said Shenoy. “If you go to a bunch of intermediate steps, wait a year or two, you might say, ‘I need to be higher. I wish we had just gotten to 800V. It would have been better, because the power levels are increasing so fast.’ I’m even having some advanced conversations about what comes after 800V.” All power-delivery components must be upgraded when switching to 800VDC to safeguard against sparks and fires. “Ensuring safety is the paramount challenge for shifting from 48V to 800VDC to support AI workloads, particularly with respect to retrofitting existing data centers (brownfields) to be hybrid or fully 800VDC,” said Hoa Tram, senior principal product engineer at Cadence. “A data center retrofit to 800VDC touches nearly every layer of the power delivery stack. Existing protection systems and switchgear at these sites are designed and rated for AC fault interruption. That means all switchgear, circuit breakers, fuses, and disconnect switches in the 800VDC distribution path must be replaced with DC-rated equipment. The telemetry specific to the new 800VDC architecture will need to be integrated with existing monitoring and alarming systems. The personnel who maintain data centers will need to be properly trained and equipped to interact with all this new equipment.” Competing architectures add complexity. “Some designs use a single-polar 800VDC bus, while others favor a bipolar ±400VDC design,” noted Tram. “This fragmentation complicates interoperability, protection schemes, and workforce training and certification.” Further, some data centers may remain on the 48V architecture while benefiting from new power conversion processes. For example, Binghamton University engineers are developing a single-stage point-of-load converter for AI data centers that steps 48V power down near GPUs, with a lab prototype demonstrating 10% to 12% higher efficiency and twice the slew rate for faster power delivery. Benefits: more efficiency and density, less copper The 800VDC architecture offers several advantages over earlier 48V architectures. “Once you start getting things to DC rather than AC, then you can get a higher power density — more power in a smaller amount of real estate,” said Keysight’s Lee. “The main reasons to use 800VDC are higher efficiency, lighter power converters, and higher power density.” TI’s Shenoy agreed that efficiency is one of the main drawcards. “If you look at one conversion stage, for example, most of the power converters that are powering these processors have a 12V input, and they go down to maybe 1 volt or below,” he said. “Just by going to a 6V input you can save roughly 2% efficiency and you can get about 30% more power for the same area, which is incredibly important because there’s limited space on these circuit boards to fit all these power converters. Efficiency is very important. Power density is very important. Also at play is the overall robustness and reliability of these devices. You need to make sure that when you have hundreds of these chips being used to power this AI processor, you want no failures to take down that processor.” An efficient energy conversion step means less power is lost. “Anytime that we can be even one percentage point more efficient in this overall energy conversion chain is incredibly valuable,” Shenoy noted. “It has a meaningful OpEx dollar impact to the various service providers, hyperscalers, or whoever’s utilizing this equipment.” Further, 800VDC uses less copper. “When you have higher voltage and smaller current, then the wires can be thinner,” said Lee. “The old power conversion method would use too much copper in an AI data center. If we look at power — voltage times current — we need the same amount of power. And if your voltage is 48, that means your current has to be a ton more, so your wires have to be thicker. Otherwise, the wires will burn. They’re not going to be able to handle all the current, and with AI, the power demand is so high that the amount of copper would be tremendous. The wires would have to be very thick.” For both data centers and automotive applications, increasing the overall voltage architecture also benefits batteries used for both backup and primary power. “Today, most of the batteries are around 400V architectures,” said Puneet Sinha, senior director and global head of battery technology at Siemens EDA. “A lot of companies are looking at, and heavily investing, in innovating to go to 800V architecture. As you increase voltage, you get improved charging. But then that has implications on the electronics that you’re working with in the system — for batteries, but also other elements that you have in your vehicle or the data center system, such as inverters.” Solid-state SiC and GaN components For 800VDC architectures, solid-state transformers are replacing traditional iron-core transformers. “AI chips can pack multiple GPUs and CPUs and demand more electric power,” said Synopsys’ Gottipati. “As the power-hungry chips also sustain highly volatile and fluctuating power loads, advanced materials like SiC and GaN at the semiconductor design level can help achieve higher power densities for increased compute performance. However, due to the faster switching nature of these devices, the thermal and cooling system design becomes more important.” A solid-state transformer can be directly connected to a high-voltage AC grid. “Depending on the architecture you choose, the output can be converted down to low-voltage AC or immediately to DC — for example, 800VDC,” said Infineon’s Wawer. Fig. 2: Implementation of the 800V architecture. Source: Infineon One of the big design challenges in switching from 48V to 800V architectures is utilizing the new technologies, including wide-bandgap semiconductors such as silicon carbide and gallium nitride. “Those are slowly taking over traditional silicon-based types of technologies, and it’s good because, if you use it correctly, SiC and GaN can switch faster,” said Keysight’s Lee. “You don’t lose as much to heat. Your overall converter can be more efficient, and for these things, even 1% efficiency, if you cascade it right over many stages, adds up to quite a bit. That’s what designers are really shooting for — even single-digit efficiency improvements in higher conversion.” However, swapping out Si for SiC or GaN isn’t as straightforward as it might sound. “Traditionally, with silicon and older stuff, they switch at maybe kilohertz, and with slower switching speeds, you don’t have to worry too much about EMI, noise, and that kind of stuff,” said Lee. “You have to think about the faster-switching signals in the higher-frequency noise, outside of and into the environment and the gates. When you drive that gate, you have to use a different kind of signal. You need a good control loop. Otherwise, it won’t be a stable circuit, and that’s even more true at higher frequencies. Also, with silicon carbide and gallium nitride, your magnetics are different. You can use much smaller and lighter magnetics.” On the plus side, wide-bandgap materials offer some area benefits. “Traditional transformers took up a lot of real estate on the board, Lee explained. “Data centers are trying to stack as many servers on top of each other as possible, so they want things that are thin and small. When you start using new technologies such as GaN, and especially SiC, you can start substituting traditional transformers with much flatter components, so I can stack more on top of each other. Instead of having 12 units on a rack, if I can minimize my footprint, I can now have 20, which also allows you to utilize cleaner transformers.” Fig. 3: A planar transformer is more efficient and more easily stackable than the old wire-wound transformers. Source: Keysight Depending on the voltage class, Si, SiC, and GaN have some overlapping use cases. “It’s not either/or. What is best depends on the application need,” said Wawer. “On the high-power side for very high voltages, you have two materials, silicon and silicon carbide. Certain applications require a lower switching speed. For the lowest switching speeds, for example, an HVDC (high-voltage direct current) converter switching 50 hertz, maybe 60 hertz, will be an IGBT (insulated gate bipolar transistor), because it’s a very cost-efficient established technology, and it will not change much over the next 10 to 20 years. The solid-state transformer requires high voltage, fast switching speed, and low switching losses. Therefore, it’s the perfect solution for silicon carbide, and nothing else works there.” Other wide-bandgap components include digital hot-swap controllers such as SiC 750V JFETs (junction field-effect transistors) and 1,200V JFETs. GaN also becomes an essential material here due to its efficiency, density, and ability to operate from low to high voltage, Wawer explained. PMIC’s role in power conversion In advanced data centers, the role of power management ICs is also changing. “PMICs have shifted from being background components to becoming central enablers of system-level performance,” said Piero Bianco, senior director of product marketing for chips at Rambus. “AI applications require higher power densities, more precise voltage regulation, and feature more aggressive load transients. At the same time, components need to fit in a smaller form factor as AI servers pack more computing, networking, and memory capabilities, and require larger cooling solutions. These requirements push for a greater adoption of PMICs to generate power rails closer to their loads, in a more decentralized way. Decentralized, PMIC-based power solutions minimize conversion losses and provide more accurate voltage regulation while better supporting aggressive load transients. These solutions also offer accurate telemetry to manage the power needs of the system in real-time, eliminate the need for external sequencing components by integrating that feature, and use board real estate more efficiently.” As AI data centers are loaded with more training and inference workloads, there is increased pressure on all components, including PMICs and memory. “Data center memory places extra pressure on PMICs,” Bianco noted. “The main hot spot for PMICs — multi-output power chips with integrated controllers, MOSFET drivers, and power MOSFETs — is the memory subsystem. With the constant increase in DRAM memory speed, the requirements for PMICs, in terms of load current, precise voltage regulation, and load transients, are getting increasingly stringent. Meeting voltage regulation and transient specs is particularly important because voltage droop can cause memory timing violations, resulting in data corruption or other malfunctions in the server. These requirements drove the decision to move the memory power solution from the main board to PMICs in the memory module with DDR5, and continue to challenge silicon-level and application-level PMIC design.” PMICs are under the same kind of requirements or constraints that everything else is, Lee said. “How to pack as much power as possible into a small footprint. I see PMICs playing a big role in this whole AI data center evolution.” Efficient GPUs and accelerators also needed To achieve more energy-efficient data centers, progress is needed on all fronts. “One challenge is to have a more computationally efficient processor,” said Shenoy. “Whatever size area there is of this GPU, if they can get more compute into it, they’re going to do it, and that’s their value proposition. You can get 10X more tokens with our latest generation chip, or whatever. That will always be one vector, and that’s important.” Such is the demand to save power that GPU IP providers that traditionally served edge and automotive markets are seeing customers explore scaling their chips towards data center use cases. “We’ve seen our efficient technology being able to scale and grow to a large GPU size,” said Matthew Bubis, director of product management at Imagination Technologies. “The other side of it is that there are a lot of customers who want to build chips to try to compete with the likes of AMD, Intel, and Nvidia, so they need much more powerful GPUs. The shift is demand-driven due to our underlying efficiency, performance, and scalability. AI traditionally meant huge, giant, data center-level cards going everywhere. But we’re seeing demand — down at the very smallest size of the IP to the very largest — that they want an AI capability embedded within those GPUs, which means it’s a lot more matrix multiply technology embedded within the GPUs.” The demand for AI is not slowing down, so all energy conversion savings and chip innovations are vital. “Compared to two years ago, customers want more intelligence,” said John Weil, vice president and general manager for IoT and edge AI processor business at Synaptics. “An example is vision language models (VLMs). Two years ago, we were running a variety of vision models, and the LLM component wasn’t important at all. Now people are saying, ‘Wait a minute. Not only can I tell there’s a person in the field of view, but I can tell that person’s intent at this very high level.’ That shows you how AI is shifting very quickly, from using AI to improve the quality of service of a system to making deterministic decisions based on the data at hand and the model that’s deployed. That’s happened in less than two years.” Conclusion A lot of money and effort are being poured into the 800VDC architecture, solid-state transformers, and solid-state circuit breaker solutions to reduce AI data center power consumption. However, infrastructure and the grid are typically very slow-moving market industries. “Putting in ACDC connections in China is fast because it is top-down directed,” said Infineon’s Wawer. “But I’m living in Germany, and that is one of the slower countries. How those things are established, and the infrastructure being built up, should not be underestimated, as they can slow down the whole process. Also, if you have very bright innovations, you have to get on the street and install them.” Every innovation will soon be needed as data centers consume an order of magnitude more power than before. “You’re going from hundreds of watts to kilowatts,” said TI’s Shenoy. “If you add that up at the IC rack level, we’re going from tens of kilowatts or hundreds of kilowatts today to megawatts in the not-too-distant future. Then at the whole data center level, the tens of megawatts or hundreds of megawatts that might be in data centers today are going to the gigawatt scale.” Data center operators have to figure out how to power them. “How do we get energy?” said Shenoy. “How do we architect that whole power delivery chain? You have power generation somewhere, and wherever you’re getting energy from, it’s at a certain voltage level, which is in a different format than what the GPU wants.” [Editor’s note: A future article will discuss human safety measures and chip reliability needs in the 800VDC architecture] Related Articles Chip Innovation Will Bridge The Gap For USA Data Center Power GenAI demand is exponential; the US power grid can’t keep up; necessity spurs invention. AI Data Centers And Auto Industry Converge On Same Issues The EV revolution relies on battery innovation, while AI data centers need a range of new energy solutions to play nice with the grid. Both sectors are taking notes. Next-Gen Batteries Require Impedance Data And Active Balancing Battery management systems are growing increasingly smarter with innovations in software and hardware that enable more accurate estimation of battery state of charge and health, along with predictive diagnostics. Moving Electrons, Not Just Vehicles Why smarter charging, battery management, and power conversion are now the real differentiators in EVs and edge systems. Batteries Charge To The Edge Chemistry is becoming the determining factor across many markets as focus shifts to power. Leave a Reply

How it works

Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.

Questions are cached — you'll always get the same 5 for this article.