Each generation of accelerators, whether GPUs, AI engines, or XPUs, consumes more power than the last, pushing the limits of what engineers have designed server architectures to support. In fact, a single inference request to a generative model today consumes ten times more energy than a conventional web search. Mu