NVIDIA discusses XPU design for AI infrastructure, emphasizing that AI factories should be built as integrated systems optimized for tokens per second, energy efficiency, and cost per token—rather than collections of individual accelerators.

To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime. That requires AI infrastructure designed and built as a full factory, not a collection of individual accelerators. Hyperscalers and AI-native companies building custom XPUs must consider […]