NVIDIA 讨论了用于 AI 基础设施的 XPU 设计,强调 AI 工厂应作为集成系统构建,优化每秒令牌数、能源效率和每令牌成本,而非单个加速器的组合。
To generate intelligence at scale, AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime. That requires AI infrastructure designed and built as a full factory, not a collection of individual accelerators. Hyperscalers and AI-native companies building custom XPUs must consider […]