Benchmarking LLM Inference at Scale with AIPerf
NVIDIA introduced AIPerf, a benchmarking tool designed to measure and evaluate large language model inference performance at scale.
Source: NVIDIA Developer BlogAI news about NVIDIA: product launches, model updates, and industry moves. Latest: Benchmarking LLM Inference at Scale with AIPerf
NVIDIA introduced AIPerf, a benchmarking tool designed to measure and evaluate large language model inference performance at scale.
Source: NVIDIA Developer BlogPawprint Studio's creature-catching game Aniimo is now available on NVIDIA's GeForce NOW cloud gaming service, alongside a path-tracing update for 007 First Light.
Source: NVIDIA BlogAI agents can be deployed to prepare and validate 3D scenes and digital twins for physical AI systems and simulation use cases.
Source: NVIDIA Developer BlogTensorRT Edge-LLM completed the MLPerf Edge Agentic Benchmark 6.4x faster on Jetson AGX Thor hardware, marking advancement in edge AI agent deployment.
Source: NVIDIA Developer BlogResearchers used agentic AI to translate CUDA tile operations from Python to Rust, leveraging AI to improve Rust-based GPU kernel development.
Source: NVIDIA Developer BlogNVIDIA's Vera Rubin NVL72 system achieved leading performance in MLPerf Inference v6.1 benchmarks, demonstrating advances in AI inference efficiency.
Source: NVIDIA BlogEmerald AI, Google, and NVIDIA announced the formation of the AI Energy Management Alliance to advance data centers with dynamic electricity management capabilities.
Source: NVIDIA BlogUniversity of Manchester uses NVIDIA Earth-2 to forecast air pollution across the UK.
Source: NVIDIA BlogNVIDIA CEO Jensen Huang presents at Salesforce Dreamforce, introducing Koa, Salesforce's first CRM reasoning model built on NVIDIA Nemotron 3 Super.
Source: NVIDIA BlogNVIDIA explains how mixture-of-experts models like Nemotron 3.5 Lightning activate only 3B parameters despite having 30B total parameters.
Source: NVIDIA Developer BlogNVIDIA discusses power management and efficiency optimization strategies for AI factories.
Source: NVIDIA BlogNVIDIA AI Infra Summit showcases AI factory efficiency advancements and tokens-per-watt optimization.
Source: NVIDIA BlogNVIDIA describes how Groq 3 LPX deterministic execution enables power-efficient inference on Vera Rubin.
Source: NVIDIA Developer BlogNVIDIA NVLink 6 provides multi-layer resiliency for large-scale AI factory operations.
Source: NVIDIA Developer BlogNVIDIA FLARE enables federated learning to scale across Docker, Kubernetes, and Slurm environments.
Source: NVIDIA Developer BlogA major children's hospital is utilizing open-source NVIDIA AI technology to enhance cardiac care capabilities.
Source: NVIDIA BlogNVIDIA announces NVLink Fusion with NVHBM technology to meet growing compute requirements for larger AI models and complex reasoning tasks in modern AI infrastructure.
Source: NVIDIA Developer BlogNVIDIA explores training robot navigation policies using cross-embodiment AI agents, enabling robots to convert perception into autonomous navigation capability.
Source: NVIDIA Developer BlogAlibaba released preview weights of the Qwen3.8-Flash-Next 176B model for developers to experiment with before the full Qwen4 architecture launch.
Source: NVIDIA Developer BlogNVIDIA introduces Shadow Engine Recovery for Dynamo, a technique to restore LLM inference capacity in seconds instead of performing lengthy cold restarts that require reloading weights and recompiling kernels.
Source: NVIDIA Developer BlogNVIDIA showcases RTX gaming support at Gamescom with new titles from publishers including EA, Embark, and Ubisoft, featuring improved anti-cheat technology and visual quality.
Source: NVIDIA BlogNVIDIA releases CUDA Python 1.0, providing stable APIs that enable Python developers to access GPU computing without requiring deep C++ expertise or complex build toolchain setup.
Source: NVIDIA Developer BlogNVIDIA Spectrum-X Ethernet infrastructure addresses data center design challenges for distributed AI model training spanning hundreds of thousands of GPUs.
Source: NVIDIA Developer BlogNVIDIA discusses XPU design for AI infrastructure, emphasizing that AI factories should be built as integrated systems optimized for tokens per second, energy efficiency, and cost per token—rather than collections of individual accelerators.
Source: NVIDIA BlogNVIDIA is extending its Vera Rubin NVL72 inference system with faster token generation capabilities to support agentic AI systems in full production.
Source: NVIDIA BlogNVIDIA's Vera Rubin NVL72 achieves up to 30x more work per watt for AI agents. The system addresses the fact that agentic workloads consume approximately 15x more tokens than simple chat requests due to activities like database queries and sub-agent invocations.
Source: NVIDIA BlogNVIDIA Vera Rubin and Blackwell platforms optimize performance per watt for agentic AI systems performing multi-step workflows with reasoning, tool invocation, and subagent coordination.
Source: NVIDIA Developer BlogNVIDIA BlueField-4 network accelerator enables new scale-in infrastructure architecture for agentic AI factories connecting diverse users and workloads.
Source: NVIDIA Developer BlogNVIDIA Vera CPU optimizes fleet economics in AI factories by improving the efficiency of converting power and capital investment into completed agent tasks.
Source: NVIDIA Developer BlogNVIDIA Groq 3 LPX is an interactive inference accelerator for the NVIDIA Vera Rubin platform, delivering ultrafast performance for long-context AI inference.
Source: NVIDIA Developer BlogNothing here yet.