NVIDIA's Vera Rubin NVL72 achieves up to 30x more work per watt for AI agents. The system addresses the fact that agentic workloads consume approximately 15x more tokens than simple chat requests due to activities like database queries and sub-agent invocations.

According to OpenRouter data, agentic AI workloads consume 15x more tokens than a simple chat request. Why? Consider what happens when an AI agent researches a company for an investment decision. The agent queries financial databases, searches news and filings, invokes a sub-agent to run peer comparisons and model valuations, then synthesizes everything into a […]