NVIDIA on August 24, 2026 said its Vera Rubin NVL72 system sets a new efficiency standard for AI agents, delivering up to 30 times more work per watt than prior generations as agentic workloads explode in complexity.
Citing OpenRouter data, NVIDIA noted that agentic AI workloads consume about 15 times more tokens than a simple chat request, because an agent researching a company for an investment decision may query financial databases, search news and call multiple tools before answering. Efficiency therefore becomes the defining economic metric.
"AI factories run continuously, and their economics are defined by delivered output: tokens per second, tokens per watt, cost per token, utilization and uptime," NVIDIA said. The Vera Rubin NVL72 is engineered to maximize those ratios for always-on agent fleets.
The efficiency claim underscores how the center of gravity in AI infrastructure is shifting from raw peak performance toward sustained, energy-aware throughput as enterprises deploy agents at scale.





