AI efficiency jumped 18x in 16 months, Stanford research finds

1 hour ago 1



Running a large language model locally used to feel like heating your apartment with a space shuttle engine. New research out of Stanford’s Hazy Research group suggests that equation has changed dramatically: the intelligence you can squeeze out of a single joule of energy has improved 18-fold in roughly 16 months. The finding comes from a paper titled “Measuring Intelligence Efficiency of Local AI” (arXiv:2511.07885), which introduces two new metrics designed to benchmark local AI inference: Intelligence per Joule (IPJ) and Intelligence per Watt (IPW). Think of IPJ as the miles-per-gallon rating for AI models running on your own hardware rather than in a distant data center. Where the 18x came from The 18x improvement in IPJ from mid-2024 to late 2025 didn’t come from a single breakthrough. It was a compound effect. Model architecture improvements contributed roughly 3.1x of the gain. Hardware and accelerator advances delivered a much larger 5.9x boost. On the model side, the research tracked a progression from earlier architectures like Mixtral-8x7B to newer mixture-of-experts variants and models like gpt-oss-120b. These designs route computation more selectively, activating only...

Read Entire Article