AI inference is pushing data-centre design beyond raw computing power
MIT Technology Review’s Insights arm says continuous AI services are making memory, storage, networking and data movement central to infrastructure planning.

The growing use of continuous, distributed AI inference is prompting enterprises to rethink data-centre infrastructure around the interaction of compute, memory, storage and networking, according to an analysis published by MIT Technology Review’s Insights arm.
The article identifies data movement, latency, memory bandwidth, caching and storage throughput as key constraints, particularly for systems that retrieve and process large volumes of information in real time. It contrasts these demands with earlier, more training-centric deployments.
Jim McGregor, founder and principal analyst at Tirias Research, says organisations need to design all four infrastructure layers together. He argues that simply selecting the fastest processors is insufficient when bottlenecks can shift between compute, memory, storage and networking.
The analysis says infrastructure planning must also weigh efficiency, cost, scalability and flexibility. Applications discussed include healthcare, robotics, financial services, customer-facing AI and edge devices, where delays may affect responsiveness, safety or trust.
It also stresses that workloads and technology are changing rapidly, making adaptability a central design consideration. The article does not quantify the cost, scale or timeframe of any industry-wide redesign, and its conclusions are analysis and expert commentary rather than evidence of a specific enterprise deployment.
The material was produced by MIT Technology Review’s custom-content arm, rather than its editorial staff.

