New Memory Technology Could Dramatically Reduce AI Energy Demands
A new memory technology called SOT-MRAM could dramatically reduce the energy demands of artificial intelligence by performing key computing tasks 5 to 100 times faster while consuming a fraction of the power of existing memory chips. Researchers at the University of Texas at Austin partnered with Taiwan Semiconductor Manufacturing Company (TSMC), the world's largest semiconductor foundry, to fabricate and test this emerging technology on various AI workloads, from neural network inference to probabilistic modeling.
What Makes This Memory Technology Different?
SOT-MRAM, which stands for spin-orbit torque magnetic random-access memory, uses magnetic properties to store and retrieve data rather than the electrical methods used in conventional memory chips. The technology can retain information even when power is turned off, making it fundamentally different from the volatile memory that powers most AI systems today. In testing, the chips completed write operations in just 2 nanoseconds while consuming only 2 picojoules of energy per write. By comparison, other memory technologies require between 5 to hundreds of times longer, sometimes taking several milliseconds, and consume hundreds of picojoules or more.
The breakthrough lies in how researchers adapted SOT-MRAM for AI applications. Traditionally, this memory type has been overlooked for artificial intelligence because it can only hold two states, 0 or 1. The research team designed a novel approach that takes advantage of this binary limitation while maintaining the accuracy needed for real-world AI tasks.
"The unique combination of speed, energy efficiency and endurance makes SOT-MRAM perfectly suited for AI applications, especially in devices where resources like power and memory are limited," said Sam Liu, first author of the new paper published in Science Advances and recent UT Austin Ph.D. graduate.
Sam Liu, First Author and Recent Ph.D. Graduate, University of Texas at Austin
How Could This Technology Transform Edge AI Devices?
The implications for on-device inference are significant. Edge devices like sensors, robotic hands, and wearables currently struggle with power constraints that limit their ability to run sophisticated AI models locally. SOT-MRAM could enable these devices to perform complex AI tasks without constantly transmitting data to cloud servers for processing. Consider a robotic hand equipped with heat sensors; with SOT-MRAM-powered AI accelerators, the hand could make quick, accurate decisions about movement without sending signals back to a central processor in the cloud.
The researchers tested the technology on multiple AI workloads to validate its potential across different use cases:
- Neural Network Inference: Running trained AI models to make predictions on new data, the most common AI task on edge devices.
- Binary Neural Network Training: Teaching simplified AI models that use only 0s and 1s, which align naturally with SOT-MRAM's binary states.
- Probabilistic Graph Modeling: Processing complex relationships in data, a technique used in recommendation systems and knowledge graphs.
Why Does This Matter for the Energy Crisis?
The urgency behind this research is real. Data centers powering AI have driven rapid increases in energy demand across Texas and worldwide. Texas is expected to become the U.S. capital of data centers within the next few years, and that growth could lead to a 5-fold increase in statewide energy usage. There are two primary strategies to address this crisis: make AI technology more efficient, and reduce reliance on centralized data centers. SOT-MRAM addresses both simultaneously.
"We show that SOT-MRAM AI accelerators can provide the energy efficiency, with enough accuracy, to eventually replace CPU-based AI accelerators in edge devices such as sensors," explained Jean Anne Incorvia, associate professor in the Cockrell School of Engineering's Chandra Family Department of Electrical and Computer Engineering and the faculty leader on the project.
Jean Anne Incorvia, Associate Professor, Cockrell School of Engineering, University of Texas at Austin
Incorvia also outlined a practical hybrid approach: when very high accuracy is needed, edge devices can still connect to GPU-based data centers in the cloud. This flexibility means developers don't have to choose between local processing and cloud power; they can use SOT-MRAM for routine, latency-sensitive tasks and offload only the most demanding workloads to centralized infrastructure.
Steps to Implement Hybrid Edge-Cloud AI Strategies
Organizations exploring on-device inference with emerging memory technologies should consider these practical approaches:
- Assess Latency Requirements: Identify which AI tasks need immediate local responses versus those that can tolerate cloud round-trip delays, then prioritize edge deployment for time-sensitive operations.
- Evaluate Power Constraints: Measure the power budgets of target devices like sensors, wearables, and robotic systems to determine whether SOT-MRAM's energy efficiency gains justify architectural changes.
- Plan for Accuracy Trade-offs: Test binary neural network approaches on your specific AI models to understand accuracy impacts, since SOT-MRAM's binary states require adapted training methods.
- Monitor Technology Maturity: Follow developments in SOT-MRAM fabrication and device variation reduction, as these refinements will determine when the technology reaches production readiness.
What's Next for SOT-MRAM Development?
The research team is continuing to refine the technology to address remaining challenges. Key priorities include optimizing the characteristics that enable the chips' speed and efficiency, and reducing variation between individual devices, which can degrade neural network accuracy. These refinements will be critical before SOT-MRAM can be deployed at scale in consumer and industrial devices.
The partnership between academic researchers and TSMC demonstrates how cutting-edge memory innovation could reshape where AI actually runs. Rather than pushing all computation to distant data centers, technologies like SOT-MRAM could enable smarter, more responsive devices that think locally and only reach out to the cloud when necessary. For industries from robotics to healthcare to agriculture, that shift could mean faster decisions, lower latency, reduced bandwidth costs, and ultimately, a more sustainable AI infrastructure.