Generative AI and Memory-Efficient Computing

The next phase of AI will be defined by systems that can reason longer, use more context, and allocate computation at test time without making inference prohibitively slow, expensive, or energy-intensive. The Sparse Intelligence Systems for Sustainable Reasoning: Memory-Efficient Computing Beyond the AI Memory Wall cluster will develop AI architectures that avoid unnecessary memory movement, making Rice a leader in memory-efficient AI systems that overcome the memory wall through learned sparsity—touching only the context and model memory each reasoning step needs to enable longer context, stronger recall and auditability, and deeper reasoning with far less energy.

Cluster Members

Anshu Slide

Research Vision

The next phase of artificial intelligence will be defined by systems that can reason longer, use more context, and allocate computation at test time without making inference prohibitively slow, expensive, or energy-intensive. Today, the limiting factor is increasingly memory movement. Modern LLMs repeatedly stream model parameters, attention state, and key-value context through constrained high-bandwidth-memory paths. Longer context windows and multi-step reasoning intensify the problem: the model spends too much time and energy moving bytes rather than performing useful computation.

This cluster will make Rice a leader in memory-efficient AI systems that overcome the memory wall through learned sparsity. Dense memory access is unnecessary for most reasoning steps. For any token, query, or intermediate reasoning state, only a small, input-dependent subset of prior context, attention heads, MLP blocks, experts, weights, and retrieval memory is useful. Rather than relying on full memory virtualization as the fundamental solution, this cluster will develop AI architectures that avoid unnecessary memory movement in the first place.

Research Thrusts

  • Learned sparse reasoning algorithms. Design token-adaptive mechanisms that select the active context, parameters, experts, and intermediate states needed for each reasoning step while preserving accuracy, recall, and in-context learning.
  • Systems for sparse execution. Build a runtime and evaluation stack that exposes memory traffic, cache behavior, bandwidth pressure, latency, and energy as first-class training and inference signals across GPU, CPU, and emerging memory-centric accelerators.
  • Hardware-aware and neuromorphic/in-memory pathways. Evaluate when sparse routing maps naturally onto in-memory, near-memory, analog, neuromorphic, or heterogeneous accelerator substrates, producing prototype kernels and design principles for useful reasoning per joule.

The expected impact is a foundation for sustainable reasoning systems: LLMs that can use longer context, think for more steps, and provide stronger recall and auditability without requiring a proportional increase in GPUs, energy, or cost. This would unlock deeper scientific reasoning, long-horizon enterprise analysis, provenance-aware responses, and local or sovereign deployment. The effort directly advances Rice priorities in Responsible AI and Sustainable Futures by attacking the energy and access bottlenecks that threaten to concentrate frontier AI in a few hyperscale infrastructures.

Rice is positioned to lead because the Ken Kennedy Institute already connects AI, systems, high-performance computing, statistics, signal processing, hardware, and translational partners. A focused cluster can convert these strengths into a national center-scale agenda: not making dense AI merely larger, but making reasoning efficient enough to be broadly deployable.