Without major advances in computing infrastructure, next-generation AI risks becoming prohibitively expensive and accessible only to a small number of organizations. The AI Systems and Hardware Co-Design for Next-Generation Foundation Models cluster is building the computational foundations to make frontier AI more efficient, sustainable, and broadly deployable by re-inventing how hardware and software work together. Their work pursues a full-stack co-design of software systems and hardware accelerators that speak the language of AI natively and can run AI models smoothly across distributed networks.
Cluster Members
- PI: Yuke Wang (CS)
- Co-Is: Wei Qiu (ECE), T. S. Eugene Ng (CS)
- Associates: Tong Geng (ECE), Jiarong Xing (CS), Hanjie Chen (CS), Hengrui Luo (Statistics)

-
Research Vision
-
Foundation models, including large language models, multimodal AI, diffusion models, and emerging scientific foundation models, are transforming science, healthcare, education, engineering, and industry. They are becoming the computational engines of future discovery and intelligent decision-making. However, their progress is increasingly constrained by soaring training and serving costs, memory and energy demands, communication bottlenecks, hardware underutilization, and growing concerns around robustness, privacy, and trustworthiness. Without major advances in computing infrastructure, next-generation AI risks becoming prohibitively expensive and accessible only to a small number of organizations.
Our long-term objective is to establish Rice University as a national leader in AI systems and hardware co-design for foundation models by creating computing foundations that make frontier AI dramatically more efficient, scalable, reliable, and broadly deployable. Rather than treating algorithms, systems, hardware, and reliability as separate layers, we pursue a full-stack co-design approach that jointly optimizes them as an integrated system.
We will advance this vision through three tightly coupled thrusts. Thrust 1: Scalable Systems for Foundation Models develops next-generation software systems for efficient training and serving across distributed clusters. We will study workload bottlenecks in LLMs, multimodal, diffusion, and scientific AI models, design adaptive parallelism strategies spanning tensor/pipeline/expert dimensions and build memory-efficient inference systems through KV-cache optimization, offloading, speculative decoding, and dynamic batching. Thrust 2: Hardware, Compiler, and Runtime Co-Design converts these bottlenecks into optimized execution stacks across modern accelerators. We will develop accelerator-aware compilers, learned runtime schedulers for GPUs/TPUs/Trainium/NPUs/heterogeneous systems, and custom kernels and architecture co-design for attention, MoE routing, sparse computation, and future chiplet-based platforms. Thrust 3: Sustainable, Reliable, and Trustworthy AI Infrastructure ensures these systems can be deployed responsibly at scale through energy-aware scheduling, resilient distributed execution, privacy-preserving inference, anomaly detection, uncertainty calibration, and confidence-guided fallback mechanisms.
Successfully achieving these objectives will lower the cost and energy barriers of foundation models, democratize access to frontier AI capabilities, and enable trustworthy deployment in science, medicine, education, and national infrastructure. It will also position Rice as a hub for next-generation AI computing and train students across computer science, electrical engineering, and statistics to lead the future AI workforce.
Our team is uniquely positioned to lead this effort through complementary expertise rarely found within one institution. Yuke Wang contributes leadership in scalable AI systems, GPU acceleration, distributed machine learning, and trustworthy AI infrastructure. Tong Geng adds strengths in computer architecture, efficient AI accelerators, and hardware–software co-optimization. T. S. Eugene Ng brings deep expertise in energy-efficient reconfigurable interconnection cluster networks and energy-efficient LLM inference optimized for mobile NPU accelerators. Jiarong Xing contributes cloud systems, GPU resource management, efficient LLM serving, and agentic AI systems. Hanjie Chen strengthens the team in NLP, interpretable machine learning, and LLM reasoning. Hengrui Luo provides expertise in statistics, uncertainty quantification, and reliable decision-making. Wei Qiu contributes leadership in AI for biomedicine and healthcare applications. Collectively, this breadth across systems, hardware, networking, machine learning, statistics, and domain AI applications enables Rice to pursue full-stack co-design at a depth few institutions worldwide can match.
