To close the loop between dynamic environments and autonomous action, we must integrate the underlying physics of sensing directly with high-level cognitive reasoning. The goal of the Multimodal Agentic Perception and Learning (MAPLE) cluster is to develop multimodal AI agents capable of continuous, active perception and real-time physical adaptation. The team will co-design hardware-level sensors and foundation models into a unified, end-to-end inference framework to enable reliable, autonomous decision-making in high-stakes environments such as healthcare and environmental science. The team unites expertise across computer vision, natural language processing, machine learning and agents, and computational imaging.
Cluster Members
- PI: Guha Balakrishnan (ECE)
- Co-Is: Vicente Ordonez (CS), Chen Wei (CS), Ashok Veeraraghavan (ECE)
- Associates: Vivek Boominathan (ECE), Anshumali Shrivastava (CS), Hanjie Chen (CS)

-
Research Vision
-
The rise of multimodal foundation models and large language models (LLMs) has shifted artificial intelligence (AI) toward agentic systems capable of complex reasoning and generative tasks. However, most current agentic AI research treats perception as a static abstraction, relying on pre-packaged data, generic digital inputs, or off-the-shelf sensors. Real-world environments are far messier. They require agents that do more than passively receive data—they must actively control, adapt to, and understand the physical modalities through which they perceive the world. To truly close the loop between dynamic environments and autonomous action, we must integrate the underlying physics of sensing directly with high-level cognitive reasoning.
While established academic leaders exist in machine learning, foundation models, and agentic reasoning, the confluence of hardware-level sensing with agentic AI is an emerging frontier without a dominant institution. This research cluster will provide the resources to conduct joint projects, prepare grant proposals, and invite distinguished speakers, positioning Rice University as a leader in Multimodal Agentic Perception and Learning. We possess the critical mass of interdisciplinary expertise and commitment to make this a reality.
This cluster evolves from our previous institute-funded "Reactive Closed-Loop Computer Vision" group. Our primary pivot is to explicitly integrate additional data modalities, particularly language, alongside agentic AI workflows. To achieve this, our team unites expertise across computer vision (Vicente Ordonez, Guha Balakrishnan, Chen Wei), natural language processing (Hanjie Chen), machine learning and agents (Anshumali Shrivastava), and computational imaging (Ashok Veeraraghavan, Vivek Boominathan).
Independently, our members have advanced efficient visual reasoning, computational sensor design, generative modeling, and agentic workflows. Collaboratively, we have begun integrating foundation models with low-level physical sensor inputs, resulting in an NSF Future CoRE proposal submitted last fall (currently pending). Moving forward, our goal with this cluster is to unify our efforts into a cohesive, sensing-to-inference agentic framework. Over the coming year, we will focus on foundational developments and partner with domain experts to apply such a framework to at least one high-impact area, such as healthcare or environmental science.
