Trustworthy AI

The Foundations of Trustworthy AI: Advancing Responsible AI through Privacy, Fairness, & Alignment (TRUST-AI) cluster aims to build a rigorous, theory-driven foundation for “AI Alignment” to design trustworthy and responsible AI systems whose behavior reflects human goals, values, and welfare. This work moves from benchmark settings to high-stakes social environments where alignment and contextual integrity is a central priority. The TRUST-AI team combines expertise in statistics, machine learning, economics, decision theory, fairness, privacy, and preference modeling.

Cluster Members

Maryam Slide

Research Vision

The first year of the cluster established a coherent research program around an increasingly important premise of trustworthy AI: privacy, fairness, collaborative learning, and preference learning, and robust data analysis and its applications to quantum computing. In the second year, we will build on this foundation and expand the cluster’s agenda toward AI Alignment: the problem of designing AI systems whose behavior reflects human goals, values, and welfare. We will focus especially on preference alignment, contextual integrity, and application-facing methods for deploying modern generative AI systems when human goals are heterogeneous, only partially observed, or strategically reported. This is a natural progression of our research agenda, in which a common statistical and conceptual core can be applied to a rapidly growing domain. It also reflects the direction of the broader field: as large language models and generative AI systems move from benchmark settings into high-stakes social environments, alignment has become a central priority for researchers, industry labs, and funding agencies. Here are two concrete example project that we will pursue:

Preference Alignment: A major direction will build on the Co-PI’s recent work on preference alignment. Large language models are commonly aligned with human preferences using pairwise- comparison data and methods such as reinforcement learning from human feedback. Standard approaches often pool binary choices from many annotators as if they reflected a single underlying preference. This abstraction is convenient but can be misleading when users have heterogeneous preferences: recent work shows that choice-only data may fail to identify even the population-average preference. To address this problem, the Co-PI and coauthors study non-choice data—inexpensive behavioral signals that enrich standard preference datasets—and show that response-time data can be combined with choice data to recover average preferences in heterogeneous populations [Echenique et al., 2025, 2026]. This approach yields consistent estimators without requiring user tracking or the assumption that all users share a common reward function.

Contextualized Integrity/Privacy: A second-year project will study contextualized privacy auditing for large language models. Di!erential privacy o!ers strong worst-case protection, but can be too stringent for LLM applications and may reduce utility. We will develop a complementary approach focused on practically meaningful leakage: cases where a model makes sensitive information more recoverable than it would be without privileged data. This can be tested through controlled audits, such as implanting synthetic “canaries” into training or retrieval data and measuring whether the model later reveals or uses them in inappropriate contexts. It can also be studied through an odds-ratio lens, asking whether model interaction substantially increases confidence about a sensitive attribute beyond what public information would imply. The goal is to build auditors and mitigation tools that detect context-inappropriate information flows while preserving utility.

The cluster has been intentionally organized to pursue this agenda. The team combines expertise in statistics, machine learning, economics, decision theory, fairness, privacy, and preference modeling, and we are extending this capacity through external collaborations with researchers working on AI alignment, human-centered AI, and preference learning. This structure allows the cluster to address alignment at multiple levels: developing the theoretical foundations needed to understand heterogeneous feedback, designing estimators and algorithms suitable for finite and noisy data, and translating these insights into methods that can inform practical AI alignment pipelines. In this sense, the second year will mark a shift from studying isolated dimensions of trustworthy AI toward building a unified, application-facing framework for AI systems that are fair, privacy-aware, and aligned with pluralistic human goals.