Research
We develop statistical learning and artificial intelligence methods for transforming complex, heterogeneous, and imperfect real-world data into reliable scientific knowledge and actionable decisions. Our research spans statistical machine learning, quality and reliability engineering, optimization, and data-driven systems science, with applications in healthcare, manufacturing, transportation, human behavior, and food systems. Across these areas, we focus on building intelligent systems that are personalized, interpretable, uncertainty-aware, and useful in practice.
Our work supports the national priority of preparing science for the Age of Intelligence. Research in multimodal data fusion, collaborative learning, digital twins, synthetic populations, and AI agents explores how artificial intelligence can integrate diverse evidence, simulate alternative interventions, and update models as new data become available. We view AI not as a replacement for scientific judgment and decision-makings, but as a tool that can strengthen discovery and decision-makings when combined with domain knowledge, human expertise, and real-world validation.
A second theme is Rigorous and Trustworthy Science. Real-world data are often noisy, incomplete, heterogeneous, and shaped by changing human and operational conditions. We develop methods for uncertainty quantification, anomaly detection, causal and longitudinal learning, privacy-preserving collaboration, and the identification of meaningful subgroups. These methods are intended to improve the transparency, robustness, and reproducibility of AI-supported research, particularly in high-consequence settings such as healthcare and engineered systems.
Another emerging direction in our research is the Participatory and Collective Intelligences for scientific and engineering intelligence. We study how distributed knowledge from experts, practitioners, patients, workers, and AI agents (i.e., wisdom of crowd) can be systematically collected, evaluated, and combined. Rather than assuming that useful knowledge must come from a single model or authority, we develop methods to identify consensus, structured disagreement, subgroup expertise, and uncertainty across diverse contributors. This work supports more inclusive and adaptive forms of data-driven enterprises, especially in settings where evidence is incomplete, expertise is dispersed, and no single source has a complete view of the problem.
Our research also contributes to the goal of reconnecting research with technological and industrial capability. Work in quality engineering, manufacturing analytics, system monitoring, digital twins, and operational decision support addresses the gap between developing an algorithm and deploying it reliably in practice. By studying AI within clinical, production, mobility, and supply-chain environments, we aim to support more effective experimentation, system improvement, and resilient operations.
Finally, we seek to ensure that scientific and technological progress responds to human and societal needs. Applications in chronic disease management, aging, mobility, food access, and personalized services are developed through collaborations across engineering, medicine, public health, industry, government, and community organizations. This collaborative approach emphasizes on cross-sector research, and the translation of research progress into broadly shared benefits.
Taken together, our research aims to help build an AI-enabled scientific and engineering ecosystem in which trustworthy data, rigorous methods, and digital models support discovery, practical innovation, and public benefit.
Highlights of Our Recent Work
- CrowdLLM: We are building LLM-based digital/synthetic populations, enabling scalable virtual experimentation and more personalized decision support with applications in healthcare, innovations, population-based research. (2026 INFORMS QSR Best Paper Competition, Honorable Mention; forthcoming at INFORMS Journal on Data Science)
- Team, Then Trim: We use teams of specialized LLM agents to generate realistic, task-relevant tabular data, helping address data scarcity, imbalance, and the high cost of real-world data collection particularly in healthcare. (2026 IISE DAIS Best Student Paper Competition, 1st Place.)
- Consensus Discovery: We develop methods that uncover meaningful agreement, disagreement, and expertise within crowds, creating a foundation for more reliable human–AI collective intelligence. (2026 IISE QCRE Best Student Paper Competition, 1st Place).
- ALARM — We integrate multimodal LLMs with uncertainty quantification, self-reflection, and ensemble reasoning to detect ambiguous, context-dependent anomalies and support trustworthy monitoring in smart homes, healthcare, and other complex environments.
- Causal Discovery through the Wisdom of Crowds: We develop methods that combine distributed knowledge from experts, practitioners, and AI agents to guide causal structure learning, reduce the search space, and make causal discovery more transparent and scientifically grounded.
Acknowledgment for Funding Support
- National Science Foundation (CMMI-1505260, CMMI-1536398, CMMI-1824623, CCF-1715027, CIS-2114260)
- Breakthrough T1D (formerly known as the Juvenile Diabetes Research Foundation)
- NIH
- AHRQ
- DARPA WASH; DARPA D3M
- AFOSR DDDAS
- USDOT
- Byrd Alzheimer’s Institute
- Royalty Research Foundation
- Helmsley Foundation
- Amazon
- Meta