Background
As wearables, smart glasses, and service robots enter daily life, ideal AI agents should continuously perceive physical and social context, understand ongoing user activities, and proactively offer help at the right moment in the right way. Yet current systems mostly remain reactive to explicit commands, far from being seamless collaborators in everyday life.
Realizing this vision requires addressing three challenges. First, context understanding must be both broad and deep enough for agents to infer user intent from multimodal input across routine and unexpected events. Second, training and evaluation data are hard to obtain: real-world egocentric recording is costly and constrained by privacy, ethics, and safety, leaving rare high-stakes scenarios poorly covered, while physics-based simulators lack the visual fidelity for sim-to-real transfer. Third, intervention decision-making remains unformalized: when, why, and how an agent should intervene still lacks a unified framework, preventing systematic training and evaluation of proactive agents.
Research Objectives
First, to develop multimodal context understanding across diverse daily-life scenarios, enabling agents to recognize activities, infer intent, and assess intervention timing from egocentric video and audio. Second, to build a scalable, controllable, and high-fidelity data infrastructure for training and evaluation, overcoming the limitations of real-world data collection so that proactive agent research can cover rare and high-stakes scenarios. Third, to formalize a research framework for intervention decision-making with a unified taxonomy of agent autonomy, making when, why, and how to intervene systematically trainable and evaluable
Methods
This project develops multimodal reasoning models that integrate visual, linguistic, and temporal information, enabling agents to recognize user activities, infer intent and needs from continuous observation, and determine intervention timing and modality based on context. The research further explores long-term memory and personalization mechanisms, allowing agents to accumulate understanding of specific users and environments across extended periods of observation.
Innovation
First, we explicitly distinguish proactive and reactive assistance as two agent modes to develop proactive AI agents. Second, we develop multimodal integration mechanisms that resolve conflicts and noise across modal signals, improving agents' reasoning robustness in real-world contexts. Third, we incorporate intervention timeliness, safety criticality, and over-alert tendency as evaluation metrics, and further internalize them as a self-assessment mechanism for agents to apply before acting.
Expected Outcomes
(1) An automated scenario data generation system that produces egocentric video covering diverse daily-life and high-stakes events from abstract scenarios. (2) A multimodal reasoning model robust to cross-modal conflicts and noise, enabling stable context understanding in real-world environments. (3) Proactive AI agents capable of risk detection, intervention planning, and tool calling to assist users.