AICoE Project

Embodied Intelligence from Multimodal Reasoning for Spatial-Temporal Understanding, Alignment, Self-Reflection and Embodied Reasoning

Project Name

Embodied Intelligence from Multimodal Reasoning for Spatial-Temporal Understanding, Alignment, Self-Reflection and Embodied Reasoning

Project Goal

This project proposes a four-year research program to advance general-purpose intelligent robots capable of reasoning, self-correction, and adaptive learning. Targeting key limitations of existing Vision–Language Models (VLMs)—including hallucination, weak reasoning, and limited transferability—we develop an integrated roadmap spanning model grounding, policy optimization, and embodied reasoning.  The project progresses from hallucination-aware spatiotemporal alignment, to self-reinforced policy learning from interaction videos, to structured reasoning for long-horizon planning, and finally to learning transferable task logic from human demonstrations within a unified Vision–Language–Action framework. This research aligns with international AI–robotics trends and is expected to significantly enhance Taiwan’s research competitiveness in intelligent robotics.


Project Description

Amid the rapid advancements in artificial intelligence and robotics, the development of general-purpose robots capable of autonomous understanding, multimodal reasoning, and flexible adaptation to complex real-world tasks has become a shared goal across global academia and industry. Notable examples include Google's Gemini Robotics and NVIDIA's GR00T-N1. In recent years, the continuous progress of vision-language models (VLMs) has enabled their deployment on robots, equipping them with preliminary semantic understanding, visual recognition, and reasoning capabilities. However, current models still face significant challenges in spatial-temporal perception, long-term task planning and self-correction, and learning skills directly from human demonstration videos. These issues—such as hallucinated descriptions, inaccurate integration of visual and action information, and lack of self-learning and reasoning in novel environments—are particularly pressing in key applications like smart manufacturing and service robotics.  In light of these challenges, this four-year project aims to begin with enhanced vision-language models and gradually progress toward building general-purpose robots capable of reasoning, self-correction, and human-level skill learning. In Year 1, we first focus on addressing hallucination in spatio-temporal perception within multimodal VLMs using self-enhanced contrastive learning, thereby strengthening the model's grounding in reality. In Year 2, we start to build upon these results to apply the improved VLMs to robot decision-making and reasoning, incorporating self-improvement strategies to reduce task execution errors. In Year 3, we further introduce structured reasoning and reinforcement learning, enabling the robot to integrate visual-language-action reasoning and acquire visual self-correction capabilities. Finally, in Year 4, we make a further leap by developing an embodied reasoning framework that allows the model to comprehend and reason about high-level task logic from human demonstration videos, and transfer that understanding to the robot's own strategy learning. This phased roadmap not only aligns with the latest international developments but also propels Taiwan’s research and practical advancements at the frontier of general-purpose intelligent robotics.