AICoE Project

Integrating External Knowledge and Physically-Grounded Multimodal World Models for Multi-Agent Robotic Control

Project Name

Integrating External Knowledge and Physically-Grounded Multimodal World Models for Multi-Agent Robotic Control

Project Goal

This project focuses on constructing a comprehensive Cyber-Physical AI (CPAI) system architecture. By pursuing the following research pillars, the project aims to enhance system flexibility and application scope, ultimately realizing a highly adaptive and versatile intelligent autonomous platform—specifically, Intelligent Multi-Agent Control Technology that integrates external knowledge with physical multi-modal world models. (1)        Learning, Compression, and Generation of visual representations. (2)        Physical World Models for understanding external knowledge. (3)        3D Scene Cognition and reconstruction. (4)        Robot Reinforcement Learning integrated with 3D world representations. (5)        Autonomous Intelligent Collaborative Defense Strategies for multi-agent UAVs in airspace security.


Project Description

In recent years, the trends of declining birth rates and an aging population have intensified, leading to an increasingly severe labor shortage. As a result, automation technologies that can replace human labor have become a key direction for future development. In response to this structural shift, the government is comprehensively promoting industrial upgrading under the premise of "industrial AI transformation.” Driven by technological advancements, autonomous mobility systems related to unmanned vehicles have developed in diverse directions. These include various platforms such as drones, self-driving cars, robotic dogs, and robots, and have emerged as viable solutions to the labor shortage problem. To reduce the cost of testing training models and to enhance adaptability to different external environments, the observation of international and industrial development trends indicates that multi-agent systems, multimodality, physical artificial intelligence, and retrieval-augmented generation (RAG) have become the main development directions for embodied intelligence machines. Based on the achievements of the previous two AI projects of NTSC of CVRC NYCU, in this stage, we will construct a virtual-physical artificial intelligence ecosystem based on these current development directions. In order to meet the rapidly changing needs of industry and enhance the breadth and flexibility of application, the system architecture proposed in this project will not be limited to drone platforms, but will be extendable to various types of autonomous mobile carriers.  The project will be divided into five major themes: 1) Visual Representations Learning, Compression, and Generation, 2) Physical model with external knowledge understanding, 3) 3D Scene Perception and Reconstruction, 4) Reinforcement Learning for Robotic Agents in 3D World Representations, and 5) Cooperative Intelligent Defense Strategies for Multi-Agent UAV Systems. Through these five research themes, the project aims to establish a minor and comprehensive system architecture for virtual-physical artificial intelligence, ultimately realizing a highly adaptable and widely applicable intelligent autonomous platform.