AICoE Project

Next-Generation Human-Centered Multimodal AI Agents with Verifiability, Controllability, and Inclusiveness

Project Name

Next-Generation Human-Centered Multimodal AI Agents with Verifiability, Controllability, and Inclusiveness

Project Goal

This research is based on multimodal agent platforms aiming to significantly enhance performance, reliability, trustworthiness, explainability, risk management, and practicality. Key focus areas are: 1.        Complex Task Reasoning: To develop reasoning frameworks, which can handle intricate, real-world scenarios. 2.        Model Output Controllability: Investigating fundamental theory of AI model interpretability to develop highly controllable LLMs. 3.        Native Visual Understanding: Developing native visual reasoning MLLMs that transcend current 'image-to-text' limitations, enabling direct visual thinking and logical deduction for tables, charts and 3D spatial environments. 4.        Speech and Emotional Intelligence: Developing speech and emotion perception to facilitate human-like interaction between AI agents and users, with applications in psychological and psychiatric disorder analysis. 5.        Human-AI Collaborative Mechanisms: Designing and testing innovative human-AI collaboration frameworks to ensure AI generates concise, explainable outputs while preventing the learning of flawed decision-making processes. 6.        Validation in Clinical Contexts: Verifying the practicality, efficacy, and safety of AI agents within real-world healthcare settings.


Project Description

This goal of the project is to develop next-generation multimodal AI agents, which are human-centric, verifiable, controllable, and inclusive. Although Large Language Models (LLMs) are very powerful, there are still some challenges need to be conquered, including reasoning gaps, output control, hallucinations, and cross-modal causal inference. Our initiative enhances agent capabilities through five pillars: •        Controllable Multi-Agent Systems: To tackle the challenges in intension ambiguity, declining interpretability, and training complexity by developing proactive demands clarification and dedicated agent training mechanisms. •        Native Visual/Multi-modal Reasoning: Advancing native reasoning in Multimodal LLMs (MLLMs) to integrate visual and textual thinking. Applications include video learning, embodied AI, and open-vocabulary 3D point cloud segmentation, bypassing the constraints of traditional "image-to-text" methods. •        Inclusive Speech AI Agent: Addresses speech heterogeneity and minority adaptation. By enhancing emotion perception and empathetic response generation, we also plan to explore some applications in emergency reporting and social behavior analysis for autism. •        Verifiable Human-AI Collaboration: Extracts precise, personalized multimodal evidence from reasoning LLMs to bolster trust in high-risk domains and buildup standardized AI evaluation benchmarks. •        Automated Disease Tracking and Management: To develop AI agents with temporal awareness, multimodal integration and reasoning capabilities, and operate on common Electronic Medical Record (EMR) platforms, ensuring interoperability with TWCDI and FHIR standards. By integrating innovative research, this project aims to enhance Taiwan’s global visibility and technical self-sufficiency in multimodal AI agents, driving the advancement of AI-powered healthcare applications.