AICoE Project

Medical Model Beyond Scaling: A Modular Foundation Model Framework for Data-Efficient and Trustworthy Medical Imaging

Project Name

Medical Model Beyond Scaling: A Modular Foundation Model Framework for Data-Efficient and Trustworthy Medical Imaging

Project Goal

This project aims to build Taiwan’s first modular, multi-task, and multi-modality training framework for medical imaging foundation models, fundamentally challenging the dominant “scaling law” paradigm. We advocate a “strategy over scale” approach that achieves strong generalization under constrained medical data conditions. Technically, we develop the MMBS (Medical Model Beyond Scaling) methodology by fusing two complementary self-supervised paradigms, MAE (global reconstruction) and DINO (local semantic alignment), and coupling them with a Transformer-style decoding design to support detection, segmentation, and classification within a unified inference pipeline. We further integrate multi-stage fine-tuning (MSFT), intermediate supervised fine-tuning (ISFT), and parameter-efficient fine-tuning (PEFT) to improve transferability and robustness across domains. Validation will start from fundus imaging and progressively extend to gastroscopy and echo, in close collaboration with clinical teams at Chang Gung Memorial Hospital, and will be consolidated into a cross-domain medical imaging AI toolkit to facilitate reproducible research and practical clinical translation.


Project Description

Medical AI is constrained by fragmented, privacy-sensitive datasets and the high cost of expert annotation, making brute-force scaling impractical. This project proposes MMBS (Medical Model Beyond Scaling), a strategy-driven framework that builds robust medical imaging foundation models without relying on billion-scale datasets. Our core idea is “strategy over scale”: we systematically combine algorithmic innovations across representation learning, fine-tuning, and modular inference to maximize transfer under limited medical data. At the representation level, we fuse two complementary self-supervised learners, MAE (capturing global anatomical structure via reconstruction) and DINO (capturing local semantic alignment), producing a unified embedding suitable for diverse clinical tasks. At the adaptation level, we implement multi-stage fine-tuning (MSFT), intermediate supervised fine-tuning (ISFT), and parameter-efficient fine-tuning (PEFT) to strengthen domain knowledge while preserving generalization and robustness across institutions. At the inference level, we design a Transformer-style decoder that supports a unified multi-task pipeline spanning detection, segmentation, and classification, enabling interpretable outputs and alignment with clinical reasoning. Validation follows a progressive cross-domain strategy: we start with fundus imaging and expand to visually challenging modalities such as gastroscopy and echo, in close collaboration with clinical teams at Chang Gung Memorial Hospital. Deliverables include reproducible training/evaluation assets and a cross-domain medical imaging AI toolkit to support model sharing, clinical translation, and trustworthy deployment within Taiwan’s medical AI ecosystem.